bims-librar Biomed News
on Biomedical librarianship
Issue of 2026–10–04
23 papers selected by
Thomas Krichel, Open Library Society



  1. Nurse Educ. 2026 Oct 01.
       BACKGROUND: Librarian expertise is an integral component of teaching, learning, and scholarship in nursing education; however, these collaborative partnerships remain underutilized.
    PURPOSE: To identify student and faculty learning opportunities that emerge through collaboration, this integrative review examined literature describing partnerships between nursing faculty and academic librarians.
    METHODS: An integrative review was conducted using Cumulative Index to Nursing and Allied Health Literature, MEDLINE, Library and Information Science Source, Google Scholar, and Primo. Data were synthesized using a thematic analysis approach.
    RESULTS: Included articles highlight diverse collaborative efforts that support nursing education and promote faculty development. We identified 3 themes: student-focused learning; faculty-focused learning and development; and varied and creative collaborations between librarians and nursing faculty.
    CONCLUSIONS: Findings demonstrate that academic librarians make important contributions to nursing education through curriculum development, information literacy instruction, faculty scholarship, and innovative educational partnerships. Leveraging librarian expertise can enrich nursing curricula and strengthen faculty development initiatives.
    Keywords:  DNP projects; embedded librarian model; evidence-based practice (EBP); open education resources (OER); scholarly communication
    DOI:  https://doi.org/10.1097/NNE.0000000000002353
  2. Med Ref Serv Q. 2026 Sep 30. 1-13
      The authors present a case describing the role and activities of an embedded medical librarian in a professional development program designed to strengthen medical and health science educators' clinical teaching skills. The embedded librarian supported the program curriculum by identifying and sharing resources on educational practices and contributing to curricular content and methods. The program resulted in faculty engagement with the librarian and perceptions of enhanced quality and efficiency in program development. This project highlights that librarians' expertise makes them valuable collaborators and aligns with increasing interprofessional collaborations including librarians in medical and health education.
    Keywords:  Embedded librarianship; case report; health educators; interprofessional collaboration; pedagogy; professional development
    DOI:  https://doi.org/10.1080/02763869.2026.2737083
  3. PLoS One. 2026 ;21(10): e0351352
      Qualitative research has shown public libraries to be well-loved institutions. But, as they are increasingly under attack in the US, quantitatively understanding the drivers of public voting support can help policy makers establish allocation priorities. Using novel voter response datasets (2010-2022) at the library system-year level nationally (N = 982), and at the precinct-year level within Washington state (N = 5,380), referendum modeling approaches are used to investigate local preferences for public libraries. Further, this study probes environmental condition effects on preferences. Results indicate that with precinct-year level estimations, voter preferences for facility access overwhelm other effects. Where a precinct contains a library building, the probability of a Yes vote increases 13%. Library support decreases 4% for every additional mile traveled to a location and is sensitive to extreme heat events. Taken together, the results demonstrate significant empirical support for libraries as critical infrastructure.
    DOI:  https://doi.org/10.1371/journal.pone.0351352
  4. Med Ref Serv Q. 2026 Sep 29. 1-9
      FaceBase serves as a premier database for human and animal research datasets within the craniofacial domains. Grounded in the principles of open data sharing, the platform not only houses critical research but also guides users on contributing their own data, creating data management plans (DMPs), and citing hosted datasets. This column explores the website's layout, search functionalities, and key use cases for researchers and information professionals.
    Keywords:  FaceBase; NIDCR; Online database; craniofacial; repository; research; review
    DOI:  https://doi.org/10.1080/02763869.2026.2739598
  5. J Med Internet Res. 2026 Sep 29. 28 e90335
       Background: Internet search engines serve as primary gateways to cancer information; yet, the commercialization of health content within organic search results remains understudied. While covert promotional content-such as native advertising and stealth marketing-has been documented in various contexts, systematic comparisons across structurally divergent search platforms are lacking.
    Objective: This study examined the prevalence, distribution, and information quality characteristics of covert promotional cancer-related content across Naver and Google, South Korea's 2 dominant search engines, which have fundamentally different platform architectures.
    Methods: A 2-phase cross-sectional content analysis was conducted. Phase 1 used natural language processing to identify 34 cancer-related keywords from 1400 preliminary posts. Phase 2 systematically collected 5848 posts in October 2023, yielding 919 unique posts (598 from Naver and 321 from Google) that covered 7 major cancer types, collectively accounting for over 70% of Korean cancer incidence. Two trained coders analyzed promotional status, intensity, institutional sources, and information quality indicators (citation practices, information depth, and source attribution), with intercoder reliability exceeding κ=0.80. Chi-square tests were used to examine associations between platform and content characteristics across cancer type.
    Results: Covert promotional content appeared in 48.6% (447/919) of analyzed posts, with a significantly higher prevalence on Google (174/321, 54.2%) than on Naver (273/598, 45.7%; χ²1=5.78; P=.02). Platform differences were pronounced. Naver promotional posts predominantly originated from blogs (262/273, 96.0%) and exhibited full promotional intensity (126/242, 52.1%), while Google posts primarily came from hospital websites (141/174, 81.0%) with simple institutional identification (52/90, 57.8%). Institutional source distribution varied significantly by platform (χ²5=209.642; P<.001). Traditional medicine institutions dominated Naver (119/120, 99.2%), whereas university-affiliated hospitals predominated on Google (96/113, 85.0%). Information quality also differed substantially. Indirect citation was more common on Google (142/174, 81.6%) than on Naver (160/273, 58.6%; χ²1=25.653; P<.001), while comparative informational depth was higher on Google (97/174, 55.7%) versus Naver (53/273, 19.4%; χ²2=64.683; P<.001).
    Conclusions: Covert promotional cancer content is pervasive in Korean search results, with platform architecture systematically shaping promotional patterns, institutional sources, and information quality rather than reflecting deliberate marketing strategies. These findings underscore the need for platform-sensitive regulation and enhanced digital health literacy to protect vulnerable cancer information seekers from commercial exploitation embedded within ostensibly neutral search environments.
    Keywords:  covert promotion; digital cancer information; digital health literacy; health information quality; medical commercialization; native advertising; pinkwashing
    DOI:  https://doi.org/10.2196/90335
  6. Cutan Ocul Toxicol. 2026 Sep 30. 1-12
       BACKGROUND: Hidradenitis suppurativa (HS) imposes a substantial biopsychosocial burden. As patients increasingly consult large language models (LLMs) for health information, their capacity to provide clinically appropriate, empathic, readable and exposome-aware counseling requires evaluation.
    OBJECTIVES: To benchmark LLM-generated responses to simulated HS consultations for clinical performance, exposome awareness, quality-of-life (QoL) coverage and readability.
    METHODS: In this blinded cross-sectional study, five LLMs (Claude-4.5 Sonnet, Gemini-3.0 Pro, ChatGPT-5.2, DeepSeek-V3.2 and Llama-4.0 Maverick 17B) generated zero-shot, single-turn responses to 47 clinically and socially complex HS cases. Three independent dermatologists rated responses across six domains: clinical accuracy, understandability, shared decision-making, actionability, exposome awareness and clinical empathy. The primary outcome was the mean six-domain Likert score, rescaled to a 20-100 normalized score. Secondary outcomes included domain scores, QoL coverage, readability and response length.
    RESULTS: Overall, 235 responses yielded 705 dermatologist evaluations (ICC[3,k] = 0.814; 95% CI, 0.797-0.831). Claude achieved the highest normalized score (87.00), followed by Gemini (81.25), ChatGPT (79.22), DeepSeek (79.20) and Llama (70.61). Clinical empathy (91.09) and understandability (90.16) scored highest, and exposome awareness lowest (52.94; p < 0.0001). Sleep, environmental factors and work/academic performance were rarely addressed. Readability was more demanding than recommended for patient education materials.
    CONCLUSIONS: LLMs provide generally empathic and clinically plausible HS counseling but show gaps in exposome awareness, QoL integration and health-literacy-sensitive communication. Within the constraints of single, zero-shot responses to simulated scenarios, these findings support a cautious, clinician-supervised adjunctive role and the need for dermatology-specific optimization before unsupervised patient-facing use.
    Keywords:  Health literacy; Hidradenitis suppurativa; Large language models; Patient-centered care; Psychosocial burden
    DOI:  https://doi.org/10.1080/15569527.2026.2740528
  7. JMIR Cancer. 2026 09 29. 12 e76471
       BACKGROUND: Patients with newly diagnosed gynecologic cancers often seek information online, but the quality of available resources may be inconsistent. Although GPT-4 may offer an alternative to traditional internet search engines, its performance remains largely under-studied in gynecologic oncology.
    OBJECTIVE: This study aimed to compare the completeness, accuracy, and reference quality of responses generated by GPT-4 with those generated by Google in clinical scenarios involving a new diagnosis of a gynecologic cancer.
    METHODS: Clinical scenarios representing early- and advanced-stage endometrial, ovarian, and cervical cancers were developed by gynecologic oncologists using publicly available patient education materials. Each scenario included 4 standardized questions addressing etiology, prognosis, treatment, and treatment efficacy. GPT-4 and Google were queried for each question, with new sessions for GPT-4 and private browsing for Google to minimize bias. Responses were independently evaluated by 4 gynecologic oncology experts who were blinded to each other's ratings. Accuracy was scored on a 6-point Likert scale; completeness and reference quality were scored on 3-point scales. Reference quality was categorized as low (commercial), medium (institutional or government), or high (peer reviewed). Optional free-text reviewer comments were collected and summarized descriptively to provide context for the quantitative findings. Descriptive statistics and Wilcoxon signed-rank tests were used for analysis.
    RESULTS: Across 6 clinical scenarios and 21 standardized questions (N=84 total responses), GPT-4 outperformed Google across all evaluated domains. The median accuracy score was 6.00 (IQR 5.00-6.00) for GPT-4 and 5.00 (IQR 4.00-6.00) for Google (P=.04). The median completeness score was 3.00 (IQR 3.00-3.00) for GPT-4 and 2.00 (IQR 1.00-3.00) for Google (P=.009). Reference quality was also higher for GPT-4, with a median score of 3.00 (IQR 3.00-3.00) compared to 2.00 (IQR 2.00-2.00) for Google (P=.009). Reviewer comments noted that GPT-4 provided more accurate, comprehensive, and personalized responses, while Google returned less detailed content from general consumer health websites rather than peer-reviewed sources.
    CONCLUSIONS: GPT-4 may serve as a reliable and high-quality resource for information in gynecologic oncology. Compared to Google, it delivered more accurate and complete content, with higher-quality references. Further research is needed to assess the readability and accessibility of GPT-4-generated content across diverse patient populations.
    Keywords:  AI; artificial intelligence; cancer communication; gynecologic cancer; gynecologic neoplasms; health literacy; large language models; patient education as topic; patient information materials
    DOI:  https://doi.org/10.2196/76471
  8. Front Public Health. 2026 ;14 1962232
       Background: Large language models (LLMs) are increasingly used by the public to obtain health information, but their ability to provide reliable information for achalasia remains unclear. We therefore compared six LLMs in answering public questions about achalasia, an uncommon esophageal motility disorder, focusing on safety, accuracy, empathy, information quality and reliability, and readability.
    Methods: In this cross-sectional comparative study, 40 patient-oriented questions about achalasia were submitted once to each of six LLMs (ChatGPT 5.5, Claude Opus 4.7, DeepSeek V4 Pro, Gemini 3.1 Pro Thinking, Grok-4.3, and Qwen3-Max), between May 7 and May 12, 2026. Three gastroenterologists independently assessed 240 responses for safety, accuracy, empathy, information quality, and readability using DISCERN, EQIP, the JAMA benchmark criteria assessing authorship, attribution, disclosure, and currency, and the Global Quality Score (GQS). Readability was evaluated using six established readability indices.
    Results: Overall, 27 of 240 responses (11.3%) were classified as potentially unsafe, with proportions ranging from 7.5 to 15.0% across models; the overall between-model difference in safety was not statistically significant. Accuracy differed significantly among models, although the effect size was small. Larger differences were observed in empathy, information reliability and quality, and readability. ChatGPT 5.5 and Gemini 3.1 Pro Thinking achieved higher empathy scores, whereas Claude Opus 4.7 showed greater reading difficulty. JAMA benchmark scores were low overall, indicating limited source transparency.
    Conclusion: Current LLMs provided generally accurate and often useful answers to public questions about achalasia. However, some differences were found in safety, empathy, transparency, information quality, and readability. Public-facing LLM responses to public questions about achalasia should be consistent with guidelines, clearly cite sources, communicate with empathy, and use plain language.
    Keywords:  achalasia; large language models; patient education; readability; safety
    DOI:  https://doi.org/10.3389/fpubh.2026.1962232
  9. Work. 2026 Sep 30. 10519815261491758
      BackgroundOffice ergonomics plays a critical role in preventing work-related musculoskeletal disorders and promoting employee well-being. As large language models such as ChatGPT are increasingly used to obtain health-related information, evaluating the quality, reliability, readability, and practical applicability of artificial intelligence (AI)-generated ergonomics content has become increasingly important.ObjectiveTo examine the reliability, quality, accuracy, readability, and clinical applicability levels of the responses given by ChatGPT-5.2 to office ergonomics search terms identified by Google Trends, through expert evaluation.MethodsNineteen keywords were selected from the terms obtained by searching "office ergonomics" on Google Trends on January 7, 2026. Responses were generated individually in incognito mode using ChatGPT-5.2 (December 2025 version) and assessed by two experts (physiotherapist, forensic medicine specialist) using Journal of the American Medical Association Benchmarking Criteria (JAMA), Global Quality Score (GQS), Modified DISCERN Score (MDS), accuracy scale, Flesch-Kincaid Grade Level (FKGL) and Flesch Reading Ease Score (FRES) and Patient Education Material Assessment Tool for printed materials (PEMAT-P); inter-rater agreement was analyzed using Intraclass Correlation Coefficients (ICC) and Cronbach's alpha.ResultsThe mean GQS score was 3.00 ± 0.74; the MDS score was 2.89 ± 0.31; and the accuracy score was 3.58 ± 0.51. The FKGL level was 7.66 ± 2.76 and the FRES score was 54.17 ± 14.91, with 47.4% of the content requiring university-level readability. The PEMAT-P comprehensibility level was 82.53 ± 11.03, and applicability was 49.21 ± 20.76. Inter-rater agreement was good-to-very high (ICC = 0.779-0.974).ConclusionsAlthough ChatGPT-5.2 generally provided acceptable levels of accuracy and content quality in this exploratory analysis, the findings should be interpreted in light of the study's methodological limitations, including the restricted keyword set, single-day assessment, and evaluation of a single model version.
    Keywords:  occupational health; office ergonomics; quality; readability; reproducibility of results; search engine
    DOI:  https://doi.org/10.1177/10519815261491758
  10. Front Public Health. 2026 ;14 1955701
       Background: Large language model (LLM) chatbots are increasingly used to obtain health information. However, fluent and clinically plausible responses may still contain safety-relevant omissions, inadequate source attribution and disclosure, or difficult-to-read text.
    Objective: To evaluate the safety, accuracy, empathy, information quality, response-level source attribution and disclosure, overall quality, and readability of five LLM chatbot interfaces answering lay-oriented questions about herpes zoster.
    Methods: This cross-sectional evaluation used 46 standardised English-language questions across seven herpes zoster domains. Each question was submitted once to ChatGPT-5.5 Instant, DeepSeek-V4-Pro, Doubao-Seed-2.0-Pro, Gemini 3.5 Flash, and Qwen3.7-Plus under consumer-access conditions on June 2-3, 2026. Five dermatologists independently assessed safety, accuracy, empathy, DISCERN, Ensuring Quality Information for Patients (EQIP), Journal of the American Medical Association (JAMA) benchmark criteria, and the Global Quality Score (GQS). Six readability indices were calculated. Paired comparisons and inter-rater agreement were evaluated using prespecified statistical methods.
    Results: The five interfaces generated 230 complete responses. Under the study conditions, 66 responses (28.7%) were classified as potentially unsafe using the prespecified ≥3/5-rater majority threshold, with no significant difference between interfaces (Cochran's Q = 1.661, df = 4, p = 0.798). Sensitivity analyses using alternative thresholds showed consistent results. Accuracy and empathy differed across interfaces (both p < 0.001; Kendall's W = 0.127 and 0.211, respectively), although effect sizes were small. Information-quality and overall-quality measures also showed outcome-specific differences. All six readability indices differed across interfaces (all p < 0.001); DeepSeek-V4-Pro generally produced text estimated to be easier to read, whereas Gemini 3.5 Flash produced text estimated to be more difficult to read. Safety agreement was substantial (Fleiss' κ = 0.682), and ICC (2,1) values for other manually rated outcomes ranged from 0.832 to 0.889.
    Conclusion: In this standardised benchmark, LLM chatbot interfaces provided generally favourable accuracy and information-quality scores but showed clinically relevant safety limitations, sparse source attribution and disclosure, and readability challenges. No interface consistently performed best across all outcomes. These findings represent single first responses obtained under the tested conditions and do not establish response stability across repeated queries. LLM-generated herpes zoster information should therefore be interpreted cautiously and should not replace individualised professional assessment or professionally reviewed patient information.
    Keywords:  Shingles; chatbots; herpes zoster; information quality; large language models; patient education; readability; safety
    DOI:  https://doi.org/10.3389/fpubh.2026.1955701
  11. Can J Surg. 2026 Sep-Oct;69(5):69(5): E402-E407
       BACKGROUND: Patients have been resorting to online content as their main source of medical knowledge, notably on orthopedic interventions. However, Web-based research findings present a reverse correlation between quality of content and popularity. We sought to evaluate whether ChatGPT could provide an alternative and safe source of medical information for arthroplasty patients.
    METHODS: We gave 5 commonly Googled questions related to hip and knee arthroplasty to board-certified arthroplasty surgeons, fellows, and orthopedic surgery residents, as well as to ChatGPT 4.0. We anonymized all answers, which were then analyzed by an independent board-certified arthroplasty surgeon. We scored the answers for accuracy of content (6-point Likert scale) and completeness (3-point Likert scale), then compared the performance of all groups.
    RESULTS: Among human responders, the mean accuracy grade was 75% (standard deviation [SD] 17%), with a mean completeness grade of 69% (SD 22%). The fellows represented the strongest subgroup, with 75% of their answers scored above 5/6 for accuracy (mean 82%, SD 16%). We found that all human-generated answers had a statistically significant correlation between accuracy and completeness. ChatGPT had a mean accuracy grade of 93% (SD 12%), with a mean completeness of 93% (SD 14%).
    CONCLUSION: ChatGPT appears to be a safe tool for patients to access general arthroplasty information online. It outperformed all human responders on both accuracy and completeness of answers. The strongest human responder group was the arthroplasty fellows. Further work is required to clarify the tool's performance against other easily accessible online patient information sources.
    DOI:  https://doi.org/10.1503/cjs.018325
  12. Cureus. 2026 Aug;18(8): e115285
      Introduction Clear, accurate, and accessible patient education is central to informed decision-making and high-quality cardiovascular care. Large language models (LLMs) are increasingly being used to generate health information, offering the potential to rapidly produce patient education materials. However, concerns remain regarding the readability, quality, clinical accuracy, and adherence to evidence-based recommendations of AI-generated content. This study compared the performance of ChatGPT (OpenAI, San Francisco, CA), Claude (Anthropic PBC, San Francisco, CA), and DeepSeek (DeepSeek Artificial Intelligence Co., Ltd., Hangzhou, China) in generating patient education leaflets for three commonly performed cardiac imaging procedures. Methodology A cross-sectional study was conducted using standardized prompts to generate patient education leaflets for cardiac magnetic resonance imaging (CMR), coronary computed tomography angiography (CTCA), and invasive coronary angiography (ICA) using ChatGPT, Claude, and DeepSeek. Nine leaflets were evaluated for readability using Flesch-Kincaid Grade Level, Gunning Fog Index, Simple Measure of Gobbledygook (SMOG), Flesch Reading Ease, word count, and sentence count. Information quality was assessed using the modified DISCERN (mDISCERN) instrument. Guideline adherence was assessed using investigator-developed checklists based on recommendations from relevant professional societies and patient education resources. Expert clinical assessment was independently performed by two reviewers using a structured 16-point scoring system. Comparisons among the three models were performed using the Kruskal-Wallis test. Results DeepSeek demonstrated the most favorable overall readability profile, with the highest Flesch Reading Ease score (70.73 ± 1.74) compared with ChatGPT (50.70 ± 7.88) and Claude (60.93 ± 5.51; H(2) = 7.20, p = 0.027). DeepSeek achieved the highest mean mDISCERN score (3.83 ± 0.29), followed by ChatGPT (3.50 ± 0.50) and Claude (2.67 ± 0.29), although the difference was not statistically significant (H(2) = 5.593, p = 0.061). ChatGPT demonstrated the highest mean guideline adherence (91.54 ± 3.39%), followed by DeepSeek (89.33 ± 3.41%) and Claude (88.07 ± 6.48%; H(2) = 0.707, p = 0.702). DeepSeek achieved the highest expert clinical assessment score (16.00 ± 0.00), followed by ChatGPT (15.67 ± 0.58) and Claude (14.33 ± 0.58), with no statistically significant difference (H(2) = 5.394, p = 0.067). A significant difference was also observed in total word count (H(2) = 7.20, p = 0.027). Conclusions All three LLMs generated high-quality patient education leaflets for common cardiac imaging procedures, with each model demonstrating distinct strengths across different evaluation domains. DeepSeek showed the most favorable readability and achieved the highest information quality and expert clinical assessment scores, whereas ChatGPT demonstrated the greatest guideline adherence. However, all three models produced content above recommended patient health-literacy standards. LLMs may therefore serve as valuable clinician-assisted tools for developing cardiac imaging education materials, but expert review and readability optimization remain essential before clinical implementation.
    Keywords:  artificial intelligence in radiology; cardiac imaging modalities; cardiac imaging-mri; health information literacy; large language models (llms); patient education material
    DOI:  https://doi.org/10.7759/cureus.115285
  13. Am J Otolaryngol. 2026 Sep 23. pii: S0196-0709(26)00164-X. [Epub ahead of print]47(6): 104948
       OBJECTIVE: To compare the quality of Google, ChatGPT, and providers in answering common questions regarding vestibular schwannoma.
    METHODS: Search engine frequency data was used to generate ten common patient questions regarding the etiology, diagnosis, and management of vestibular schwannoma. Answers to these questions were generated from compiled Google search, ChatGPT-4, and provider responses. Response quality was graded by three blinded, independent, board-certified otolaryngologists using the DISCERN instrument. Readability scores were calculated. Inter-rater agreement was determined using Fleiss's kappa.
    RESULTS: Inter-rater agreement was high with a Fleiss's kappa score of 0.73. Total DISCERN scores among Google, ChatGPT, and expert response were 23.7, 30.4, and 28.4, respectively, with a standard error of 0.85. Average word count for ChatGPT (214 ± 98.2) responses were significantly longer (p < 0.05) than both Google (68.8 ± 72.3) and provider responses (108 ± 91.7) responses. ChatGPT had the highest readability score with a Flesch Reading Ease score of 33.1 ± 10.9, followed by Google (30.3 ± 20.7) and provider response (15.8 ± 8.70), with ChatGPT and Google performing significantly better than providers (p < 0.05).
    CONCLUSION: The overall quality of ChatGPT responses to common questions about vestibular schwannoma was comparable to that of providers. Moreover, these responses were calculated to represent the highest readability levels when compared to both providers and Google search derived information.
    LEVEL OF EVIDENCE: III.
    Keywords:  Artificial intelligence; Lateral skull base; Vestibular schwannoma
    DOI:  https://doi.org/10.1016/j.amjoto.2026.104948
  14. Am Surg. 2026 Sep 29. 31348261494138
      BackgroundThe proliferation of Large Language Models necessitates evaluating their performance in communication accessibility. This study compares the linguistic architecture, readability, and automated psycholinguistic text properties of AI-generated breast cancer materials across varying clinical complexities.MethodsTwenty standardized questions across diagnosis and treatment categories were queried through three independent iterations. Readability was analyzed using Flesch-Kincaid Grade Level, SMOG, and Gunning Fog indices. Linguistic dimensions were evaluated via LIWC-22 software as computational style metrics. Statistical analyses utilized Wilcoxon, Mann-Whitney U, Spearman's correlation, and with False Discovery Rate (FDR) corrections. Cross-iteration stability was assessed with the intraclass correlation coefficient (ICC(2,1), two-way random-effects, and absolute agreement).ResultsGemini had better readability than ChatGPT (FKGL 9.23 vs 10.07, q = 0.022); both exceeded the 6th-grade threshold. Treatment queries increased linguistic complexity for both (ChatGPT FKGL q = 0.012; Gemini FKGL q = 0.047). Gemini scored higher on Analytic (q = 0.001) and Authentic (q = 0.021) dimensions; ChatGPT scored higher on I-words (q = 0.001) and Cognitive Processes (q < 0.001). Treatment-specific shifts in Gemini Positive Tone (raw P = .041, q = 0.367) and ChatGPT Authentic (raw P = .025, q = 0.114) lost significance after FDR adjustment. ChatGPT showed a negative correlation between readability and Authentic score (ρ = -0.63, q = 0.026). Cross-iteration ICC values were poor-to-moderate for both models (mean ICC: ChatGPT = 0.58, Gemini = 0.59).ConclusionsBoth LLMs create practical information gaps for patients with limited health literacy, particularly when addressing complex treatment pathways. Finer-grained psycholinguistic claims did not withstand statistical correction, and cross-iteration stability was modest. These automated analyses highlight structural limitations of current LLMs, reinforcing the necessity of clinician-guided curation before deploying them for patient education.
    Keywords:  ChatGPT; Gemini; breast cancer; digital empathy; health literacy; readability
    DOI:  https://doi.org/10.1177/00031348261494138
  15. Dental Press J Orthod. 2026 ;pii: S2176-94512026000400601. [Epub ahead of print]31(4): e262657
       OBJECTIVE: This study aimed to evaluate, using the DISCERN questionnaire, the quality, reliability, and decision-support capacity of aligner-related information published on orthodontists' websites.
    METHODS: A cross-sectional study was conducted in July 2024. Google (www.google.com.br) was searched using Portuguese-language keywords ("tratamento ortodôntico", "alinhadores", and "ortodontia") in incognito mode from a Brazilian IP address. The first 130 non-sponsored websites were screened; 113 met the inclusion criteria. Two calibrated orthodontists independently evaluated each website using the validated Brazilian Portuguese version of the DISCERN instrument. Inter-rater reliability was assessed using the Intraclass Correlation Coefficient (ICC). Descriptive statistics and Spearman rank-order correlations were calculated using SPSS v.25.
    RESULTS: The mean overall DISCERN score was 26.3 (SD = 5.8; range 15-52) out of a maximum of 75. The ICC was 0.87 (95% CI = 0.82-0.91), indicating good-to-excellent inter-rater reliability. Among the 113 websites, 54.3% were rated as very poor, 30.7% as poor, 11.5% as fair, 3.5% as good, and none as excellent. The lowest-scoring items were Q4 (sources of information, 99.1% very poor or poor), Q5 (publication date, 87.6%), Q8 (areas of uncertainty, 95.6%), Q11 (treatment risks, 98.2%), and Q12 (consequences of no treatment, 100%).
    CONCLUSIONS: The quality of clear aligner information on Brazilian orthodontists' websites is critically poor. Orthodontists should provide evidence-based, transparent, and balanced content - including treatment risks, alternatives, and references - to support informed patient decision-making.
    DOI:  https://doi.org/10.1590/2177-6709.31.3.e262657.oar
  16. Kidney360. 2026 Sep 28.
       BACKGROUND: Clear information is essential for patients navigating kidney transplantation, yet patient-facing materials may be difficult to understand. Patient education guidelines recommend reading levels near a sixth-grade level to support comprehension across diverse populations. This study evaluated readability of patient-directed content on United States kidney transplant center websites.
    METHODS: Patient-facing webpages from United States kidney transplant programs were identified using a previously assembled national dataset. Website text was extracted and processed to isolate patient-directed educational content, excluding non-English pages and pages with fewer than five sentences. Readability was assessed using validated indices, including the Flesch-Kincaid Grade Level, Simple Measure of Gobbledygook (SMOG), Gunning Fog Index, and a pretrained transformer-based readability model. Summary statistics characterized reading difficulty and proportion of webpages exceeding recommended levels.
    RESULTS: A total of 92,461 English-language webpages from 194 transplant center websites were included. Median readability exceeded recommended levels across measures. At the webpage level, the median SMOG score was 14.69 (interquartile range, 13.01-16.63); >99% of webpages exceeded the recommended sixth-grade level, 98% exceeded tenth-grade level, 86% exceeded twelfth-grade level, and 61% exceeded fourteenth-grade level. Among the 144 eligible hosts, the median SMOG score was 14.92 (interquartile range, 14.11-15.45). Similar patterns were observed across the Flesch-Kincaid Grade Level, Gunning Fog Index, and transformer-based readability model, demonstrating consistent findings across methods.
    CONCLUSIONS: Patient-facing kidney transplant center websites were written at reading levels substantially exceeding recommended standards. This widespread, modifiable barrier in transplant center communication may limit patients' ability to understand transplant-related information.
    DOI:  https://doi.org/10.34067/KID.0000001401
  17. Health Promot Pract. 2026 Oct 01. 15248399261488364
      Digital health information is an important source of preventive health advice during pregnancy. Its public health value depends on whether advice is guideline-concordant, safe, transparent and understandable. This study aimed to assess the guideline concordance, content coverage, safety information, transparency and readability of Australian consumer-facing webpages providing advice about physical activity during pregnancy. We conducted a systematic, cross-sectional audit of English-language, freely accessible webpages affiliated with Australian organisations that provided substantive advice about physical activity or exercise during pregnancy. Searches were undertaken between June and July 2025, with data collection continuing until September 2025. Data were extracted using a structured checklist mapped to Australian pregnancy physical activity guidelines. Transparency was assessed using Journal of the American Medical Association (JAMA) benchmarks and readability using Flesch Reading Ease (FRE), Flesch-Kincaid Grade Level (FKGL) and Simplified Measure of Gobbledygook (SMOG). Sixty webpages were included. Most encouraged physical activity during pregnancy, but practical guidance and essential safety information were often incomplete. Fourteen webpages (23.3%) reported weekly exercise volume consistent with current guidelines. Fourteen (23.3%) included no warning signs, relative contraindications or absolute contraindications, and only four (6.7%) covered all three safety domains. Transparency was limited: authorship was reported by 12 webpages (20.0%), attribution by 16 (26.7%), disclosure by three (5.0%), and currency by none. Readability exceeded recommended levels for consumer materials; no webpage met the SMOG threshold of 8 or below. Australian online information about physical activity during pregnancy is often not fit for purpose as public-facing preventive health guidance. Minimum standards for safety content, transparency and readability may improve the quality and accessibility of digital antenatal health information.
    Keywords:  digital health; health literacy; health promotion; online health information; physical activity; pregnancy; readability
    DOI:  https://doi.org/10.1177/15248399261488364
  18. Pregnancy (Hoboken). 2026 Nov;2(6): e70476
       Introduction: In obstetrics, patients with insufficient health literacy have a higher likelihood of adverse perinatal outcomes. The American Medical Association (AMA) recommends all patient-facing health information be written at or below a sixth-grade reading level. The reading level of standardized obstetric patient education materials embedded into Epic via "Patient Pass/Exit Care" package published by Elsevier is unknown. Since the Spanish versions of the materials have not been back-translated, the readability of those materials also requires further assessment. In this study, our aims were to (1) characterize overall readability of Epic's obstetric patient education documents, (2) assess differences in readability between English and Spanish materials, and (3) understand differences in readability by subject material, stratified by language.
    Methods: Using the "Clinical References" tab in Epic, obstetric patient education materials were extracted in both English and Spanish. Materials were evaluated using metrics for assessing readability that have been validated in English and/or Spanish: Flesch-Kincaid Grade Level (FKGL), Fernandez-Heurta (FH), and Simple Measure of Gobbledygook (SMOG). Education was coded by content category. Paired t-tests were conducted to evaluate differences in readability scores (FKGL vs. FH) and SMOG scores across language pairs. Mean differences (95% confidence intervals [CIs]) and one-way ANOVA with Bonferroni-adjusted post hoc tests assessed language readability gaps across clinical categories.
    Results: Of 235 total texts analyzed, grade level for English was 6.2 (SD 1.0) using the FKGL metric versus 4.2 (SD 0.5) for Spanish content that used the validated FH. Overall, the English documents were 2.0 grade levels higher than the corresponding Spanish documents (95% CI, 1.8-2.1). SMOG mean grade level for English materials was 9.7 (SD 0.9) versus 6.2 (SD 0.6) for Spanish materials (p < 0.01). Across all perinatal topics, English texts scored significantly higher than Spanish texts on both SMOG and language-specific readability metrics. A total of 41.2% of Spanish materials were at or below a 6th grade level using SMOG, but none of the English materials were.
    Conclusion: Existing obstetric patient education materials embedded in the Epic electronic medical record (EMR) do not meet AMA standards for readability and comprehension. Differences between English and Spanish materials highlight a potential for disparities in patient understanding of their plans of care.
    Keywords:  OB/GYN; health literacy; patient education; readability
    DOI:  https://doi.org/10.1002/pmf2.70476
  19. J Patient Cent Res Rev. 2026 ;13(3): 108-114
       Purpose: Avascular necrosis (AVN) is a progressive pathology affecting pediatric and adult patients. Given the risks of joint collapse and surgery, patients frequently research their condition online. The average reading level in the United States is approximately the eighth-grade level. The National Institutes of Health (NIH) recommend that patient materials be written at or below a sixth-grade level. This study aimed to assess the readability and quality of online AVN resources.
    Methods: A Google search using the terms "avascular necrosis" and "osteonecrosis" was conducted, and first-page results were aggregated. Flesch-Kincaid Grade Level, Gunning Fog Index, and Flesch Reading Ease tests were performed to assess the reading levels of these resources. DISCERN scores were determined by two reviewers.
    Results: Sixteen original general educational articles emerged. Of these, one article (6.25%) was at or below the NIH-recommended reading level. Five articles (31.25%) were at or below the average reading level in the United States. The mean Flesch-Kincaid Grade Level was 9.3 ± 1.8, the mean Gunning Fog Index was 10.6 ± 2.3, and the mean Flesch Reading ease was 46.7 ± 11.3. The mean DISCERN score was 59.7 ± 7.8, which is classified as "good quality."
    Conclusions: The readability of first-page Google results for AVN-related patient resources is at an education level higher than that recommended by the NIH and higher than the average national reading level. This study highlights the necessity for developing high-quality, accessible patient educational materials about AVN that are tailored to an appropriate level for patients and their families.
    Keywords:  avascular necrosis; google; osteonecrosis; patient education; readability
    DOI:  https://doi.org/10.17294/2330-0698.2229
  20. JMIR Infodemiology. 2026 Sep 30. 6 e98768
       Background: Patients with seizure clusters and their families often use YouTube videos as resources for information about the illness and how to use concomitant rescue medications; however, many videos include misinformation or incorrect information that could potentially cause harm.
    Objective: This study aims to examine YouTube videos intended as educational and instructional resources for seizure clusters and treatment options posted since the approval of intranasal benzodiazepine rescue medications in 2019 and to assess their accuracy and quality using a 3-item rating scale.
    Methods: We searched YouTube for 10 specific terms related to seizure clusters and treatment options and analyzed videos from the first 20 results for each to mimic internet user behavior when choosing results with search engines. Videos were rated on 3 criteria: accurate seizure cluster definition, appropriate US Food and Drug Administration (FDA)-approved treatment recommendation, and correct visual demonstration of rescue medication use. A single author rated the videos based on the presence of specific subcriteria of all 3 primary criteria in each video. Score assignment was based on the number of present subcriteria.
    Results: Seizure clusters and trade names generated the highest number of videos meeting criteria; other search terms, including nonproprietary names, garnered <8 videos each. Only 19.5% (n=39) of the 200 collected videos provided accurate descriptions of seizure clusters, while 61.0% (n=122) offered no definition, 37.0% (n=74) recommended at least 1 FDA-approved rescue medication, and 37.5% (n=75) offered no rescue therapy at all; 26.5% (n=53) provided a full demonstration of rescue medicine, while 62.5% (n=125) did not. Only 4.5% (n=9) of videos included all subcriteria of the 3 primary criteria. Search results for specific trade names yielded more results, while less specific terminology returned unrelated or unhelpful videos.
    Conclusions: Accurate YouTube videos of seizure cluster management are limited and represent a knowledge gap that should be addressed to avoid patient harm. To facilitate better navigation of educational and instructional resources, clinicians should have a ready list of videos available to give to patients or provide a list of search terms that are known to produce accurate results to facilitate better searches. Reputable health care sources could also produce brief, helpful videos of each rescue therapy with search terms that direct patients to their content.
    Keywords:  YouTube; caregivers; diazepam; education; intranasal; midazolam; rescue medication; seizure clusters
    DOI:  https://doi.org/10.2196/98768
  21. J Maxillofac Oral Surg. 2026 Oct;25(5): 1507-1517
       Background: Cleft lip and palate (CLP) is among the most common congenital anomalies, necessitating complex, multidisciplinary care from infancy through adolescence. Caregivers often turn to online resources, particularly YouTube, for supplemental guidance on presurgical management, feeding techniques, and long-term care. However, the reliability and educational value of such freely available video content remain uncertain.
    Methods: A systematic literature search was conducted in PubMed/MEDLINE, Scopus, Cochrane, and relevant grey literature from 2014 to 2026. Eligible studies were cross-sectional, observational and qualitative study analyses of YouTube videos related to cleft lip and/or palate, reporting content quality using standardized tools such as the Global Quality Score (GQS), DISCERN, PEMAT-AV, VIQI, MQ-VET, or custom indices.
    Results: A total of eight studies that met the eligibility criteria were included in the systematic review. Seven YouTube video-analysis studies that involved quantifiable methods contributed extracted data from 373 YouTube videos, whereas one study involved qualitative methods, providing narratives on the preferences of parents and their unmet educational needs. GQS was reported in two studies and analyzed exploratorily, resulting in a pooled average of 2.50/5 (95% CI 2.11-2.90; I² = 68.8%), which shows the presence of moderate to low educational quality.
    Conclusions: The quality and comprehensiveness of YouTube videos on cleft lip and palate remain suboptimal, with considerable variability across topics and uploader types. There is a critical need for improved, peer-reviewed, and user-centered educational resources to better support families and caregivers. Collaborative efforts among clinicians, patient advocates, and digital content creators are essential to elevate the standard of online health information in this domain.
    Supplementary Information: The online version contains supplementary material available at https://doi.org/10.1007/s12663-026-03261-9.
    Keywords:  Cleft Palate; Cleft lip; Patient education; Social media; YouTube
    DOI:  https://doi.org/10.1007/s12663-026-03261-9
  22. Front Med (Lausanne). 2026 ;13 1917426
       Objective: To evaluate the quality, reliability, and educational value of videos related to precancerous cervical lesions across four major short-video platforms in China (TikTok, Bilibili, Kwai, and Xiaohongshu), using multiple validated evaluation tools.
    Methods: A total of 1842 short videos were collected, assessed using GQS, JAMA, modified DISCERN, PEMAT, and HONcode tools. The videos were analyzed for quality, reliability, educational value, and engagement metrics (likes, comments, collections, shares).
    Results: Videos on TikTok consistently achieved the highest scores for quality, reliability, and transparency across all scoring systems. In contrast, Bilibili, Kwai, and Xiaohongshu exhibited variable performance. Videos created by medical practitioners, particularly gynecologists, scored higher compared to those uploaded by non-medical practitioners. Content focusing on basic disease knowledge had higher educational value than prevention or post-treatment videos. Audience engagement did not always correlate with higher quality.
    Conclusion: Short-video platforms may serve as valuable channels for health information dissemination. Among the sampled videos, TikTok demonstrated higher quality scores than the other evaluated platforms. Future efforts may focus on improving health information quality assurance, encouraging healthcare professional participation in content creation, and optimizing video formats for actionable health guidance.
    Keywords:  platform governance; precancerous cervical lesions; public health education; short-video platforms; video quality assessment
    DOI:  https://doi.org/10.3389/fmed.2026.1917426
  23. BMC Geriatr. 2026 Sep 08. pii: 1217. [Epub ahead of print]26(1):
       BACKGROUND: With advances in information technology and the widespread availability of online health resources, access to digital health information has become increasingly common. In this context, the ability of older adults to effectively access, evaluate, and use online health information has emerged as a critically important competence. Since the internet provides easy and often low-cost access to a wide range of information, older adults are increasingly using the internet for health-related purposes. This study aimed to examine the demographic, economic, and personal factors associated with online health information seeking behaviour among older adults in Türkiye.
    METHODS: This study used the microdata from the Household Information Technology (IT) Usage Survey conducted by the Turkish Statistical Institute (TurkStat) in 2023 and 2024. Binary logistic regression analysis was performed.
    RESULTS: Compared with older adults surveyed in 2023, those surveyed in 2024 had lower odds of online health information seeking behaviour (AOR = 0.760; 95% CI = 0.644-0.898), and men had lower odds of engaging in this behaviour than women (AOR = 0.620; 95% CI = 0.516-0.745). Higher educational attainment, particularly university education (AOR = 4.002; 95% CI = 2.405-6.660), searching for product and service information (AOR = 4.546; 95% CI = 3.783-5.462), social media use (AOR = 1.994; 95% CI = 1.646-2.415), and internet banking use (AOR = 1.649; 95% CI = 1.352-2.011) were associated with greater odds of online health information seeking behaviour.
    CONCLUSION: The findings may help health information providers and healthcare professionals better understand the online health information seeking behaviour of older adults and inform the development of digital health strategies tailored to their needs.
    Keywords:  Digital health; Household Information Technology Usage Survey; Older adults; Online health information seeking behaviour; Türkiye
    DOI:  https://doi.org/10.1186/s12877-026-08237-5