bims-librar Biomed News
on Biomedical librarianship
Issue of 2026–09–06
23 papers selected by
Thomas Krichel, Open Library Society



  1. Cochrane Evid Synth Methods. 2026 Sep;4(5): e70105
       Introduction: Climate change has major health impacts to which researchers and policymakers need to respond. The volume and multidisciplinary nature of climate-health evidence pose challenges to its comprehensive identification. Search filters are currently not available for this topic. Our aim was to empirically develop sensitivity-maximizing climate-health search filters for two interfaces of the MEDLINE database.
    Methods: Climate change impacts health via several exposure pathways involving distinct mechanisms: extreme weather events, heat stress, air quality, water quality and quantity, food supply and safety, vector distribution and ecology, and social factors. Using the relative recall method, we established a gold standard by extracting included studies of 110 evidence syntheses and classifying them into exposure pathways. From this set, we empirically derived and validated filters per pathway.
    Results: Based on 1572 primary studies indexed in PubMed, we developed filters with a sensitivity of 95%, 97% and 99% for six of the seven major climate-health exposure pathways.
    Conclusion: We designed the first empirically derived search filters for climate-health pathways, enabling sensitive retrieval despite acknowledged limitations. Our filters can be applied to different types of research questions at the intersection of climate and health.
    Keywords:  MEDLINE; climate change; evidence synthesis; information storage and retrieval; public health; search filters; systematic reviews
    DOI:  https://doi.org/10.1002/cesm.70105
  2. J Vet Sci. 2026 Jul 08.
       IMPORTANCE: The Veterinary Neurologic Disease Library has been created as an educational resource library for veterinary professionals related to breed-associated neurologic diseases in dogs.
    OBJECTIVE: To develop an open-access library that serves as a curated collection repository of peer-reviewed literature pertaining to dog breeds and neurologic diseases.
    METHODS: Dog breeds were identified using the American and United Kennel Club registered breed lists. Disease terms were collected from published texts discussing veterinary neurologic diseases. Publicly available resources were searched by inputting keywords, including dog breeds, specific neurologic diseases, clinical signs, and additional related terms. Retrieved information was stored in a Zotero web-library, and automated searches continue to be conducted monthly. Original peer-reviewed manuscripts (in English) were categorized for inclusion with an individual unique breed was identified with a neurologic disease.
    RESULTS: The library is currently active and accessible online. Four hundred breeds have been included for peer-reviewed citations of neurologic diseases. At the time of this writing, over 1,900 scholarly citations have been identified and categorized.
    CONCLUSIONS AND RELEVANCE: This centralized, open-access, repository of breed-associated neurologic diseases can be accessed to search breed and neurologic disease to assist primary care clinicians in the identification breed-associated neurologic disease in dogs and locating original source publications.
    Keywords:  Dogs; breed-related; library; neurologic disease
    DOI:  https://doi.org/10.4142/jvs.25316
  3. Indian J Palliat Care. 2026 Jul-Sep;32(3):32(3): 304-311
       Objectives: Palliative care aims to enhance the quality of life for patients facing terminal or life-threatening illnesses. Unfortunately, many patients encounter difficulties in understanding their condition, treatment processes and available care options due to limited access to clear and relevant information. This often leads to confusion, disrupted decision-making and increased anxiety. This study aimed to explore patient experiences and information needs in the context of palliative care.
    Materials and Methods: A qualitative descriptive design was employed, involving 20 participants, consisting of 10 palliative patients and 10 family members, to triangulate data sources. Participants were selected through purposive sampling based on predefined inclusion criteria. Data were collected through semi-structured face-to-face interviews and analysed using thematic analysis.
    Results: Five main themes with 10 subthemes were identified: (1) Sources of information about palliative care: information from physicians and nurses, information from family members and information from friends or community; (2) Alternative sources of health information included the internet (e.g., Google, websites) and educational videos on social media platforms (e.g., YouTube and TikTok); (3) Facilitators in accessing health information: Direct information provided by physicians and support from family members in obtaining information; (4) Barriers in understanding health information: difficulty understanding medical terminology and (5) Expectations for the Use of Digital Technology: easily accessible online health information and Digital applications or platforms for communication with healthcare professionals.
    Conclusion: Patient experiences in accessing palliative care information are shaped by interactions with doctors, family members and communities, while the internet and social media serve as additional sources. Most patients reported no difficulties due to the support of healthcare providers and their families. However, medical terminology created barriers to comprehension. Patients expressed strong expectations for hospitals to implement digital technologies to enhance access to information and continuity of care.
    Keywords:  Digital technology; Information access; Palliative care; Patient experiences
    DOI:  https://doi.org/10.25259/IJPC_351_2025
  4. J Med Internet Res. 2026 Sep 02. 28 e93944
       BACKGROUND: Generative AI (GenAI) tools powered by large language models (LLMs) are increasingly used by the public to seek health information. Unlike traditional web search, these systems generate conversational responses that may alter how users assess credibility, manage uncertainty, verify information, and decide whether to consult clinicians. As GenAI becomes more embedded in everyday health information practices, a clearer synthesis of the emerging empirical evidence is needed.
    OBJECTIVE: This scoping review mapped and synthesized empirical research on consumer and patient health information seeking using GenAI and LLM tools, with a focus on study contexts, outcome constructs, and the facilitators and barriers shaping use, reliance, and verification.
    METHODS: The review adhered to Joanna Briggs Institute guidance for scoping reviews and reported using PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews), with search reporting additionally guided by PRISMA-S (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Search Extension). We searched PubMed, Scopus, PsycINFO, Web of Science, IEEE Xplore, ACM Digital Library, Google Scholar, ERIC, EBSCO, and ProQuest for English-language studies published 2022 onward. The final updated search was conducted on January 8, 2026. Eligible studies were empirical quantitative, qualitative, or mixed methods studies examining health information seeking mediated by GenAI and LLM systems, wherein an LLM served as the interface or source for obtaining health information. Data were charted using a structured extraction form capturing study characteristics, populations, health contexts, GenAI tool types, outcomes, and factors shaping use.
    RESULTS: The review included 27 studies. GenAI was used for symptom appraisal, condition understanding, treatment options, and care navigation. Reported facilitators included convenience and clarity, particularly efficiency and access (n=8, 29.6%), comprehensibility and presentation quality (n=11, 40.7%), personalization and specificity (n=5, 18.5%), and affective or interpersonal comfort (n=5, 18.5%). Reported barriers were dominated by credibility and trust concerns (n=13, 48.1%), particularly when accuracy cues or citations were missing or difficult to interpret. Additional barriers included perceived unsuitability for complex, urgent, or emotionally charged situations (n=5, 18.5%); privacy or data security concerns (n=4, 14.8%); limited prompting skills (n=2, 7.4%); and modality or interaction constraints that hindered credibility assessment and information comparison (n=5, 18.5%). Six (22.2%) studies reported literacy-related capability was, and 5 (18.5%) reported verification-supporting features, such as visible sourcing, transcripts, and save, revisit, or share functions.
    CONCLUSIONS: This review is innovative in focusing on health information seeking as a user practice rather than on technical performance or clinical implementation alone. Unlike prior reviews, it maps how the emerging literature conceptualizes use, trust, reliance, and verification. It contributes a structured synthesis of the main facilitators, barriers, and verification-related features reported on GenAI-mediated health information seeking. In practice, the findings suggest that safer use may depend on not only model quality but also users' ability to interpret, verify, and act on AI-generated responses.
    Keywords:  LLM; PRISMA-ScR; Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews; adoption; consumer health informatics; credibility; generative AI; health information seeking; large language model; scoping review; trust; verification
    DOI:  https://doi.org/10.2196/93944
  5. JMIR AI. 2026 Aug 31. 5 e107588
      
    Keywords:  AI; LLMs; health information; large language models; patient education material; readability
    DOI:  https://doi.org/10.2196/107588
  6. J Stomatol Oral Maxillofac Surg. 2026 Sep 02. pii: S2468-7855(26)00272-7. [Epub ahead of print] 102972
       OBJECTIVE: To compare the quality and readability of responses from five generative artificial intelligence chatbot platforms to clinician-oriented questions on open temporomandibular joint (TMJ) surgery against guideline-based reference answers.
    MATERIAL AND METHODS: Forty questions across eight domains were submitted on 16 July 2026 to ChatGPT (GPT-5.5), Claude (Opus 4.8), Gemini (3.1 Pro), Grok (4) and Perplexity (Pro), each via its paid tier at default settings (200 responses). Two blinded oral and maxillofacial surgeons applied the Quality Analysis of Medical Artificial Intelligence (QAMAI) tool, the Global Quality Score (GQS) and a five-point overall quality rating using a priori key elements and written anchors. Readability was assessed with the Flesch-Kincaid Grade Level (FKGL) and Flesch Reading Ease Score (FRES); platforms were compared with the Friedman test and Bonferroni-corrected Wilcoxon post-hoc tests; inter-rater reliability used intraclass correlation coefficients (ICC).
    RESULTS: Inter-rater reliability was good to excellent (average-measure ICC 0.93-0.99). All outcomes except clarity differed among platforms (P < .001). Perplexity achieved the highest QAMAI total (27.5 ± 1.4; Kendall's W = 0.75), largely through retrieval-based source provision; excluding this domain, Claude, Perplexity and Gemini converged. Claude had the highest GQS (4.5 ± 0.6), Gemini the highest overall rating (4.3 ± 0.7); Grok scored lowest. Claude and Gemini were least readable (median FKGL 22.7 and 26.1; FRES -11.6 and -7.5). Length did not explain scores within platforms.
    CONCLUSIONS: Performance varied substantially across platforms and dimensions; no platform optimized all outcomes. These findings describe informational quality, not clinical safety or decision-making, and support specialist verification before clinical use.
    Keywords:  Arthroplasty; Artificial intelligence; Clinical decision-making; Generative artificial intelligence; Large language models; Temporomandibular joint
    DOI:  https://doi.org/10.1016/j.jormas.2026.102972
  7. Rev Assoc Med Bras (1992). 2026 ;pii: S0104-42302026000802212. [Epub ahead of print]72(8): e20260541
       OBJECTIVE: The aim of this study was to evaluate the responses generated by ChatGPT-5, Gemini, and Grok to the "six most frequently asked patient questions" on cryptorchidism published by the European Association of Urology, assessing them in terms of quality, understandability, actionability, and readability.
    METHODS: ChatGPT-5, Gemini, and Grok were asked these six frequently asked questions listed on the European Association of Urology patient information page on cryptorchidism. The quality of the responses was evaluated using the Quality Assessment Tool for Patient Health Information instrument, understandability and actionability were assessed using Patient Education Materials Assessment Tool scores, and readability was measured with the Coleman-Liau Index. All evaluations were performed by four urologists.
    RESULTS: Among the three artificial intelligence-based large language models, Gemini achieved the highest mean Quality Assessment Tool for Patient Health Information and Patient Education Materials Assessment Tool for Printable Materials scores. Kruskal-Wallis analysis demonstrated a statistically significant difference in Quality Assessment Tool for Patient Health Information scores among the groups (p=0.005); pairwise comparisons revealed that Gemini scored significantly higher than ChatGPT-5 (p=0.001). No significant differences were observed among the models for Patient Education Materials Assessment Tool-Understandability Section or Patient Education Materials Assessment Tool-Actionability Section scores (p>0.05). In the readability analysis, Grok had the highest Coleman-Liau Index value (Coleman-Liau Index=14.68), and all models produced texts requiring a university-level reading ability. Although the Gemini model achieved higher overall quality scores, all artificial intelligence-based large language models provided good-quality but difficult-to-read information.
    CONCLUSION: The responses generated by all three models demonstrated high levels of understandability and strong actionability. We anticipate that future, more advanced versions of artificial intelligence-based large language models will further improve these outcomes and contribute positively to the existing literature.
    DOI:  https://doi.org/10.1590/1806-9282.20260541
  8. Childs Nerv Syst. 2026 Sep 04. pii: 353. [Epub ahead of print]42(1):
       PURPOSE: In recent years, artificial intelligence-based language models have emerged as a means of rapid access to health-related information. This study aimed to evaluate the quality, reliability, understandability, and readability of ChatGPT's responses to frequently asked questions by families of children with cerebral palsy (CP).
    METHODS: Responses generated by the free version of ChatGPT to the ten most frequently asked questions posed by families of children with CP were obtained. These responses were evaluated across four dimensions: quality, reliability, understandability, and readability. Content quality was evaluated using the DISCERN instrument. Reliability was assessed by 20 physiotherapists holding MSc or PhD degrees using a 5-point Likert scale. Understandability and actionability were assessed using the Patient Education Materials Assessment Tool for Printable Materials (PEMAT-P), while readability was analyzed using the Flesch-Kincaid Grade Level (FKGL).
    RESULTS: The median DISCERN score was 45, reflecting average content quality. The median Likert scale score across all questions was 4 out of 5, indicating reliable responses. No significant differences were observed between PhD and MSc expert raters in their Likert scale ratings. The median PEMAT-P was 69.23; however, only 40% of answers exceeded the threshold for understandability, and none met actionability threshold. All responses were written at a level exceeding high school, indicating limited readability for the general public.
    CONCLUSION: ChatGPT has the potential to provide accurate information to families of children with CP; however, improvements in understandability, actionability, and readability are needed to better support families. Further development is required in order to support decision-making and care processes.
    Keywords:  Artificial intelligent; Cerebral palsy; ChatGPT; Large language models; Reliability
    DOI:  https://doi.org/10.1007/s00381-026-07444-0
  9. J Dent Sci. 2026 ;21(3): 1942-1947
       Background: /purpose: The use of artificial intelligence (AI) powered chatbots in dental education is becoming increasingly widespread. Evaluating their performance and the reliability of their sources is essential to understand their educational value. The aim of this study was to evaluate the performance of AI-powered chatbots in addressing orthodontic questions from the Dental Specialty Exam (DUS) and to assess the accuracy and reliability of the information sources on which they rely.
    Materials and methods: A total of 129 orthodontic questions from the exam administered between 2012 and 2021 were categorized according to Bloom's taxonomy. Each question was individually entered into ChatGPT-5, Claude 3.7, and Copilot, and their performances were comparatively evaluated. The sources referenced by the chatbots while generating their answers were also assessed. The data were analyzed using Pearson's chi-squared test.
    Results: ChatGPT-5, Claude 3.7, and Copilot achieved accuracy rates of 82.2 %, 83.7 %, and 85.3 %, respectively. Copilot performed best on scenario-based questions (100 %) but performed worst on visual analysis questions (33.3 %). Citation analysis showed that, ChatGPT-5.0 used reliable academic sources, whereas Claude cited few and less credible references, and Copilot relied mainly on moderately reliable materials.
    Conclusion: Chatbots exhibited strong text-based reasoning abilities but limited visual interpretation skills. While ChatGPT-5.0 provided more reliable and well-referenced responses, other models showed weaker citation practices. These underscored both the potential and the current limitations of AI-based systems in orthodontic education and clinical practice.
    Keywords:  Artificial intelligence; Bloom’s taxonomy; Chatbots; Dental specialty exam; Orthodontics; Source evaluation
    DOI:  https://doi.org/10.1016/j.jds.2025.07.3563
  10. Front Public Health. 2026 ;14 1880639
       Background: Patients increasingly rely on large language models (LLMs) for health information, yet their suitability for decision-critical conditions such as acute pancreatitis remains unclear. Given that acute pancreatitis requires timely symptom recognition, severity assessment, treatment decision-making, recurrence prevention, and follow-up management, LLM-generated information should demonstrate reliability, transparency, and readability.
    Objectives: To evaluate the informational quality, visible transparency-related features, and readability of English-language responses generated by five publicly accessible LLMs to standardized patient-facing questions on acute pancreatitis.
    Methods: This cross-sectional benchmark study developed 24 English-language, single-intent questions on acute pancreatitis across six clinical domains using public search intents and guideline-derived decision-critical content. The analysis focused exclusively on English-language patient-facing responses. Each question was submitted once to GPT-5.4 Thinking, DeepSeek-V3.2, Gemini 3.1 Pro, Grok 4.3, and Qwen3.6-Max-Preview, generating 120 responses. Anonymized responses were independently evaluated by two blinded gastroenterologists using DISCERN, Ensuring Quality Information for Patients (EQIP), Global Quality Scale (GQS), and Journal of the American Medical Association (JAMA) benchmark criteria. Readability was assessed using six established formulas. A structured response-level safety analysis was added to evaluate factual inaccuracies, clinical hallucinations, and clinical safety-risk severity. Between-model differences were analyzed with Friedman tests followed by post hoc paired Wilcoxon signed-rank tests with Holm correction.
    Results: Significant between-model differences were observed across all quality, transparency-related, and readability outcomes. Grok 4.3 achieved the highest mean DISCERN, GQS, EQIP, and JAMA scores, reflecting the strongest informational quality profile according to the predefined quality instruments, although visible transparency cues remained limited across all models. In the added safety analysis, factual inaccuracies and clinical hallucinations were each identified in 16 of 120 responses (13.3%), whereas clinical safety-risk signals were identified in 7 responses (5.8%), all of which were adjudicated as score 1 (low risk) under the predefined clinician-rated rubric; no moderate- or high-risk clinical safety event was adjudicated. DeepSeek-V3.2 demonstrated the most favorable readability profile, with the highest Flesch Reading Ease score (44.42 ± 12.27), which nevertheless remained substantially below the recommended threshold of ≥ 80. None of the 120 responses satisfied all six predefined readability targets. All model-level readability distributions differed significantly from recommended thresholds in the direction of poorer readability.
    Conclusion: Publicly accessible LLMs generated English-language responses on acute pancreatitis with variable informational quality, limited visible transparency cues, and consistently inadequate readability under standardized default public-interface conditions. Because the primary analysis used a single-generation cross-sectional design, model rankings should be interpreted as performance snapshots rather than definitive or temporally stable hierarchies. Higher quality scores did not necessarily establish factual accuracy, clinical safety, patient-education-level readability, or stronger response-level transparency. Current public-interface LLM outputs may serve as clinician-reviewed drafts for patient education but should not function as standalone patient resources. Future AI-based health information systems should strengthen clinical completeness, plain-language communication, visible evidence support, actionability, and clinician-supervised safeguards.
    Keywords:  acute pancreatitis; artificial intelligence; health information quality; large language models; patient education; readability; transparency
    DOI:  https://doi.org/10.3389/fpubh.2026.1880639
  11. Allergy Asthma Proc. 2026 Sep 01. 47(5): 343-347
      Background: Aspirin-exacerbated respiratory disease (AERD), also known as nonsteroidal anti-inflammatory drug-exacerbated respiratory disease, is a chronic condition that is both clinically complex and often difficult for patients to understand. As patients increasingly turn to online tools for medical information, it is important to evaluate the quality of the responses they may receive. Objective: This study assessed the medical accuracy and readability of responses generated by ChatGPT 5.1, Gemini 2.5 Flash, and Claude Sonnet 4.5 to 12 common questions that patients often ask their providers about AERD. Methods: The 12 questions were developed by clinicians and patients; the resulting chatbot responses were de-identified and reviewed by an expert panel of 11 AERD specialists who rated each response for medical accuracy on a 10-point scale. Readability was measured by using readability and grade-level analysis programs. Results: Gemini's responses had the highest overall ratings for medical accuracy, followed by Claude and ChatGPT. The models generated responses written at a postgraduate reading level, which is substantially higher than recommended targets for patient-education materials. Questions with more complex responses showed greater variability in expert ratings, potentially reflecting differences in clinical interpretation. Conclusion: Analysis of these findings suggests that, whereas chatbots may improve access to medical information, their current outputs require clinician review and substantial simplification before they can be reliably used for patient education.
    DOI:  https://doi.org/10.2500/aap.2026.47.260052
  12. J Neurol Sci. 2026 Sep 01. pii: S0022-510X(26)00438-7. [Epub ahead of print]490 126156
       BACKGROUND: Public understanding about brain death/death by neurologic criteria (BD/DNC) is generally poor. With the rising popularity of utilization of large language model (LLM) chatbots to answer medical questions, we sought to determine the quality of information about BD/DNC provided by ChatGPT 4o-mini and Gemini 3 Flash (Fast).
    METHODS: With the assistance of a family advocate, we developed 45 open-ended questions about BD/DNC and submitted them to ChatGPT 4o-mini and Gemini 3 Flash (Fast) in January 2026. We recorded response word count, Flesch-Kincaid Readability Score and source reputability (non-reputable sources were defined as nonmedical, nongovernmental, nonlegal and not affiliated with an organ donation organization). Two authors of the 2023 BD/DNC guidelines independently assessed response accuracy relative to accepted medical standards and a third adjudicated discrepancies.
    RESULTS: Most responses were ≥ 10th grade level [ChatGPT 4o-mini: 44/45 (98%), Gemini 3 Flash (Fast): 39/45 (86%)]. After adjudication, 22/45 (49%) responses from Gemini 3 Flash (Fast) and 20/45 (44%) from ChatGPT 4o-mini were considered completely correct (p = 0.673). There was no relationship between accuracy and: word count; readability; or citation of at least one non-reputable source.
    CONCLUSION: ChatGPT 4o-mini and Gemini 3 Flash (Fast) responses to questions about BD/DNC may include inaccuracies. This could promote confusion and distrust. There is remarkable potential for integration of artificial intelligence in public education about healthcare, but there is a need for improvement to ensure responses are accurate and readable. The healthcare team should be prepared to address misconceptions about BD/DNC based on use of LLMs.
    Keywords:  Artificial intelligence; Brain death; Chatbot; Communication; Education; Ethics
    DOI:  https://doi.org/10.1016/j.jns.2026.126156
  13. Front Public Health. 2026 ;14 1900281
       Objective: To systematically evaluate the quality and readability of health information generated by four large language models (LLMs) in response to inquiries regarding type 2 diabetes mellitus (T2DM), using an authoritative Chinese clinical guideline as the reference standard.
    Methods: A total of 124 standardized questions were extracted from the Chinese Type 2 Diabetes Popular Science Guidelines. Six endocrinologists and diabetes specialists conducted independent, blind evaluations using the CLEAR tool (Completeness, Lack of false Information, Evidence, Appropriateness, Relevance) and PEMAT-P (Patient Education Materials Assessment Tool for Printable materials). Response characteristics were also recorded. Between-model differences were tested using the Kruskal-Wallis H test with Bonferroni pairwise comparisons.
    Results: All four models achieved total CLEAR scores within the "very good" range (19-25), with no significant differences seen between models (χ2 = 1.985, p = 0.576). No significant differences were observed in the dimensions of Lack of false information (χ2 = 7.644, p = 0.054), Evidence (χ2 = 2.309, p = 0.511), and Relevance (χ2 = 7.516, p = 0.057). However, significant differences emerged in Completeness (χ2 = 47.661, p < 0.001) and Appropriateness (χ2 = 88.360, p < 0.001). Claude-4.0 received the lowest score in Completeness (median 4.00, IQR 3.00-5.00) but achieved the highest ranking in Appropriateness (median 4.00, IQR 4.00-5.00). On the PEMAT-P, understandability differed significantly across models (χ2 = 159.120, p < 0.001), yet all models surpassed the 70% threshold, with ChatGPT-4.1 highest (median 91.91%, IQR 91.91-100.00%). However, despite significant differences among the various models (χ2 = 354.023, p < 0.001), only ERNIE Bot 4.5 Turbo (median 75.00%, IQR75.00-75.00%) surpassed the 70% threshold, with no single model demonstrating consistent superiority across all dimensions.
    Conclusion: Although the four LLMs generally provide accurate and pertinent information regarding type 2 diabetes, enduring limits in actionability and inconsistencies among models in content completeness and understandability restrict their effective use in diabetic patient education. Future development should prioritize stronger step-by-step behavioral guidance and differentiated, scenario-specific model deployment to enhance their value in patient-facing diabetes self-management support.
    Keywords:  health information quality; large language models; patient education; readability; type 2 diabetes mellitus
    DOI:  https://doi.org/10.3389/fpubh.2026.1900281
  14. Rev Assoc Med Bras (1992). 2026 ;pii: S0104-42302026000802206. [Epub ahead of print]72(8): e20260388
       OBJECTIVE: Older adults increasingly use artificial intelligence-based tools to obtain health information. Although artificial intelligence chatbots such as ChatGPT may enhance access, the quality, readability, and patient safety of fall-prevention information remain uncertain. This study aimed to evaluate the quality, readability, and patient safety implications of ChatGPT-generated responses to common questions about fall risk and home safety in older adults.
    METHODS: Ten frequently asked fall-related questions were submitted to ChatGPT (version 5.2). Responses were independently assessed by a multidisciplinary panel including physiotherapists, a geriatrician, a physical medicine and rehabilitation physician, an occupational therapist, and an orthopedic specialist. Quality was evaluated using the Mika classification. Readability was measured with the Flesch-Kincaid Grade Level. Interrater reliability was analyzed using a two-way random-effects intraclass correlation coefficient model with absolute agreement (intraclass correlation coefficient [2,k]).
    RESULTS: Three responses were rated as "excellent," while seven responses were rated as "satisfactory requiring minimal clarification." No response received a rating corresponding to "moderately satisfactory" or "unsatisfactory." The mean Flesch-Kincaid Grade Level was 8.4 (range 4.3-11.9). Five responses exceeded the readability levels commonly recommended for patient education materials. Interrater reliability demonstrated fair agreement (intraclass correlation coefficient [2,k]=0.72; 95%CI 0.64-0.80).
    CONCLUSION: While ChatGPT provided generally acceptable clinical information, variability in readability and expert ratings raises patient safety concerns. AI-generated health content should be reviewed and tailored to older adults' health literacy needs before clinical use.
    DOI:  https://doi.org/10.1590/1806-9282.20260388
  15. Npj Viruses. 2026 Aug 31. pii: 41. [Epub ahead of print]4(1):
      Bacteriophage therapy is re-emerging as a potential strategy to address antimicrobial resistance, but standardized patient education materials are limited. Large language models (LLMs) are increasingly used for patient-facing medical information. The quality of LLM-generated responses to 20 patient-relevant questions was evaluated by 12 clinicians and research experts in bacteriophage therapy independently rated each response for accuracy, completeness, clarity, and tone/empathy using 5-point Likert scales. Expert suggestions for improvement were recorded. A total of 960 ratings were analyzed. Adjusted mean scores ranged from 3.36 to 3.96 across domains, indicating generally favorable evaluations for all models. Significant differences among LLMs were observed for completeness and tone/empathy (Holm-adjusted p = 0.042 for both), but not for accuracy or clarity. Differences were small in magnitude (Cohen's d = 0.12-0.29). Claude scored significantly lower than the other models for completeness and tone/empathy, while Perplexity achieved the highest completeness scores. Experts recommended improvements for 34-40% of responses; wrong information was given in 20%. The best responses were revised into an expert-informed patient guide provided as Supplementary Material, presenting a hybrid model in which LLMs generate draft patient information that is subsequently refined by clinical experts, particularly in rapidly evolving therapeutic domains lacking standardized educational resources.
    DOI:  https://doi.org/10.1038/s44298-026-00224-2
  16. JDR Clin Trans Res. 2026 Sep 04. 23800844261478633
       INTRODUCTION: As generative artificial intelligence (GenAI) chatbots increasingly mediate online health information seeking, this study aimed to assess the readability of chatbot-generated responses and characterize negatively framed content by estimating its prevalence and identifying its thematic patterns across different communicative scenarios and prompt formats.
    METHODS: This cross-sectional infodemiological study evaluated fluoride-related responses generated by 5 freely accessible GenAI chatbots. To support this assessment, prompts were developed to simulate 3 communicative scenarios through which users may request or evaluate fluoride-related information: fact-checking, opinion, and explanation. Each prompt was formulated in both organic and structured versions, representing simple user-like questions and formulations expanded using prompt-engineering principles. Responses were segmented into sentences, which were independently classified by 2 trained and calibrated evaluators as positively/neutrally or negatively framed with respect to fluoride use for dental caries prevention. Readability was assessed using distinct indices, and topic modeling was performed to characterize thematic patterns.
    RESULTS: A total of 1,326 sentences were analyzed, of which 1,040 were classified as positively/neutrally framed and 286 as negatively framed. The prevalence of negatively framed sentences did not differ significantly across chatbots or prompt formats. However, it varied according to the communicative scenario, with fact-checking prompts generating a higher prevalence of negatively framed content than opinion and explanation prompts. Topic modeling showed that negatively framed sentences were primarily organized around risk-oriented themes, including dental fluorosis, excessive fluoride exposure in children, toxicity, potential neurodevelopmental concerns, and skeletal fluorosis. Although structured prompts improved readability across most grade-level indices, chatbot-generated responses generally remained within difficult readability levels.
    CONCLUSION: GenAI chatbots have the potential to support communication about fluoride use; however, the frequent emphasis on fluoride-related risks in fact-checking responses and the generally difficult readability of the generated texts may hinder lay audiences' interpretation of evidence-based recommendations.Knowledge Transfer Statement:Although generative artificial intelligence (GenAI) chatbots can provide adequate fluoride-related information, the presence of unfavorable content and high readability demands limit their practical usefulness. Because patients use these tools for health queries, clinicians must proactively address AI-driven misconceptions during consultations. Additionally, public health policymakers should advocate for digital oversight to ensure AI platforms deliver accurate, unbiased, and broadly accessible information regarding dental caries prevention.
    Keywords:  artificial intelligence; dental caries; digital health; fluorides; infodemiology; internet
    DOI:  https://doi.org/10.1177/23800844261478633
  17. Turk J Med Sci. 2026 ;56(4): 973-982
       Background/aim: Neuropathic pain management requires high patient adherence and complex dose-titration regimens, making the readability of patient information leaflets (PILs) a critical patient safety factor. This study aimed to compare the clinical content quality and linguistic readability of PILs for neuropathic pain medications between Türkiye (TİTCK) and the United States (DailyMed).
    Materials and methods: A total of 11 pairs of innovator medications were analyzed. Clinical content was assessed using the 15-item Keystone Clinical Content Checklist (KCCC). Readability was assessed using the Ateşman index for Turkish texts and the Flesch Reading Ease Score (FRES) for English texts. Information quality and reliability were evaluated using DISCERN, the Global Quality Score (GQS), and JAMA benchmarks.
    Results: TİTCK leaflets exhibited significantly higher word counts (2652.27 ± 591.6) and sentence counts (180.55 ± 30.7) than the US group (564.09 ± 75.71 and 28.18 ± 4.48, respectively; p < 0.001). While the Turkish group demonstrated superior clinical exhaustiveness (KCCC: 14.82 ± 0.40 versus 11.91 ± 1.22; p < 0.001), both groups remained within the "Difficult/Academic" readability category, with no significant difference between Ateşman and FRES scores (p = 0.4056). Notably, the US group demonstrated significantly higher educational utility (GQS: 4.27 ± 0.47 versus 3.27 ± 0.47; p = 0.0049), despite its lower content volume.
    Conclusion: The comparison of regulatory frameworks revealed a structural trade-off: while the TİTCK model offered superior clinical exhaustiveness, the FDA model provided significantly higher educational utility. Despite these differing priorities, neither model provided a readability advantage, as both remained within the "Difficult" category for the target patient population. These findings suggest that the high information density of the Turkish model acts as a significant barrier to effective health communication.
    Keywords:  Health literacy; neuropathic pain; patient information leaflets; readability
    DOI:  https://doi.org/10.55730/1300-0144.6341
  18. Nat Sci Sleep. 2026 ;18 599776
       Purpose: This study aimed to evaluate the quality and reliability of social media videos addressing weight loss for OSA across three major platforms (YouTube, TikTok, Bilibili) and to identify key determinants of trustworthy content.
    Methods: We conducted a cross-platform content analysis of publicly available videos on weight loss for OSA from YouTube, TikTok, and Bilibili, representing major video-sharing platforms with different regional audiences and content formats. The top 100 default search results from each platform were screened using predefined criteria, yielding 235 eligible videos. Video and uploader characteristics, engagement metrics, and quality scores were extracted. Videos were assessed using PEMAT, VIQI, GQS, mDISCERN, and JAMA benchmarks. The primary outcome was overall video quality assessed by the Global Quality Score (GQS), while other assessment tools were considered secondary outcomes. For TikTok, we used English-language search terms ("obstructive sleep apnea and weight loss" and "OSA and weight loss") to capture content comparable with YouTube. The Chinese-language TikTok counterpart (Douyin) was not included because it operates under a distinct platform ecosystem with different content algorithms, regional audience, and moderation policies.
    Results: YouTube videos scored significantly higher in understandability, actionability, and overall quality compared with TikTok and Bilibili. Content from verified institutions and healthcare professionals consistently received higher ratings across all tools. Engagement metrics (likes, shares, comments) showed weak correlations with information quality. In regression models, production quality (VIQI) and source transparency (JAMA benchmarks) emerged as the strongest predictors of content reliability, while video duration had minimal effect.
    Conclusion: Within digital health environments, popular engagement does not equate to informational quality for OSA-related weight-loss content. Platform- and uploader-level characteristics-particularly verification status and production standards-are key markers of reliability. These findings underscore the need for platform-level governance that prioritizes verified health sources, algorithmic transparency, and integrated digital health literacy efforts to safeguard patients navigating lifestyle advice for OSA online.
    Keywords:  Bilibili; TikTok; Youtube; obstructive sleep apnea; public health; social media
    DOI:  https://doi.org/10.2147/NSS.S599776
  19. Rev Assoc Med Bras (1992). 2026 ;pii: S0104-42302026000801001. [Epub ahead of print]72(8): e20260493
      
    DOI:  https://doi.org/10.1590/1806-9282.20260493
  20. Clin Appl Thromb Hemost. 2026 Jan-Dec;32:32 10760296261486123
      BackgroundShort-video platforms have become important sources of health information, but the quality and reliability of venous thromboembolism (VTE)-related short videos remain unclear. This study evaluated VTE-related short videos on TikTok and Bilibili and conducted an exploratory examination of factors associated with video likes.MethodsOn February 27, 2025, the top 150 videos from each platform were retrieved according to the platforms' default ranking algorithms. Video quality and reliability were assessed using the Global Quality Score (GQS), modified DISCERN (mDISCERN), and Medical Quality Video Evaluation Tool (MQ-VET). In an exploratory analysis, an eXtreme Gradient Boosting (XGBoost) model was developed to examine factors associated with video likes.ResultsA total of 184 videos were included, comprising 81 from Bilibili and 103 from TikTok. Bilibili videos were longer, covered more VTE-related topics, and had higher median numbers of saves and shares, as well as higher GQS, mDISCERN, and MQ-VET scores than TikTok videos. TikTok videos received more likes and comments, while videos from professional uploaders on TikTok demonstrated higher quality and reliability. In the exploratory XGBoost analysis, follower count contributed the most to predicting video likes, followed by days since upload and the use of subtitles.ConclusionThe overall quality and reliability of VTE-related short videos were suboptimal. Greater professional involvement and clearer reporting of evidence sources, references, and update information may enhance the educational value of online VTE-related short videos.
    Keywords:  bilibili; eXtreme gradient boosting; information quality; tiktok; venous thromboembolism
    DOI:  https://doi.org/10.1177/10760296261486123
  21. Health Informatics J. 2026 Jul-Sep;32(3):32(3): 14604582261485153
      BackgroundShort-video platforms increasingly influence patients' access to glioma treatment information; however, the quality and reliability of such content remain unclear.MethodsThis cross-sectional study analyzed 200 high-visibility Chinese-language videos on TikTok and Bilibili (TikTok, n = 103; Bilibili, n = 97). Videos were evaluated for engagement metrics, uploader characteristics, and content coverage. Information quality and reliability were assessed using the Global Quality Score, modified DISCERN, Journal of the American Medical Association benchmark criteria, and the Patient Education Materials Assessment Tool.ResultsTikTok videos were shorter and received greater user engagement than Bilibili videos, while also achieving higher Global Quality Score and modified DISCERN scores. Professional uploaders provided more reliable and evidence-based content, whereas non-professional videos were longer and attracted significantly more comments (p < 0.001); differences in likes, collections, and shares were not statistically significant. Only 30.0% of videos addressed grade- or molecular-specific treatment pathways, and prognostic factors were discussed in only 15.5% of videos; recurrence management and long-term care were seldom covered. Engagement metrics were weakly correlated with quality scores, indicating a mismatch between popularity and information quality.ConclusionsChinese-language short videos on glioma treatment provide accessible but uneven patient education. Verified medical authorship, guideline-based content labeling, and quality-weighted recommendation strategies may improve the visibility of trustworthy information.
    Keywords:  bilibili; digital health; glioma; patient education; tiktok
    DOI:  https://doi.org/10.1177/14604582261485153
  22. Naunyn Schmiedebergs Arch Pharmacol. 2026 Sep 04.
      Baloxavir marboxil is a novel antiviral agent; however, the accuracy of medication-related information available on social media remains uncertain. This study aimed to systematically evaluate the content coverage, quality, and reliability of short videos related to baloxavir marboxil on Douyin and Xiaohongshu. The top 100 baloxavir marboxil-related videos were collected from both Douyin and Xiaohongshu. Quality and reliability were assessed using the JAMA criteria, modified DISCERN (mDISCERN), Global Quality Score (GQS), and the medication information items (MII) checklist. Correlations between video quality and user engagement metrics were also analyzed. A total of 176 eligible videos were included. While videos on Xiaohongshu demonstrated higher quality and reliability than those on Douyin, Douyin videos showed greater engagement. Median JAMA scores were 2.00 (Xiaohongshu) and 1.00 (Douyin); mDISCERN scores were 2.00 on both; GQS scores were 3.00 on both; and MII values were 3.00 and 2.00, respectively. A substantial proportion of videos scored poorly on JAMA-2, JAMA-4, mDISCERN-2, and mDISCERN-4. Physicians and individual users were the main uploaders. Most videos employed a monologue format and focused on indications. Videos uploaded by pharmacists had higher quality and greater medication information coverage, and were the most engaging. Although video quality was not correlated with popularity, MII was a significant predictor of high-quality videos. Misinformation mainly involved exaggerated efficacy, pharmacological misunderstandings, inappropriate self-medication, incorrect indications, misleading safety claims, and commercial exaggeration. The quality and reliability of baloxavir marboxil-related videos on Douyin and Xiaohongshu remain suboptimal, with notable deficiencies in source referencing. Significant differences in video quality exist between platforms, highlighting the need for users to exercise discernment when accessing related content. This study underscores potential disparities in access to high-quality health information across user groups on short video platforms, as well as the associated health equity issues. Future efforts should focus on characterizing misinformation related to baloxavir marboxil on short video platforms. Additionally, social media platforms should establish policies to review and monitor health education content to ensure the public receives accurate and trustworthy health information.
    Keywords:  Baloxavir marboxil; Douyin; Influenza; Xiaohongshu
    DOI:  https://doi.org/10.1007/s00210-026-05865-x
  23. PLoS One. 2026 ;21(9): e0355294
      Health literacy is essential for making informed healthcare decisions, yet ethnic minority communities often face disparities due to limited access to culturally relevant information. This qualitative study explores how South Asian and Black communities in the United Kingdom (UK) assess health information credibility, examining cultural influences and expectations. Sixty-eight participants (42 South Asian, 26 Black) were recruited using purposive sampling; data collection involved semi-structured interviews (n = 30) and community observations (n = 38). Findings revealed that cultural expectations shape perceptions of healthcare professionals and health information sources. Trusted sources included personal contacts and social media. Health information lacking personal or culturally relevant content led participants to seek alternative sources. While some credibility-seeking behaviours overlap with those observed in the general population, the unique social, cultural, and historical experiences of these populations shape how and why such behaviours are enacted, resulting in distinct patterns of trust and information use. Providers must recognise these dynamics to create personalised, culturally sensitive resources that enhance trust, health literacy, and equity. Tailored approaches, including trusted intermediaries and relevant communication channels, are crucial, while digital platforms must address misinformation.
    DOI:  https://doi.org/10.1371/journal.pone.0355294