bims-librar Biomed News
on Biomedical librarianship
Issue of 2026–08–09
twenty-six papers selected by
Thomas Krichel, Open Library Society



  1. Soc Work. 2026 Aug 06. pii: swag039. [Epub ahead of print]
      Library social work is an emerging field. In rural communities, where resources and infrastructure are often limited, public libraries serve as vital community hubs. However, little is known about rural contexts, particularly community needs and librarians' experiences. This study explored rural residents' perspectives on libraries and librarians' experiences in addressing their community's psychosocial needs. The authors used mixed-methods data from a large project aimed at developing a social service model for rural libraries. As part of this project, the authors conducted a statewide survey of rural residents (N = 1,145), followed by qualitative interviews with rural librarians (N = 11). The survey reported a wide range of psychosocial needs and showed their high expectations for local libraries to address them through various services, including resource referrals, online consultations, and in-library experts (e.g., social workers). Community members generally viewed psychosocial service provision as a fundamental role of local libraries, recognizing them as essential spaces for community well-being. This study underscored the gap between rural communities' psychosocial needs and the capacity of libraries to respond effectively. Social work can help bridge this gap through core skills such as community mobilization, resource navigation, direct service provision when needed, and training and consultation for librarians.
    Keywords:  community engagement; interprofessional collaboration; library social work; mixed-methods data; rural communities
    DOI:  https://doi.org/10.1093/sw/swag039
  2. Prev Sci. 2026 Aug 07.
      The Research Methods Resources (RMR) section of the National Institutes of Health (NIH) website provides information on design, analysis, and sample size for designs used in NIH-supported clinical trials including randomized controlled trials, individually randomized group-treatment trials, parallel group- or cluster-randomized trials, group- or cluster-stepped wedge trials, and group regression discontinuity designs. The material provided for each design includes background information, links to webinars and training opportunities, a list of frequently asked questions and answers, the relevant CONSORT or TREND statement, key references for the designs and for state of the practice reviews, methods for sample size and power, and common errors. Sample size calculators are readily available for randomized controlled trials, but not for the other four designs; as a result, the RMR website provides a sample size calculator for each, set up as a tutorial in 8-9 steps. Each step includes text and references that can be explored if the user is new to the calculator or left largely hidden if the user is familiar with the calculator. Each sample size calculator provides worked examples for variations for the design including those based on a simple difference vs a net difference, or on a cross-sectional design vs a cohort design. The worked examples are available for download as an.pdf document and present the formulae used for the calculations. The calculators can be used to estimate sample size requirements for a new trial or to confirm a calculation made by others. The sample size calculators are illustrated with three examples.
    Keywords:  Clustered design; Power analysis; Sample size; Tutorial
    DOI:  https://doi.org/10.1007/s11121-026-01955-7
  3. Sci Data. 2026 Aug 07. pii: 1150. [Epub ahead of print]13(1):
      Scientific data is a critical input into scientific research. Yet the research data landscape is constantly changing as new datasets emerge, others are retired, or some disappear altogether. Without a systematic way to track how datasets are used across a research field, researchers have no reliable method for identifying relevant data resources or locating communities that work with them. Data-usage descriptors can substantially advance research productivity by reducing the time that researchers spend finding new and relevant datasets in their research field, and the communities that use them. This paper describes how to generate data-usage descriptors by finding how datasets are used in publications and then linking the dataset information to the publication metadata. It also shows how usage descriptors can be used to find other related datasets and their usage. It concludes by arguing that the approach represents a critical piece of foundational infrastructure that could be deployed in repositories as part of a referenceable, navigable, and contextual data framework. This article contains a reproducible workflow for constructing data-usage descriptors, based on analyzing the full text of publications in the Dimensions database. The illustrative use case is research on food security. The illustrative repository is the National Data Platform.
    DOI:  https://doi.org/10.1038/s41597-026-07753-8
  4. Int J Popul Data Sci. 2026 ;11(5): 3581
      This presentation demonstrates how the ODISSEI research infrastructure enables interoperable, population-scale social science through the integration of data, compute, and services. At the core of the architecture is access to high-quality administrative microdata via Statistics Netherlands (CBS), combined with secure storage and linkage through the CBS Data Storage environment. Discovery and contextualisation of these data are facilitated through the ODISSEI Data Access Broker (DAB) and the ODISSEI Portal, which together streamline dataset findability, access requests, and metadata harmonisation across providers. The analysis leverages population-scale networks constructed within the Secure Analytics Environment (SANE), enabling computationally intensive network operations while complying with legal and ethical constraints. Interoperability across services is achieved through shared standards, coordinated identifiers, and integrated workflows across ODISSEI's service stack, including the ODISSEI Secure Supercomputer (OSSC) for scalable analysis and the ODISSEI Code Library for transparent documentation, versioning, and reuse of analytical pipelines. By linking administrative registers, derived network layers, and reproducible code within a single federated infrastructure, ODISSEI lowers barriers to complex, multi-source research while increasing robustness and replicability. The presentation focuses on the architectural design rather than the empirical results, showing how ODISSEI functions as an integrated ecosystem rather than a collection of standalone tools and illustrates how national research infrastructures can operationalize FAIR principles, support population-scale analytics, and provide a blueprint for interoperable social science infrastructures in Europe.
    DOI:  https://doi.org/10.23889/ijpds.v11i5.3581
  5. Int J Popul Data Sci. 2026 ;11(5): 3693
      While data linkage has traditionally used bespoke, customer led methods, an increased demand for linked data has led to the development of indexes representing people, businesses and locations. Data are linked to indexes through generalised methods, and an index ID is appended to each statistical entity in the dataset. Users request relevant data and join them on index IDs. This process is called indexing. This approach has clear benefits, but there is still uncertainly about the quality of indexing and the quality of datasets linked through index IDs. The Indexing First Research programme provides evidence to better understand the quality outputs of these processes and give clear precedents to direct the future application of indexing. It takes existing bespoke linkages and compares them to an indexing approach on the same project. This provides evidence on the coverage of the indexes, precision and recall of different methods and bias in each linkage method. This paper will set out the aims of the research programme and the progress to date. It will discuss the outcomes of indexing hard to link populations (ie. homeless people and prisoners), of comparing bespoke linkages and linkages via the indexes (ie. births-deaths linkages), and the accuracy of indexing data through generalised methods (ie. nursing data linked to the persons index). This research has implications for the approach taken to data linkage and gives direction for when indexing is appropriate and when bespoke linkage is required.
    DOI:  https://doi.org/10.23889/ijpds.v11i5.3693
  6. Int J Popul Data Sci. 2026 ;11(5): 3683
      The linkage evaluation toolkit project aims to develop tools and criteria that can be used to evaluate data linkage pipelines and methods. This toolkit will not only allow assessment of a single pipeline, but will also facilitate the comparison between pipelines; enabling analysts to identify their strengths and weaknesses. There are currently limited standardised evaluation metrics for linkage pipelines, except for the precision, recall, and bias of the linked datasets. Eight key evaluation areas were identified; accuracy, bias in linkage error, flexibility in input datasets, scalability, efficiency, platform suitability, ease of use, and the ease and transparency of quality assurance. Evaluation criteria for each area were identified and quantified. Tools are being developed to enable users to assess pipelines across these criteria. These include tools for stress-testing, testing the pipelines under different parameters, assessing bias, and capturing and quantifying qualitative feedback. We are liaising with quality teams for standardised tools to assess precision and recall. Once finalised, the toolkit will be used to compare linkage pipelines when linking the same administrative datasets, to identify which pipeline performs best across each criterion. Future use will include identifying suitable linkage methods to support the UK's 2031 Census. The evaluation toolkit will allow analysts to evaluate the strengths and weaknesses of linkage pipelines, in both their use and impact on linked outputs. The methods can be used to improve existing pipelines, and to support the development and evaluation of new methodologies in future.
    DOI:  https://doi.org/10.23889/ijpds.v11i5.3683
  7. PLoS One. 2026 ;21(8): e0349147
      The global move toward open education has led to greater receptivity to the concept of Open Educational Resources (OER) in terms of being able to foster inclusivity and innovation. However, it remains challenging to implement OER in the institutional context. This study examined teachers' perceptions on the adoption, challenges and future of OER at a Hong Kong higher education institution. A qualitative interpretivist approach was taken, and eight faculty members across different disciplines participated in semi-structured interviews. The major findings of this study are as follow: First, there is a distinct conceptual gap where faculty equate "free" commercial tools with "open" licensed resources; adoption is driven by "pedagogical pragmatism" rather than an ideological commitment to openness. Second, significant barriers persist, including the time burden of curation, quality concerns, and a lack of institutional incentives. Third, the study highlights the disruptive potential of Artificial Intelligence (AI) on open education. Participants view AI not merely as a tool but as a future "co-creator" that could automate content generation. These findings mean the shift for faculty is becoming a facilitator and course designer not just a content disseminator. The study also suggests that institutional policy must move beyond basic technical support to provide holistic professional development focused on copyright literacy, open licensing, and the integration of AI within open educational ecosystems.
    DOI:  https://doi.org/10.1371/journal.pone.0349147
  8. Nat Mach Intell. 2026 Jul;8(7): 1142-1156
      Compared with generic artificial intelligence agents, deep research agents perform longer-horizon reasoning and deeper literature exploration to investigate complex questions. Here we present DeepEvidence, a deep research agent for evidence exploration and synthesis across heterogeneous biomedical knowledge sources. DeepEvidence advances deep research through coordinated multi-agent collaboration combining breadth-first and depth-first research strategies to search, explore and aggregate evidence from multiple biomedical knowledge bases and literature. It also incrementally constructs an evidence graph of key entities and observations to support transparent tracking, attribution and validation of the research process. DeepEvidence substantially outperforms generic artificial intelligence agents across four open benchmarks. We further establish seven benchmark tasks spanning major stages of biomedical discovery, including drug discovery, preclinical experimentation, clinical trial development and evidence-based medicine. DeepEvidence demonstrates substantial improvements in systematic evidence exploration and synthesis. These results highlight the potential of deep research agents to accelerate biomedical discovery and translational research.
    DOI:  https://doi.org/10.1038/s42256-026-01266-0
  9. Eur J Vasc Endovasc Surg. 2026 Aug 03. pii: S1078-5884(26)00739-2. [Epub ahead of print]
       OBJECTIVE: Patients frequently seek health information and medical advice from chatbots instead of consulting their physicians or referring to credible patient education resources provided by medical societies. A study to evaluate the quality, readability, and clinical appropriateness of ChatGPT generated answers to common vascular surgery questions was designed.
    METHODS: Sixteen questions, written in a layperson style, were developed covering four major vascular conditions: abdominal aortic aneurysm; carotid stenosis; peripheral arterial occlusive disease; and varicose veins. The questions were categorised across four domains (signs/symptoms, natural history, medical advice, and best treatment) and further grouped for analysis into symptom related and treatment related queries. They were addressed to ChatGPT (gpt-3.5-turbo-0125) in dedicated sessions. Tone, complementarity, urgency, and uncertainty were assessed using adapted QUEST, DISCERN, and urgency scales. The Flesch Reading Ease (FRE) and Flesch-Kincaid Grade Level (FKGL) were used to measure readability. The accuracy, comprehensiveness, and clarity of the information was rated on a Likert scale by three board certified vascular surgeons.
    RESULTS: No significant differences were found in tone or complementarity across disease categories or question types. Urgency did not differ significantly between symptom related and treatment related queries overall; however, urgency varied significantly across subtypes of symptom related questions (p = .03), consistent with inconsistent escalation recommendations. The mean FRE was 32.3 ± 12.1 and FKGL 13.5 ± 2, corresponding to a university level reading requirement. Symptom related responses were more readable than treatment related ones (FRE 38.1 ± 10.3 vs. 26.4 ± 11.4; p = .025). Among the 16 outputs, only seven (44%) were judged clinically appropriate by all reviewers; clarity was rated adequate in 81% of responses, whereas only 50% reached acceptable accuracy and 69% acceptable comprehensiveness. Treatment related answers were particularly weak, with only 25% deemed appropriate. In symptom related questions, misalignment of urgency recommendations emerged as a potential patient safety concern.
    CONCLUSION: In this specialist evaluation, ChatGPT outputs related to vascular surgery were clear, but mostly clinically inappropriate and inaccurate, and required 13th grade reading skills. The combination of low accuracy and poor accessibility presents a serious patient safety concern for unsupervised patient education. As they stand, generalist large language models cannot provide patient facing information about vascular surgery. A rigorously validated, domain specific, and knowledge locked artificial intelligence system based on curated vascular guidelines may be more appropriate to ensure safety and comprehension.
    Keywords:  Artificial intelligence; Chatbot; Large language model; Natural language processing; Vascular surgery; Virtual assistant
    DOI:  https://doi.org/10.1016/j.ejvs.2026.07.060
  10. Indian J Orthop. 2026 Aug;60(8): 1949-1956
       Objective: This study aimed to compare the reliability of references generated by three different AI-based conversational agents (ChatGPT, Gemini, Perplexity) within the rotator cuff literature.
    Methods: Designed as a cross-sectional study, 30 subtopics were posed to each bot in two formats (letter to the editor and original article), resulting in a total of 3150 references. Reference reliability was assessed using the Reference Hallucination Score (RHS), which included the following criteria: existence/verifiability, bibliographic accuracy, PMID validity, and topical relevance.
    Results: When all formats were analyzed together, ChatGPT had the lowest mean RHS (1.81 ± 3.40) and thus emerged as the most reliable model. Gemini scored 4.01 ± 4.89, while Perplexity had the highest score at 6.51 ± 4.89 (p < 0.001). In the letter format, ChatGPT, Gemini, and Perplexity scored 1.81 ± 3.20, 3.81 ± 4.83, and 6.43 ± 4.78, respectively. In the article format, ChatGPT scored 4.02 ± 1.70, Gemini 4.13 ± 1.57, and Perplexity 6.31 ± 1.68. ChatGPT was more reliable than Gemini and significantly superior to Perplexity.
    Conclusion: AI-based conversational agents may play a supportive role in academic writing; however, they exhibit critical limitations in reference accuracy. While ChatGPT was relatively the most reliable model, Perplexity had the highest hallucination rate. Researchers must refrain from using references generated by these tools without verification, as doing so poses significant concerns for scientific integrity and ethics.
    Supplementary Information: The online version contains supplementary material available at 10.1007/s43465-026-01807-0.
    Keywords:  Artificial intelligence; Reference hallucination; Rotator cuff
    DOI:  https://doi.org/10.1007/s43465-026-01807-0
  11. Front Endocrinol (Lausanne). 2026 ;17 1895366
       Background: Diabetes-related foot disease requires timely recognition of neuropathic risk, ulceration, infection, ischemia, offloading needs, and recurrence risk. Publicly accessible large language models (LLMs) may provide patient-facing information, but reproducible prompt construction for benchmarking such outputs remains insufficiently characterized.
    Objective: This study aimed to develop and apply a domain- and source-balanced prompt framework for benchmarking patient-facing diabetic foot information generated by publicly accessible LLMs under default single-turn public-interface conditions.
    Methods: A 24-item benchmark prompt set was generated using a domain- and source-balanced framework incorporating public-query sources and guideline-derived decision-critical content. Six clinical domains were crossed with four source categories: Google Trends, Baidu Zhidao, a PubMed-indexed Chinese diabetic foot guideline, and PubMed-indexed international diabetic foot guidelines. Each prompt was submitted once to GPT-5.5 Thinking, DeepSeek-V4, Gemini 3.1 Pro, Grok 4.3, and Qwen3.6-Max-Preview, yielding 120 responses. Response quality was assessed using DISCERN, EQIP, and GQS; visible transparency-related features were evaluated using JAMA benchmark criteria; readability was assessed using six formulas; and an exploratory potential clinical-risk flag (PCF) screened for overt short-term harm signals. Formal claim-level factual-accuracy review, guideline-concordance adjudication, and hallucination-frequency analysis were not performed.
    Results: Significant metric-specific differences were observed across models. Grok 4.3 recorded the highest observed mean DISCERN, EQIP, GQS, and JAMA-based visible transparency-related scores, whereas DeepSeek-V4 showed the lowest observed mean scores for several readability-grade metrics and the highest mean FRES. Visible transparency-related scores remained low across models. No response was rated as PCF 1 or PCF 2. No response met all predefined readability targets.
    Conclusions: The proposed prompt framework provides a structured basis for public-interface LLM benchmarking in diabetic foot education. Default responses showed metric-specific variation, limited visible transparency, and inadequate readability, and should not be relied upon independently for high-risk diabetic foot decision-making without clinician oversight.
    Keywords:  benchmarking; diabetic foot; domain- and source-balanced prompt framework; large language models; patient-facing information
    DOI:  https://doi.org/10.3389/fendo.2026.1895366
  12. Urol Ann. 2026 Jul-Sep;18(3):18(3): 215-224
      Shockwave lithotripsy (SWL) is a noninvasive method for removing stones, primarily used for small, uncomplicated urinary stones, as the number of kidney stone cases continues to rise. With rising patient numbers, online information must be accurate and easy to understand. We assessed the quality and readability of SWL information available online and compared it with leaflets created by generative artificial intelligence (AI). The top 20 search results for four SWL-related terms on Google, Yahoo, and Bing were collected and duplicates removed. Generative Artificial Intelligence (GAI) platforms such as ChatGPT, Gemini, and DeepSeek generated a patient leaflet. Two authors evaluated quality using DISCERN and JAMA scores and readability using Flesch Reading Ease Score (FRES) and Flesch-Kincaid Grade Level (FKGL). Later, ChatGPT reviewed the AI-generated information with DISCERN. A total of 68 websites were included. The quality assessments showed average websites with AI-leaflet DISCERN scores of 52.4 (±8.91) and 58.7 (±5.03), JAMA Benchmark scores of 2.22 (±1.10) and 0, FRES of 52.1 (±13.8) and 51.5 (±8.05), and FKGL scores of 9.04 (±2.36) and 8.93 (±1.10). No significant differences were seen in DISCERN (P = 0.224), FRES (P = 0.652), and FKGL (P = 0.866), but differences in JAMA scores were significant (P = 0.0010). A difference was also noted in the author-rated and the AI-rated DISCERN scores. There is a deficiency in high-quality SWL information that meets global readability standards. As more individuals use GAI-resources trained on online data, it is crucial to improve their quality for patients and carers. Doing so enables them to make informed decisions and promotes their well-being during treatment, resulting in improved comfort and outcomes.
    Keywords:  Generative artificial intelligence; online information; patient information; quality; readability
    DOI:  https://doi.org/10.4103/ua.ua_42_26
  13. Orthop Traumatol Surg Res. 2026 Aug 04. pii: S1877-0568(26)00235-5. [Epub ahead of print] 104814
       BACKGROUND: Patients increasingly turn to AI chatbots for medical information, including before complex orthopaedic procedures such as progressive collapsing foot deformity (PCFD) surgery. Whether these tools deliver content of sufficient quality and accessibility for preoperative patient education remains unclear, particularly across competing platforms. This study addressed three questions: (1) Do ChatGPT, Perplexity AI and Google Gemini differ in the accuracy, comprehensiveness and clarity of their responses to PCFD-related patient questions? (2) Do these platforms produce content meeting recommended readability thresholds for patient education? (3) Does the level of agreement among blinded foot and ankle surgeons rating the quality of chatbot responses vary depending on the platform used?
    HYPOTHESIS: The three AI chatbot platforms produce responses of comparable accuracy but differ significantly in readability, with none reaching the recommended readability thresholds for patient education materials.
    PATIENTS AND METHODS: Cross-sectional comparative study. Twenty frequently asked questions regarding PCFD, covering disease understanding, conservative management, surgical planning and postoperative recovery, were submitted verbatim to ChatGPT (GPT-4o mini), Perplexity AI and Google Gemini (free versions, March 25, 2026). The 60 resulting responses were rated by three blinded foot and ankle surgeons on three 5-point Likert scales (accuracy, comprehensiveness, clarity). Readability was assessed using the Flesch Reading Ease (FRE) and Flesch-Kincaid Grade Level (FKGL). Inter-rater agreement used Kendall's W; differences between platforms were analysed using Kruskal-Wallis tests with Bonferroni-corrected pairwise comparisons.
    RESULTS: All platforms produced responses rated accurate to very accurate. Perplexity achieved significantly higher accuracy than ChatGPT (p = 0.0003) and higher accuracy and clarity than Gemini (p = 0.0065 and p = 0.0075). No platform reached the recommended FRE ≥ 60 or FKGL ≤ 6 thresholds: median FKGL ranged from 13.1 (ChatGPT) to 21.4 (Perplexity), with ChatGPT producing the most readable and Perplexity the least readable content (p < 0.001). Inter-rater agreement was fair to substantial across platforms, lowest for Perplexity.
    DISCUSSION: AI chatbots produce generally accurate baseline information on PCFD surgery, with Perplexity showing significantly higher expert-rated accuracy and clarity than the other platforms-contrary to our hypothesis of comparable accuracy-while readability remains uniformly inadequate for all platforms, as hypothesized. These tools may serve as a supplementary source of information, but their inadequate readability suggests they are not yet suited to replace tailored, surgeon-led patient education.
    LEVEL OF EVIDENCE: III; cross-sectional comparative study.
    Keywords:  Artificial intelligence; Chatbot; Flatfoot; Patient education; Progressive collapsing foot deformity; Readability.
    DOI:  https://doi.org/10.1016/j.otsr.2026.104814
  14. Can J Ophthalmol. 2026 Aug 03. pii: S0008-4182(26)00246-2. [Epub ahead of print]
       OBJECTIVE: Ophthalmological patient education materials (PEMs) are written at a reading level of grades 10.4-12.6, higher than the recommended 8th-grade level. This study quantitatively assesses the readability of large language model (LLM)-generated chatbot responses to patient queries, de novo PEM creation, and simplification of existing PEMs.
    METHODS: This systematic review and meta-analysis article was registered in PROSPERO (CRD420251123401). We searched Ovid MEDLINE, Embase, CINAHL, Web of Science, and Scopus (August 11, 2025). The primary outcome was readability, measured by the Flesch-Kincaid Grade Level (FKGL). Random-effects meta-analysis compared ChatGPT-generated PEM readability with existing control PEMs (i.e., clinician- or organization-authored websites, brochures, and clinical guidelines). A binomial test evaluated the proportion of studies achieving an 8th-grade reading level. Joanna Briggs Institute and Grading of Recommendations Assessment, Development, and Evaluation tools assessed risk of bias and quality.
    RESULTS: Of 1660 studies screened, 57 were included, with 11 contributing to the meta-analysis. Pooled FKGL scores were 9.82 (95% CI: 7.93-11.70) for ChatGPT PEMs and 9.54 (95% CI: 8.68-10.40) for control PEMs. There was no significant difference in FKGL readability between ChatGPT and control PEMs (mean difference = 0.21; 95% CI: -1.52 to 1.94; p = 0.81; I2 = 97.08%). Seventeen studies (29.82%) included LLM(s) at or below 8th-grade reading level. Grade-level prompting significantly improved LLM-generated PEM readability. Risk of bias was low to moderate. Reasons for downgrading were inconsistency and imprecision.
    CONCLUSIONS: In the context of high heterogeneity and low certainty, it appears that LLMs do not produce ophthalmology PEMs that are more readable than existing materials.
    DOI:  https://doi.org/10.1016/j.jcjo.2026.07.006
  15. Otolaryngol Head Neck Surg. 2026 Aug 04.
       OBJECTIVE: To evaluate the accuracy, completeness, clarity, source transparency, and readability of leading AI chatbot responses to patient questions about tracheostomy and to determine whether AI tools can reliably support patient education where high-quality guidance is critical for safety.
    STUDY DESIGN: Cross-sectional content analysis.
    SETTING: Virtual study environment using publicly accessible AI platforms, with expert evaluation conducted via Qualtrics-based distribution.
    METHODS: Twelve frequently asked questions about tracheostomy care were identified using search-listening tools and clinician input, then submitted to 5 AI chatbots - ChatGPT4, Google Gemini 2.0, Microsoft Copilot, DeepSeek V3, and Grok 3 - and to a senior laryngologist. Three blinded laryngologists independently evaluated each response using the Quality Analysis of Medical Artificial Intelligence instrument. Readability was assessed using nine metrics.
    RESULTS: Gemini 2.0 achieved significantly higher completeness scores than physician responses (P < .001), with DeepSeek and Grok 3 (P < .05) also outperforming (P < .05). Accuracy did not differ significantly between AI- and expert-generated responses. On average, the AI models outperformed physician in clarity, completeness, and usefulness based on QAMAI scoring (P < .05). All AI and expert responses exceeded the NIH-recommended 6th-grade reading level, ranging from 10th-13th grade (P < .001). Inter-rater reliability was 78%.
    CONCLUSION: AI chatbots can generate accurate and comprehensive responses to common tracheostomy care questions, demonstrating potential to support patient education. However, they continue to lack guaranteed, verifiable sourcing, and this study did not assess actual patient comprehension of the AI-generated responses. Future efforts should focus on adapting AI-generated education materials to meet health literacy standards and evaluating their direct impact on patient understanding and outcomes.
    Keywords:  AI; AI chatbots; large language models; patient education; tracheostomy
    DOI:  https://doi.org/10.1002/ohn.70359
  16. Int J Clin Pharm. 2026 Aug 06.
       INTRODUCTION: Oral anticoagulants are frequently associated with preventable hospital admissions, and patient education is essential to their safe and effective use. Patients increasingly use large language models (LLMs) for medication information, yet few studies have compared AI-generated outputs with regulatory materials or examined oral anticoagulants, a class in which misunderstanding can cause bleeding or thrombosis.
    AIM: To assess the informational quality, usability, readability, and output reproducibility of AI-generated patient information leaflets (PILs) for oral anticoagulants compared with FDA-referenced patient materials.
    METHOD: PILs for five oral anticoagulants (warfarin, apixaban, dabigatran, rivaroxaban, and edoxaban) were generated by ChatGPT, Gemini, and DeepSeek using a standardised zero-shot prompt based on FDA labelling templates; FDA-approved leaflets served as the comparator. Materials were anonymised and brand-blinded. Three clinical pharmacists independently evaluated each PIL using the Patient Education Materials Assessment Tool for Print Materials (PEMAT-P) and a modified DISCERN (mDISCERN), which assess informational quality, understandability, and actionability, but not pharmacological accuracy or clinical safety. Readability was assessed using seven validated indices. Output reproducibility was examined within the same day and at Day 1, 14, and 28.
    RESULTS: ChatGPT produced PILs of informational quality similar to FDA-referenced materials, both achieving a median mDISCERN score of (41/65); DeepSeek (34/65) and Gemini (33/65) scored significantly lower than ChatGPT; Gemini also scored significantly lower than the FDA-referenced materials. Understandability was acceptable for all sources, whereas actionability was limited, including in FDA materials, with no leaflet providing a patient summary or decision-support tool. All materials exceeded the recommended sixth- to eighth-grade range on most indices, with FDA leaflets the most complex. Output was stable across generations, with no significant within-model differences and the largest mean difference of 0.23 points.
    CONCLUSION: Among the evaluated models, ChatGPT most closely matched FDA-referenced materials in the assessed informational domains and demonstrated stable output over 28 days. Shared deficiencies across AI-generated and FDA leaflets suggest broader limitations in written health information. This comparability was limited to the assessed informational domains and does not establish equivalence in pharmacological accuracy, clinical correctness, or patient safety. AI-generated PILs may serve as clinician-reviewed supplementary materials and should not be used as standalone tools without independent professional review and verification of clinical content.
    Keywords:  Anticoagulants; DISCERN; Large language models; PEMAT; Patient information leaflets; Reproducibility of results
    DOI:  https://doi.org/10.1007/s11096-026-02203-2
  17. Clin Pediatr (Phila). 2026 Aug 06. 99228261475551
      Bright Futures Parent and Patient Handouts from the American Academy of Pediatrics are widely used after pediatric well-child visits, yet their content has not been systematically evaluated. We assessed all fourth-edition handouts using standard readability formulas for readability and the Patient Education Materials Assessment Tool for Printable Materials (PEMAT-P) for understandability and actionability. Patient Handouts averaged a 4th-grade level (ages 7-8), 5th-grade level (ages 9-10), and 6th- to 7th-grade level (age 11+). Parent Handouts ranged from about 5th to nearly 7th grade. The PEMAT-P scores across all 23 handouts showed high understandability (92.3%) but lower actionability (60%). Overall, Bright Futures Handouts appear readable and understandable, supporting their use in pediatric practice, although opportunities remain to improve actionability. When clinicians create or evaluate patient education materials, attention to plain language, straightforward delivery, visual aids, and checklists helps improve readability, understandability, and actionability. These strategies may support more effective patient education and promote pediatric health outcomes.
    Keywords:  Bright Futures; PEMAT-P; generative AI; health literacy; patient education; patient education material; readability
    DOI:  https://doi.org/10.1177/00099228261475551
  18. Jt Dis Relat Surg. 2026 May 21. pii: jdrs.2026.2628. [Epub ahead of print]37(3): 723-730
       OBJECTIVES: This study aims to evaluate the content quality and reliability of YouTube videos on limb lengthening surgery.
    MATERIALS AND METHODS: On July 5th, 2025, a YouTube search was performed using the keywords "limb lengthening surgery" and "leg lengthening surgery." The first 100 videos were reviewed; duplicates and those without English audio were excluded, resulting in 53 videos for analysis. Basic characteristics including views, upload time, duration, comments were recorded. Videos were categorized by source as "physician," "speaker," or "patient," and by theme as "general information" or "patient testimony." Video content quality, including accuracy, reliability, and comprehensibility of information, was assessed using the DISCERN, Journal of the American Medical Association (JAMA) Benchmark Score, Global Quality Score (GQS), and a researcher-developed Limb Lengthening Scoring System (LLSS). Two orthopedic surgeons independently evaluated all videos.
    RESULTS: The median view count of the 53 videos was 48,036 (range, 1,229 to 3,734,177), and the mean number of days since upload (as of July 5th, 2025) was 1,331 ± 735 days. The mean DISCERN was 24.4 ± 6.9, median JAMA 2 (range, 1 to 2), median GQS 1.5 (range, 1 to 3.5), and mean LLSS 1.1 ± 0.9. Physician-generated videos achieved significantly higher DISCERN and GQS scores (p = 0.031 and p = 0.040, respectively) than patient-generated videos. Speaker-generated videos had higher LLSS scores than patient-generated videos (p = 0.013). General information videos scored higher than patient testimonies for DISCERN (p = 0.035) and GQS (p = 0.025). Video duration and comment count positively correlated with LLSS (p < 0.001 and p = 0.047, respectively), whereas view counts and ratios showed no significant association with quality scores.
    CONCLUSION: YouTube videos on limb lengthening surgery are usually of low quality, with limited scientific accuracy and educational value. Physician-produced videos receive higher scores; however, their overall quality still remains limited, while patient-generated content shows the lowest reliability. These findings highlight that inaccurate or incomplete online information may influence patient expectations prior to consultation and complicate shared decision-making. Greater involvement of orthopedic surgeons and academic institutions is needed to provide clear, evidence-based, and reliable online educational content to improve patient understanding and clinical communication.
    DOI:  https://doi.org/10.52312/jdrs.2026.2628
  19. Medicine (Baltimore). 2026 Aug 07. 105(32): e50044
      Short videos on health information related to Attention-Deficit/Hyperactivity Disorder (ADHD) on mainstream platforms in China are increasing; however, the quality of relevant content has not been systematically evaluated. A search was conducted on Douyin, Bilibili, and Xiaohongshu from August 18 to 22, 2025, using the keywords 'Attention Deficit Hyperactivity Disorder ', ' Hyperactivity Disorder ', and 'ADHD'. The top 100 videos under each keyword were collected. Two researchers independently evaluated them using GQS, mDISCERN, PEMAT-A, PEMAT-U and JAMA benchmark criteria. A total of 228 eligible videos were analyzed. Most of the videos focused on disease-related knowledge (n = 124, 54.4%) and were published by medical professionals (n = 89, 39.0%). Senior professionals had the highest content quality score (GQS median: 3). The median of mDISCERN: 2)The scores of Bilibili videos on GQS, mDISCERN and PEMAT-U were significantly higher than those of Douyin and Xiaohongshu (all P < .05). However, the overall information reliability (mDISCERN median: 1) and operability (PEMAT-A median: 25%) were low. It is worth noting that there was no significant positive correlation between user engagement indicators (likes, comments, sharing) and any quality score (all | r |<0.2), indicating weak associations with limited practical importance, while video duration was significantly but weakly positively correlated with PEMAT-A score (R = 0.187, P < .001, suggesting limited practical value). In short, although ADHD-related videos are widely prevalent on mainstream short video platforms in China, their popularity cannot fully reflect scientific validity. Overall video quality and reliability remain to be enhanced. In the future, relevant parties may consider optimizing platform algorithms, standardizing content publishers" obligations and enhancing public media literacy, to facilitate sound development of ADHD-related public health communication.
    Keywords:  attention-deficit/hyperactivity disorder (ADHD); health information; quality assessment; short videos; social media; user engagement
    DOI:  https://doi.org/10.1097/MD.0000000000050044
  20. Am J Clin Pathol. 2026 Aug 04. pii: aqag065. [Epub ahead of print]166(2):
       OBJECTIVES: The use of #pathology and #pathologist is increasing on TikTok, a short-form video platform with over 1 billion users. As of July 2024, #pathology had over 21 000 posts and 468 million views, while #pathologist had over 3000 posts. It could be a useful tool for education and recruitment if the content is accurate and engaging.
    METHODS: To assess the accuracy, engagement, and educational value of this growing content, we conducted a cross-sectional study analyzing 105 English-language TikTok videos identified using these keywords over a 72-hour period. Videos were evaluated using the Patient Education Assessment Tool for audiovisual material (PEMAT-AV) for audiovisual quality, the Global Quality Scale (GQS) for overall content quality, and a modified JAMA benchmark score for information accountability. Additionally, a harm-benefit score categorized educational impact.
    RESULTS: Statistical analysis revealed that educational content demonstrated significantly higher quality scores on GQS (P = .0001) and PEMAT-AV (P = .0007) than other content types. In contrast, medical profile type was associated with higher PEMAT-AV (P = .0147), JAMA (P = .0332), and harm-benefit (P = .0004) scores.
    CONCLUSIONS: While the volume of pathology-specific content is currently limited, the high average engagement metrics-298 100 followers, 1.5 million views, and 67 884 likes per video-indicate substantial user interest. The higher-quality scores for educational and medical content suggest that pathology content creators are generally knowledgeable and accurately represent the field on TikTok, highlighting TikTok's potential as a valuable platform for disseminating reliable pathology-related information.
    Keywords:  #pathologist; #pathology; TikTok; medical education; pathology; social media
    DOI:  https://doi.org/10.1093/ajcp/aqag065
  21. Otolaryngol Head Neck Surg. 2026 Aug 04.
       OBJECTIVE: To evaluate the accuracy, quality, and potential harm or benefit of TikTok videos related to ankyloglossia and to assess how uploader type influences informational reliability and educational value.
    METHODS: A cross-sectional analysis of TikTok videos was conducted using 4 search terms related to ankyloglossia. Videos posted between January 2021 and August 2024 were screened, yielding 66 eligible videos. Videos were categorized by uploader type (medical provider, influencer, layperson) and content type. Three independent reviewers assessed factual accuracy based on AAO-HNSF Clinical Consensus Statements, classified videos using a harm-benefit scale, and evaluated quality using the PEMAT-AV and Global Quality Scale (GQS).
    RESULTS: Of 66 videos, 78.8% were educational, but only 63.5% were factual. Medical providers produced significantly more factual content than influencers and laypersons (85% vs 50%, P = .045). Medical provider videos demonstrated the highest harm-benefit scores (P = .025) and the highest PEMAT-AV and GQS averages, though quality score differences were not statistically significant. Influencer-generated content achieved higher view counts despite lower accuracy. Views and shares did not correlate with quality or accuracy metrics. Lifestyle and medical advice videos were significantly associated with more negative harm-benefit scores.
    DISCUSSION: Although TikTok videos on ankyloglossia reach large audiences, misinformation remains prevalent, particularly among non-medical uploaders. Popularity metrics were poor indicators of educational quality, while videos aligned with the clinical consensus statement demonstrated higher clarity and benefit.
    IMPLICATIONS FOR PRACTICE: Medical professionals should play a proactive role in creating accurate, accessible social media content to counter misinformation and guide patient understanding of ankyloglossia. However, given that engagement metrics do not favor accuracy, broader platform-level strategies and collaborations may be necessary to meaningfully counter health misinformation on social media.
    Keywords:  Ankyloglossia; Health Misinformation; TikTok; patient safety; quality improvement; social media
    DOI:  https://doi.org/10.1002/ohn.70371
  22. Health Commun. 2026 Aug 05. 1-13
      With the growing use of artificial intelligence (AI) chatbots and online medical consultation, individuals seeking health information increasingly face not only the question of whether to seek information, but also which source to consult. Integrating the risk information seeking and processing (RISP) framework with the concept of source credibility, this study examines how risk-related motivations and credibility-related evaluations shape preferences for AI chatbots versus online human doctors. In an online experimental survey (N = 248), participants were randomly assigned to a higher-sensitivity topic (sexually transmitted diseases) or a lower-sensitivity topic (seasonal allergies). Regression analyses revealed an overall preference for online human doctors. Informational subjective norms emerged as the strongest predictor of source choice, increasing the likelihood of preferring online human doctors, while positive channel beliefs in AI significantly predicted preference for AI agents across conditions. Importantly, higher perceived information-gathering capacity was associated with greater preference for AI in low sensitive context. Topic sensitivity influenced source preference indirectly through subjective norms, whereas affective responses and knowledge-related factors showed limited effects. These findings suggest that RISP-related motivations explain when health information seeking becomes salient, whereas source credibility considerations help explain how individuals translate those motivations into source preferences. The study contributes to research on digital health communication by clarifying how human and AI sources are evaluated in contemporary health information environments.
    DOI:  https://doi.org/10.1080/10410236.2026.2711078
  23. Aust N Z J Obstet Gynaecol. 2026 Aug;66(4): e70167
       BACKGROUND: Online information in Australia about abortion has not been assessed, despite its capacity to impact the trajectory of care without face-to-face medical expertise.
    AIM: To evaluate the online information that women in Australia can access for an unintended pregnancy.
    STUDY DESIGN: Mixed methods were employed in four phases. A search identified abortion portals (defined as a webpage under a single domain, providing access to other abortion websites). DISCERN scores evaluated portals quality standard. A Student's t-test was employed to assess the difference in quality by type of organisation (portals ending in .gov.au vs. .org.au). A one-way ANOVA compared portals according to the population size they targeted: ≤ 3 million, > 3 million to 8.2 million and > 8.2 million. Summative content analysis explored the framework of written words within portals associated with decision-making.
    RESULTS: Two hundred and thirty-seven results were identified, with 24 selected. Portals were generally high quality (61.25 out of a possible 80). Quality standard did not significantly differ by type of organisation (.gov.au: 64.9, .org.au: 59.1; p = 0.28), nor by population size targeted (≤ 3 million: 64.2, > 3 million to 8.2 million: 58.3, > 8.2 million: 61.25; p = 0.59). Summative content analysis identified five categories associated with decision-making: I am pregnant now what?, emotional support, facts that impact choices, practicalities and procedural knowledge.
    CONCLUSION: The rigorous methods employed to review the online abortion sites provided women with the information required to better prepare for and navigate abortion services and make decisions expeditiously. We now need to assess if women are using these resources to make decisions about their care.
    DOI:  https://doi.org/10.1111/ajo.70167