bims-librar Biomed News
on Biomedical librarianship
Issue of 2026–08–23
twenty-one papers selected by
Thomas Krichel, Open Library Society



  1. Med Ref Serv Q. 2026 Aug 17. 1-23
      This scoping review aims to explore and map the existing literature that encompasses academic health sciences libraries (AHSLs) and maker learning spaces. This research will identify and provide an in-depth analysis on how AHSLs engage with maker programming, technologies, tools, and services. The literature examined will include both scholarly and gray literature from 2005 to the present day and will seek to highlight maker initiatives implemented specifically in these settings. This aims to inform future implementation and research initiatives by identifying gaps, trends, and opportunities for maker learning in AHSLs.
    Keywords:  Academic health sciences libraries; maker movement; makerspaces; scoping review
    DOI:  https://doi.org/10.1080/02763869.2026.2718912
  2. J Integr Bioinform. 2026 Aug 17.
      Maintaining scientific databases that depend on continuous curation of research literature often requires labor-intensive, slow, and error-prone annotation processes. To address these challenges, we present a pipeline that integrates text mining with expert supervision to support database expansion. Using the BRENDA enzyme database as a case study, we compiled a relation extraction dataset by aligning document-level annotations with literature references through distant supervision. We then developed a neural model that performs entity recognition and relation classification, enabling the extraction of enzyme-strain associations from full-text articles. To close the loop between machine learning and expert curation, we designed a web-based interface that allows annotators to review and refine predicted relations. While preliminary, our initial experiments show the potential of combining weak supervision and human-in-the-loop validation to accelerate the integration of literature-derived information into knowledge bases.
    Keywords:  data annotation; database expansion; named-entity recognition; relation extraction; text mining
    DOI:  https://doi.org/10.1515/jib-2025-0058
  3. Urology. 2026 Aug 17. pii: S0090-4295(26)00517-0. [Epub ahead of print]
       OBJECTIVE: To evaluate the quality of responses from four publicly available LLMs (ChatGPT-4o, Claude 3.7, Gemini 2.5, and Copilot) to frequently asked questions (FAQs) in pediatric urology.
    METHODS: FAQs were generated using standardized prompts and submitted to each LLM using parent-centered instructions. Two board-certified pediatric urologists independently rated responses across seven domains: Accuracy, Completeness, Safety, Clarity, Actionability, Conciseness, and Global Quality Score using a 5-point Likert scale. Emotional tone was analyzed using the NRC Word-Emotion Association Lexicon, and readability was assessed using seven established metrics.
    RESULTS: All LLMs produced generally high expert ratings for accuracy (mean 4.27), safety (4.36), and clarity (4.42). Claude achieved the highest overall quality score (4.35), followed by Gemini (4.03), ChatGPT (4.00), and Copilot (3.80). Emotional tone was predominantly positive and supportive across models. Reading corresponded to an 11th-15th grade reading level, exceeding the recommended 6th-8th grade patient-education standards.
    CONCLUSION: LLMs can provide accurate, safe, and supportive information for parents, but their usefulness is limited by gaps in completeness and high readability levels. Claude achieved the highest overall quality, whereas ChatGPT demonstrated high safety with lower completeness. Future development would prioritize plain-language optimization, context-aware emotional framing, and parent co-design to improve comprehension.
    DOI:  https://doi.org/10.1016/j.urology.2026.08.006
  4. J Neurogastroenterol Motil. 2026 Aug 20.
       Background/Aims: : Gut directed hypnotherapy can assist with irritable bowel syndrome, but access remains limited. This feasibility study conducted an exploratory comparison of 4 large language models (LLMs) to evaluate the readability and information quality of model generated responses to common patient questions about pre- hypnotherapy support.
    Methods: : Four LLMs (ChatGPT-4o, Claude 3.7 Sonnet, Gemini, and DeepSeek-R1) were queried using a pre-specified set of 14 standardized patient-oriented prompts. Responses were assessed for readability using validated indices (FKGL, FRES, and CLI) and for information quality using the DISCERN instrument and 5 item Likert scale, rated by a single evaluator. Paired statistical tests were used, and results are reported with effect sizes and confidence intervals.
    Results: : In this exploratory evaluation, DeepSeek-R1 and ChatGPT-4o generated responses with relatively higher scores across readability metrices (Flesch-Kincaid Grade Level, Flesch Reading Ease Score, and Coleman-Liau Index) and information quality measures (DISCERN and Likert), which are distinct constructs assessed by separate instruments. Claude 3.7 Sonnet produced denser with lower readability and reliability scores, and Gemini showed intermediate performance. Interpretation is limited using automated readability metrics and single-rate design. As the DISCERN and Likert scores were generated by a single rater, statistical comparisons involving these measures cannot be considered definitive. The results below are presented solely to illustrate observable trends within this exploratory dataset.
    Conclusion: LLMs show potential for generating patient oriented pre-hypnotherapy information, but this feasibility study is not sufficient to determine clinical applicability. The findings primarily highlight the need for expanded evaluation with multiple expert raters, patient participants, and more robust quality assessment methods before considering any practical or clinical use.
    Keywords:  Artificial intelligence; Hypnosis; Irritable bowel syndrome; Mental health services; Natural language processing
    DOI:  https://doi.org/10.5056/jnm25145
  5. Front Public Health. 2026 ;14 1921819
       Background: Publicly accessible large language model (LLM) chatbots are increasingly used to seek nutrition advice. Although vegetarian and vegan nutrition advice is often perceived as low risk, it may involve supplementation, vulnerable life stages, chronic disease, and symptoms that require clinical assessment.
    Objective: To compare the safety, accuracy, empathy, reliability, information quality, and readability of five publicly accessible LLM chatbot products in responding to public-facing vegetarian and vegan nutrition advice questions.
    Methods: In this cross-sectional comparative study, 58 predefined public-facing vegetarian and vegan nutrition advice questions were submitted once in English to ChatGPT, Gemini, Microsoft Copilot, DeepSeek, and Doubao through their official web interfaces. The final dataset included 290 chatbot responses. Five blinded raters with clinical nutrition training evaluated safety, accuracy, and empathy using predefined reference-answer anchors, and assessed reliability and information quality using four established instruments: the DISCERN instrument, Ensuring Quality Information for Patients (EQIP), Journal of the American Medical Association (JAMA) benchmark criteria, and Global Quality Score (GQS). Readability was assessed using six formula-based indices: the Automated Readability Index (ARI), Coleman-Liau Index (CL), Flesch-Kincaid Grade Level (FKGL), Gunning Fog Index (GFI), Simple Measure of Gobbledygook (SMOG), and Flesch Reading Ease Score (FRES). Model comparisons were paired by question.
    Results: Inter-rater agreement was high for all human-rated outcomes. Potentially harmful responses occurred in all models, although adjusted pairwise safety comparisons were not statistically significant. Recurring harm patterns involved vitamin B12 source reliability, uncontrolled iodine or selenium intake, vulnerable life stages, chronic disease, symptom triage, and poorly traceable or overconfident statements. Accuracy, empathy, reliability, information quality, and readability differed significantly across models. ChatGPT, Copilot, and DeepSeek showed higher accuracy; ChatGPT showed higher empathy; Copilot performed best on DISCERN and EQIP; and Gemini produced the most readable responses.
    Conclusion: Publicly accessible LLM chatbots can provide useful general information on vegetarian and vegan nutrition, but their performance is uneven and safety limitations remain. Chatbot-generated advice should be interpreted cautiously, especially for supplementation, vulnerable groups, chronic disease, and symptoms requiring clinical assessment. Future systems should improve safety guidance, referral cues, source traceability, and plain-language communication.
    Keywords:  accuracy; chatbot health advice; empathy; large language models; readability; reliability; safety; vegetarian and vegan nutrition
    DOI:  https://doi.org/10.3389/fpubh.2026.1921819
  6. J Stomatol Oral Maxillofac Surg. 2026 Aug 20. pii: S2468-7855(26)00256-9. [Epub ahead of print] 102956
       INTRODUCTION: Large language models (LLMs) are increasingly utilized for medical and dental information retrieval, yet their ability to interpret authentic, patient-style inquiries remains insufficiently investigated. This study compared the performance of ChatGPT, Claude, and Gemini in responding to patient-oriented queries related to periodontal and peri-implant diseases.
    MATERIALS AND METHODS: Unlike traditional investigations using expert-generated questions, this study employed 40 realistic, patient-oriented queries designed to simulate the post-examination cognitive state, blending colloquial language with partially retained clinical jargon. Each query was submitted to GPT-4o, Claude Sonnet 5 and Gemini 2.5 Pro generating 120 responses. Three blinded periodontists independently evaluated scientific accuracy, completeness, clinical safety, and overall quality using a 5-point Likert scale. Automated text analysis assessed readability metrics (Flesch Reading Ease, Flesch-Kincaid Grade Level, Gunning Fog Index) and linguistic characteristics. Statistical protocols included Friedman, Bonferroni-adjusted Wilcoxon signed-rank, and Model Dominance analyses.
    RESULTS: Significant performance differences were observed among the models across all expert-rated domains (all p < 0.001). Gemini achieved the highest expert ratings for scientific accuracy (4.77 ± 0.22), clinical safety (4.94 ± 0.15), and overall quality (4.85 ± 0.20), and was identified as the most frequently top-ranked platform via dominance analysis. Conversely, Claude performed significantly better regarding response completeness (4.77 ± 0.22) and demonstrated the most favorable overall readability profile, yielding the lowest Flesch-Kincaid Grade Level (7.54 ± 1.27). GPT-4o consistently received the lowest expert ratings across all evaluated domains.
    DISCUSSION: While all evaluated LLMs generated high-quality responses to realistic periodontal queries, their functional strengths were highly multidimensional. Gemini demonstrated superior clinical precision and safety, whereas Claude provided more comprehensive and readable explanations. These findings support the integration of LLMs as pragmatic, high-ecological-validity complementary tools for patient education, while emphasizing the persistent necessity for professional clinical oversight.
    Keywords:  Artificial intelligence; Large language models; Patient education; Peri-implant diseases; Periodontology
    DOI:  https://doi.org/10.1016/j.jormas.2026.102956
  7. Musculoskelet Sci Pract. 2026 Aug 11. pii: S2468-7812(26)00152-9. [Epub ahead of print]85 103636
       BACKGROUND: Patellofemoral Pain Syndrome (PFPS) is a highly prevalent musculoskeletal condition affecting young adults and athletes. Patients increasingly turn to AI chatbots for medical information, yet the reliability, safety, and readability of these tools for PFPS remain unclear.
    OBJECTIVE: To evaluate accuracy, clarity, completeness, consistency, readability, and health advice disclaimers in responses from four AI chatbots (ChatGPT, Gemini, Claude, Perplexity).
    METHODS: On February 18, 2026, thirty common PFPS questions were submitted to four AI models. Anonymized responses were independently evaluated for: information quality (accuracy, clarity, completeness, consistency) using a 4-point Likert scale; readability via seven indices benchmarked against the sixth-grade level; and safety signaling by dichotomous coding of health advice disclaimers.
    RESULTS: All models achieved median scores of 4.00 for completeness (P = 0.296) and consistency (P = 0.1). Significant differences emerged in accuracy (P = 0.019) and clarity (P < 0.001) overall, though adjusted pairwise accuracy differences were non-significant. Perplexity (median 3.00) was significantly inferior in clarity compared to other models (median 4.00). No model met the sixth-grade readability benchmark (P < 0.001); Gemini and ChatGPT were most readable, while Claude and Perplexity produced the most complex text. Health advice disclaimers appeared in 46.7% of ChatGPT and 40.0% of Claude responses, but only 16.7% for Gemini and Perplexity.
    CONCLUSIONS: AI chatbots generate accurate, complete, and consistent PFPS information but uniformly fail readability benchmarks. Disclaimer rates remain low, particularly for Gemini and Perplexity. These findings suggest AI chatbots currently function better as supplementary educational tools, highlighting the need for linguistic simplification and improved safety signaling.
    Keywords:  AI; Large language models; Patellofemoral pain syndrome; Patient education
    DOI:  https://doi.org/10.1016/j.msksp.2026.103636
  8. World J Methodol. 2026 Sep 20. 16(3): 116022
       BACKGROUND: Large language models (LLMs) are increasingly accessed by patients for gastrointestinal health information. Despite their growing use, concerns persist regarding accuracy, empathy, actionability, and readability of responses generated by LLMs.
    AIM: To assess the responses generated by ChatGPT-5, Gemini-2.5, and Claude-4 for common patient questions on "acidity" (heartburn/dyspepsia/gastroesophageal reflux disease).
    METHODS: Thirty-nine frequently asked questions were submitted to each model. Responses were independently rated by three gastroenterologists for accuracy, comprehensiveness, empathy, and actionability; and by 20 patients for empathy, comprehensiveness, actionability, compassion, and usefulness. Readability indices were also analyzed.
    RESULTS: Significant inter-model differences were observed across multiple physician-rated domains. Gemini-2.5 and Claude-4 achieved higher mean scores for accuracy, comprehensiveness, and actionability compared with ChatGPT-5 (P < 0.05), while Claude-4 demonstrated the highest empathy scores. Patient ratings indicated uniformly high comprehensibility across all models; however, Gemini-2.5 and Claude-4 responses were perceived as more actionable than those generated by ChatGPT-5. Readability analysis showed that ChatGPT-5 produced the most accessible responses, corresponding approximately to a high-school reading level, whereas Gemini-2.5 and Claude-4 generated more linguistically complex content.
    CONCLUSION: These findings underscore the need for careful model selection and suggest that hybrid approaches integrating complementary model strengths may optimize safe and effective artificial intelligence -assisted patient education in gastroenterology.
    Keywords:  Actionability; Empathy; Gastroesophageal reflux; Large language models; Patient education
    DOI:  https://doi.org/10.5662/wjm.v16.i3.116022
  9. Int Urogynecol J. 2026 Aug 20.
       INTRODUCTION AND HYPOTHESIS: Artificial intelligence (AI) tools such as ChatGPT are increasingly used by patients to seek health information, yet concerns remain regarding the accuracy and reliability of AI-generated content. Integrating evidence-based resources into these models may improve their educational effectiveness. In a previous pilot study, a retrieval-augmented ChatGPT model was found to outperform a standard ChatGPT model under controlled conditions, highlighting the need for further testing in pragmatic clinical contexts. The objective of this second pilot was to evaluate whether a retrieval-augmented AI model improves response quality and usability for patient-representative urogynecology queries, as rated by experts using validated instruments.
    METHODS: We developed a retrieval-augmented ChatGPT model grounded in the American Urogynecologic Society's (AUGS) patient education materials. Urogynecology specialists were recruited through professional networks and submitted patient-representative questions. Responses were assessed across six domains using the Quality Analysis of Medical Artificial Intelligence (QAMAI) tool, with usability measured by the System Usability Scale (SUS). Quantitative data were analyzed descriptively, and qualitative feedback was examined thematically.
    RESULTS: Twenty-two questions were posed, with 11 out of 20 respondents providing complete responses. The median QAMAI score was 28 (IQR 24-29.8) and the mean SUS score was 81.8 ± 11.2. Respondents highlighted patient-friendly language, ease of use, and integration of AUGS links as strengths, while noting citation consistency, clinical depth, and comprehensiveness as areas for improvement.
    CONCLUSIONS: A retrieval-augmented ChatGPT model trained on AUGS materials generated high-quality responses with excellent usability. Though not a substitute for clinician counseling, such models may supplement provider-patient communication.
    Keywords:  Artificial intelligence; Patient education; Pelvic floor disorders; Retrieval-augmented; Urogynecology
    DOI:  https://doi.org/10.1007/s00192-026-06834-x
  10. Eur Spine J. 2026 Aug 17.
       PURPOSE: To evaluate whether large language models (LLMs) can provide accurate, complete, and audience-adapted answers to common spine-surgery-related questions for patients and family practitioners.
    METHODS: Ten frequently asked spine-surgery questions were collected at a level 1 trauma center and simplified linguistically. Five LLMs (ChatGPT, Claude 3.5 Sonnet, Gemini Advanced 1.5 Pro, Copilot Pro, and DeepSeek V3) were queried using zero-shot prompting with persona-specific instructions for family practitioners and middle-aged patients. Responses were assessed by spine surgeons and non-medical raters for correctness, completeness, adaptability, and empathy using five-point Likert scales. Readability was quantified using the Flesch Reading Ease Score (FRES).
    RESULTS: All LLMs generated largely correct and usable responses. ChatGPT and Claude showed the highest correctness and completeness, particularly for practitioner-directed answers. Gemini and Copilot achieved superior readability and empathy for patient-facing responses. DeepSeek demonstrated balanced performance across all domains. Readability differed substantially between practitioner- and patient-oriented outputs.
    CONCLUSION: LLMs can support communication and education following spine surgery when used with structured prompting. Clinical oversight remains essential to mitigate risks related to inaccuracies and hallucinations.
    LEVEL OF EVIDENCE: III.
    Keywords:  Artificial intelligence; ChatGPT; Large language models; Spine surgery
    DOI:  https://doi.org/10.1007/s00586-026-10280-0
  11. Front Public Health. 2026 ;14 1923536
       Objective: To evaluate the performance of large language models (LLMs) and expert clinicians in optimizing Chinese patient education materials (PEMs) for temporomandibular disorders (TMD) across readability, accuracy, actionability, and cultural adaptability, and to determine whether a human-AI collaboration model can achieve an optimal balance among these dimensions.
    Methods: Fifteen TMD education topics were selected through a Delphi consensus process. Four groups of PEMs were generated: original texts (Group A), LLM-rewritten texts (Group B, Claude Opus 4.6), clinician-rewritten texts (Group C), and human-AI collaboration texts (Group D). Blinded assessments were conducted by a professional panel (n = 3) and a health literacy-stratified patient panel (n = 6). Readability was measured by sentence length and common character proportion, accuracy by 5-point expert ratings against standardized checklists, actionability by the PEMAT instrument, and cultural adaptability by qualitative thematic analysis. Prompt sensitivity was tested across three distinct styles.
    Results: LLM-rewritten texts demonstrated significantly shorter sentences (21.1 vs. 33.9 characters, p < 0.001) and substantially higher actionability (95.6% vs. 16.7%, p < 0.001) compared with original texts. Clinician-rewritten texts achieved the highest accuracy (4.48 vs. 3.23, p < 0.001) but low actionability (46.7%). The human-AI collaboration model matched clinician accuracy (4.49, p = 1.000) while preserving LLM actionability (95.6%), with 72% less clinician time. Sensitivity analysis confirmed actionability robustness under empathetic and authoritative prompts but attenuation under a concise prompt. Qualitative analysis identified four cultural adaptability themes, with LLMs excelling in terminology localization and behavioral structuring but showing limitations in TCM conceptual depth.
    Conclusion: The human-AI collaboration model represents a promising and balanced approach for producing Chinese TMD patient education materials, combining LLM strengths in structural optimization with expert clinician oversight for accuracy and cultural appropriateness.
    Keywords:  Chinese; actionability; cultural adaptability; human-AI collaboration; large language models; patient education; temporomandibular disorders
    DOI:  https://doi.org/10.3389/fpubh.2026.1923536
  12. J Hand Surg Glob Online. 2026 Nov;8(6): 101100
       Purpose: Online health information enhances education and connection to care for patients seeking hand surgery in the United States but can be challenging to navigate for patients who speak languages other than English. This study sought to assess the availability and accessibility of online Spanish-language resources at academic hand surgery institutions.
    Methods: We performed a cross-sectional study of 95 institutions offering hand surgery fellowships listed by the American Society for Surgery of the Hand (2024-2025). Institutional websites were evaluated for (1) patient educational content in English and Spanish, (2) language toggle features, (3) availability of language services information, and (4) the proportion of hand surgeon profiles offering Spanish. Readability of carpal tunnel syndrome resources in English and Spanish was assessed using the Fry and Gilliam-Peña-Mountain readability formulas, respectively, and differences were assessed using Student t tests. Associations between Hispanic population and language accessibility features were analyzed using point-biserial correlation.
    Results: Of 93 institutional websites included in the analysis, 73.1% offered English patient resources, 18.2% provided language toggle options, 11.8% offered translated Spanish resources, 29% included a link to language assistance services, and 8.6% listed a direct phone number for interpreter access. Among 736 identified hand surgeons, 6.7% listed Spanish as an offered language. State-level Hispanic population was not significantly associated with language accessibility features. At the city level, Hispanic population was significantly associated with presence of a website language toggle (r = 0.27, P = .01) and Spanish-speaking providers (r = 0.32, P = .003). English and Spanish carpal tunnel resources (P = .45) were written above recommended reading levels, with no significant differences between languages.
    Conclusions: Academic hand surgery websites demonstrate limited availability of Spanish-language patient education, language assistance information, and Spanish-speaking providers. Development of language-concordant online resources represents a cost-effective opportunity to promote equitable access to care.
    Clinical relevance: This study promotes equitable patient education in hand surgery.
    Keywords:  Hand surgery; Health literacy; Language barriers; Patient education
    DOI:  https://doi.org/10.1016/j.jhsg.2026.101100
  13. Indian J Cancer. 2026 Aug 18.
       BACKGROUND: Cervical cancer is a leading cause of cancer-related deaths in women in low-and middle-income countries. Awareness plays a crucial role in early diagnosis and prevention. Among social media platforms, YouTube offers extensive health awareness content, including information on cervical cancer. However, the validity, reliability, and usefulness of this content are not known. This study aimed to evaluate the quality and validity of widely viewed Hindi YouTube videos on cervical cancer.
    METHODS: YouTube was searched using the term "cervical cancer-Hindi." The 50 most-viewed Hindi videos were manually coded and statistically analyzed. Reliability was assessed using the DISCERN scale, while quality, credibility and usefulness were evaluated using the Journal of the American Medical Association (JAMA) and Global Quality Score (GQS) tools. Two experienced public health researchers independently rated the videos, and their agreement correlation scores were calculated. Videos from professionals, news organizations, and individuals were compared.
    RESULTS: Among the 50 videos, 78% were uploaded by professionals ( N = 39), 14% by news organizations ( N = 7), and 8% by individuals ( N = 4). The DISCERN median 36.50 (31-42), GQS median 8.00 (6.00-10.00), and JAMA median 6.00 (4.00-8.00) were reported. Observers' agreement for JAMA, GQS, and DISCERN was statistically significant. However, no significant differences were observed between professionals, individuals, and news organizations, or between genders and content types.
    CONCLUSION: Valid, reliable, and good-quality cervical cancer videos are available on YouTube and widely viewed. However, as no creator type showed significantly better quality, more efforts by professionals and news organizations are needed to enhance video reliability and awareness outcomes.
    Keywords:  Cervical cancer; Hindi videos; YouTube; health awareness; health communication; social media
    DOI:  https://doi.org/10.4103/ijc.ijc_238_26
  14. J Craniofac Surg. 2026 Aug 19.
      YouTube is widely used as a source of medical information, but the educational value of videos on orbital fractures remains unclear. This cross-sectional content analysis evaluated the quality, reliability, and orbital fracture-specific completeness of YouTube videos. YouTube was searched on April 3, 2026, using "orbital fracture," "blowout fracture," and "orbital floor fracture." After duplicate removal and predefined exclusions, 123 English-language videos were included. Two ophthalmologists independently assessed each video using DISCERN, Journal of the American Medical Association benchmark criteria, Global Quality Score, and an orbital fracture-specific score. Inter-rater reliability ranged from moderate to excellent. Overall quality was moderate but heterogeneous, with median DISCERN, JAMA, GQS, and OFSS scores of 46.0, 1.5, 3.5, and 3.5, respectively. Medical professional/institutional videos had higher JAMA and OFSS than nonmedical videos in unadjusted analyses, although after false discovery rate adjustment, only the OFSS difference remained significant. Across content categories, differences in DISCERN and OFSS remained significant after adjustment, with surgical/academic videos showing the highest scores. In exploratory multivariable models, video duration remained independently associated with all 4 quality scores, while professional/institutional source and surgical/academic content were most consistently associated with OFSS. Most popularity metrics were not associated with quality; view count, daily viewing rate, and Video Power Index showed weak associations only with JAMA score. YouTube videos on orbital fractures provide only moderate, inconsistent information, and popularity should not be taken as a proxy for educational value.
    Keywords:  Blowout fracture; YouTube; online health information; orbital fracture; patient education
    DOI:  https://doi.org/10.1097/SCS.0000000000013302
  15. Eur J Pediatr. 2026 Aug 21. pii: 684. [Epub ahead of print]185(9):
      Cow's milk protein allergy (CMPA) is a common pediatric condition frequently discussed on social media. Because symptoms are often nonspecific and parental anxiety is high, inaccurate online information may contribute to overdiagnosis and unnecessary use of hypoallergenic formulas. This study evaluated the quality, reliability, and clinical accuracy of short-form social media videos on CMPA. In this cross-sectional content analysis, short-form videos from Instagram Reels, TikTok, and YouTube Shorts were searched in April 2026 using the hashtag #cowsmilkproteinallergy. The first 50 eligible English-language videos from each platform were included. Educational quality and reliability were assessed using the Global Quality Scale (GQS), modified DISCERN (mDISCERN) score, and a CMPA Information Accuracy Scale (CMPA IAS) developed by allergy/immunology experts based on current guidelines. Of the 150 videos, 69 were produced by medical creators and 81 by non-medical creators. Overall, 25 videos (16.6%) contained potentially harmful or misleading information. TikTok had the highest engagement metrics but also the highest proportion of non-medical creators and harmful content. Videos produced by medical creators had significantly higher GQS, mDISCERN, and CMPA IAS scores, whereas non-medical creators generated significantly greater engagement. In multivariable analysis, Instagram videos scored higher than TikTok videos for GQS and mDISCERN. Videos produced by medical creators also scored significantly higher across all three instruments. Additionally, longer video duration was independently associated with higher scores.
    CONCLUSION: Within this hashtag-based sample, short-form social media videos on CMPA frequently contained incomplete or potentially harmful information, particularly on highly engaging platforms and among non-medical creators. Our findings suggest that greater healthcare professional involvement and strategies to improve the visibility of evidence-based content may help improve the quality of online CMPA information.
    WHAT IS KNOWN: • Short-form social media platforms are increasingly used by parents to obtain health information, but the quality and reliability of this content are highly variable. • Cow's milk protein allergy (CMPA) is particularly vulnerable to misinformation because of its nonspecific symptoms and high levels of parental anxiety.
    WHAT IS NEW: • This is the first cross-platform study evaluating the clinical accuracy, quality, and reliability of short-form videos on CMPA using a guideline-based CMPA Information Accuracy Scale. • Highly engaging platforms and non-medical creators were more likely to share potentially harmful information, highlighting the need for greater healthcare professional involvement.
    Keywords:  Cow’s milk protein allergy; Health information quality; Misinformation; Pediatric allergy; Short-form video; Social media
    DOI:  https://doi.org/10.1007/s00431-026-07358-8
  16. J Am Acad Orthop Surg Glob Res Rev. 2026 Aug 01. 10(8):
       BACKGROUND: TikTok has become an increasingly popular source of health information. This study aimed to assess the quality and content of TikTok videos discussing glenohumeral dislocation. The hypothesis being that videos produced by healthcare professionals (HCPs) would be superior.
    METHODS: This cross-sectional study used the keyword "shoulder dislocation" to retrieve videos, and the first 173 videos found were reviewed. Inclusion criteria was (1) English language, and exclusion criteria were (1) video following a trend, (2) no association with shoulder dislocation, and (3) being no longer available at the time of the analysis. Videos were classified into two groups: those by general users and those by HCPs. Basic information was extracted, and the content classified by two independent raters into six types: definition, signs/symptoms, risk factors/prevention, evaluation, management, and outcome. The same raters assessed the quality of the information presented with the validated DISCERN instrument. Any disagreements were resolved by a third rater. Pearson correlation and nonparametric Wilcoxon-Mann-Whitney tests were used for statistical analysis.
    RESULTS: One hundred two of the 173 videos reviewed met the inclusion criteria; 70 from general users and 32 from HCPs. The most discussed topic was management (73), while definition (15) was the least discussed. A significant difference was observed in DISCERN score between videos by HCPs (28.78 ± 7.16) and general users (19.00 ± 5.99; P < 0.01).
    CONCLUSION: Videos by HCPs were of higher quality than those by general users, although overall content quality was poor.
    DOI:  https://doi.org/e24.00337
  17. Front Public Health. 2026 ;14 1826485
       Background: Semaglutide, a glucagon-like peptide-1 receptor agonist originally developed for glycemic control, has recently gained widespread attention for its weight-loss effects. With the rapid rise of short-video platforms, these media have become important channels for disseminating health information. However, the quality, reliability, and underlying determinants of semaglutide-related content on such platforms have not been systematically characterized.
    Methods: A cross-sectional content analysis was conducted on semaglutide-related videos from Bilibili and TikTok. Videos were collected on February 4, 2026. Two independent reviewers evaluated the quality of each video using the Global Quality Scale (GQS), modified DISCERN (mDISCERN), and Journal of the American Medical Association (JAMA) benchmark criteria. Multivariable regression analyses were performed to identify factors associated with information quality.
    Results: A total of 225 videos were included (107 from Bilibili and 118 from TikTok). Most videos focused on weight loss (82.2%), while only 2.7% primarily discussed hypoglycemic effects. The overall median scores were 4.0 (3.0-4.0) for GQS, 4.0 (3.0-5.0) for mDISCERN, and 2.0 (1.0-3.0) for JAMA. TikTok videos had significantly higher JAMA scores than Bilibili videos (median 3 vs. 1, p < 0.001), while GQS and mDISCERN scores were comparable between platforms. In multivariable analyses, videos uploaded by medical professionals were consistently associated with higher quality scores across all three instruments, whereas commercial videos were associated with lower mDISCERN and GQS scores. Higher numbers of likes were independently associated with lower mDISCERN and GQS scores overall, and a higher number of comments on TikTok was similarly associated with lower quality scores, whereas a higher number of shares on Bilibili was positively associated with mDISCERN and JAMA scores.
    Conclusion: Semaglutide-related videos on Chinese short-video platforms demonstrated relatively high subjective quality and usability (mDISCERN and GQS), but markedly lower adherence to formal transparency benchmarks (JAMA), with substantial variability driven primarily by uploader professional background. Content produced by medical professionals tends to provide more reliable information, while engagement metrics do not consistently reflect information quality.
    Keywords:  Bilibili; DISCERN; TikTok; semaglutide; short video
    DOI:  https://doi.org/10.3389/fpubh.2026.1826485
  18. Medicine (Baltimore). 2026 Aug 21. 105(34): e50408
      Cerebral hemorrhage is a critical public health issue marked by high incidence, disability, and mortality. In China, short-video platforms have become the major channel for the public to acquire relevant health knowledge. However, the online cerebral hemorrhage-related health information varies greatly in quality, with complex dissemination mechanisms and a lack of systematic academic evaluation, which greatly weakens the authenticity and practicality of popular science content for the public. This study aimed to compare the content ecosystems of cerebral hemorrhage-related short videos on TikTok and Bilibili, explore the core factors associated with video dissemination effects, and evaluate the classification performance and feature correlation of machine learning models in identifying high-quality health popularization videos. A total of 200 cerebral hemorrhage-related short videos were collected from the 2 platforms, with their metadata, interaction indicators, and content quality scores recorded. Statistical tests and regression analyses were adopted to compare platform differences and variable correlations. Nine machine learning models were built for high-quality video identification. Area under the curve, calibration curves, and ridgeline plots were used to evaluate model performance and visualize the distribution characteristics of core research indicators. The 2 platforms presented obvious differences in content ecological characteristics. Bilibili had longer-duration, institution-produced, and higher-quality videos, while TikTok featured short videos, individual creation, and stronger user interactivity. Light Gradient Boosting Machine showed the optimal analytical performance, and video duration was the core stable feature affecting video quality. Videos created by medical professionals had significant advantages in information credibility, actionability, and understandability. This study identified a notable "quality-dissemination paradox," revealing the essential conflict between platform algorithm logic and public health information quality. Nonmedical creators need to enhance professional capabilities, and platforms should optimize recommendation algorithms to boost the efficient and accurate dissemination of high-quality cerebral hemorrhage health information.
    Keywords:  cerebral hemorrhage; content quality; health communication; machine learning; short videos
    DOI:  https://doi.org/10.1097/MD.0000000000050408
  19. Health Expect. 2026 Aug;29(4): e70826
       BACKGROUND: Caregivers may encounter short videos about paediatric general anaesthesia and neurodevelopment before consultation. We assessed their quality, reliability, usability and misinformation patterns, and considered the implications for parent-centred risk communication.
    METHODS: On 15 August 2025, we used newly created accounts to retrieve the first 100 results from Douyin and Bilibili for two Chinese search terms. Two anaesthesiologists independently assessed eligible videos using GQS, MQ-VET, JAMA benchmarks, mDISCERN and PEMAT-A/V. A pre-search PPIE consultation, conducted through brief one-to-one discussions with 10 family contributors, informed study framing and parent-facing interpretation.
    RESULTS: Of 200 retrieved videos, 95 were included (Douyin, n = 74; Bilibili, n = 21). Douyin videos attracted more engagement, but quality did not differ by platform. Videos uploaded by anaesthesiologists, hospitals and news agencies tended to score higher than those uploaded by non-anaesthesiologists in several domains, although these subgroup comparisons were exploratory. Misinformation occurred in 51 videos (53.7%); false safety assurances were most common (43/95, 45.3%). After false-discovery-rate correction, stratified analyses showed no significant differences by platform or uploader. Engagement was not a reliable indicator of quality, whereas video duration was positively correlated with mDISCERN, GQS and MQ-VET.
    CONCLUSIONS: Short videos were readily accessible but uneven in reliability and transparency. Popularity did not indicate quality. Preoperative communication should address online information, correct binary claims, explain exposure-related modifiers and direct families to trustworthy clinician-produced resources. Caregiver exposure and decision outcomes were not measured.
    PATIENT OR PUBLIC CONTRIBUTION: Ten family contributors took part in brief one-to-one PPIE consultations during pre-anaesthesia assessment. Their advice supported the use of lay search language and the parent-facing interpretation of omissions and false reassurance. Contributors did not review or score videos, and their views were not treated as individual-level research data.
    Keywords:  caregiver decision‐making; health information quality; misinformation; paediatric anaesthesia; parent‐centred care; short‐video platforms
    DOI:  https://doi.org/10.1111/hex.70826
  20. Front Public Health. 2026 ;14 1898275
       Background: The intensification of global population aging has led to a growing demand for health information among older adults. Proactively seeking health information is crucial for enabling older adults to make informed health decisions. However, the factors influencing older adults' health information seeking behavior (HISB) have yet to be systematically summarized.
    Objectives: This study aims to identify the barriers and facilitators of HISB among older adults using the Capability, Opportunity, Motivation-Behavior (COM-B) model, and to provide a theoretical basis for developing intervention strategies.
    Methods: Following PRISMA guidelines, we conducted a search in seven databases, including PubMed, Embase, PsycINFO, CINAHL, Web of Science Core Collection, Scopus, and the Cochrane Library. The initial search period covered the period from the establishment of each database to 2 November 2025. The Mixed Methods Appraisal Tool was used to assess the methodological quality of the included studies. Researchers used convergent integration methods to collect, analyze, and integrate data. The COM-B model was applied to identify potential barriers and facilitators.
    Results: A total of 20 studies were included, comprising 16 quantitative descriptive studies, three qualitative studies, and one mixed-methods study. The findings highlight a complex and multifaceted range of factors influencing HISB, which were categorized into six domains: physical capability (skills to access information, age-related physiological decline and healthy lifestyle habits), psychological capability (limited ability to evaluate health information, health literacy and trust in the source of information), physical opportunity (credibility of accessible health information, access to information sources and financial difficulties), social opportunity (family support and social support), reflective motivation (self-efficacy, perceived usefulness of health information and intention to maintain health status), automatic motivation (personality trait).
    Conclusions: This review identified 15 factors influencing HISB among older adults and categorized them using the COM-B model. The analysis underscores the importance of addressing barriers and leveraging facilitators related to capability, opportunity, and motivation as a foundation for intervention design. Future research should further examine and validate these determinants and use them as a theoretical basis for developing targeted interventions to promote HISB among older adults.
    Systematic review registration: https://www.crd.york.ac.uk/PROSPERO/view/CRD42024600497, identifier: CRD42024600497.
    Keywords:  barriers; facilitators; health information seeking behavior; older adults; systematic review
    DOI:  https://doi.org/10.3389/fpubh.2026.1898275