bims-librar Biomed News
on Biomedical librarianship
Issue of 2026–08–16
33 papers selected by
Thomas Krichel, Open Library Society



  1. Med Ref Serv Q. 2026 Aug 13. 1-10
      The Rural Health Information Hub (RHIhub) is a database for trusted information related to rural health. RHIhub offers a wide range of resources for researchers, individuals, community organizations, and other stakeholders committed to improving health outcomes in rural communities. This column highlights how readers can effectively search and use RHIhub's resources and tools to address the unique healthcare needs of rural populations.
    Keywords:  Community health; funding; online database; review; underserved populations
    DOI:  https://doi.org/10.1080/02763869.2026.2715986
  2. J Evid Based Soc Work (2019). 2026 Aug 13. 1-16
       PURPOSE: Locating relevant studies is the first step of evidence-based practice, yet most searching relies on keyword matching. Artificial intelligence (AI) tools called embedding models search by meaning, but the best-known options are paid commercial services. The study asked which free embedding models best search the social work literature, whether they match the commercial standard, and whether rerankers are needed.
    MATERIALS AND METHODS: Twelve free embedding models and two commercial OpenAI models were tested on 64,956 social work records (1989-2025) using 150 curated queries. Two frontier AI judges, from families unrelated to every tool evaluated, made 50,328 blind head-to-head comparisons (nDCG@10). A judge-free known-item test (496 queries) and a blind 120-pair expert human instrument provided validation.
    RESULTS: Free tools matched or beat the commercial standard. Two free models outperformed the paid flagship; a free 300-million-parameter model beat the paid default, essentially tied for first at finding specific papers (91.9%), and with a reranker was the best configuration overall (.846). Keyword search trailed every embedding model (.604 vs. .680-.842). Rerankers rescued weak models but added nothing to the strongest. Judges agreed on 85.5% of comparisons, and committee-to-rater agreement (69-76%) matched or exceeded rater-to-rater agreement (69-71%).
    DISCUSSION: Score differences among leading models are too small to change what a searcher sees; tool choice should rest on size, speed, cost, and privacy.
    CONCLUSION: High-quality, meaning-based search of the social work literature is achievable with free tools on an ordinary computer: no subscription, no queries sent to an outside company.
    Keywords:  Literature search; evidence-based practice; information retrieval; open-source artificial intelligence; research infrastructure; semantic search
    DOI:  https://doi.org/10.1080/26408066.2026.2718944
  3. J Hist Dent. 2026 Summer/Fall;74(2):74(2): 87-95
      Like many major organizations, the British Dental Association has had a number of bases over the years, from 1880 in private homes, next in shared and rented accommodations and then, for many years, leasing a whole building in London's medical district. The latter included a library and museum, both of major interest to dental historians. Sadly, that phase came to an end in 2025. This paper explores why that happened and its worrying implications for dental historians.
    Keywords:  British Dental Association; Dental Museum Early dental professionalization; Dental library
    DOI:  https://doi.org/10.58929/jhd.2026.074.02.87
  4. J Biomed Semantics. 2026 Aug 14. pii: 16. [Epub ahead of print]17(1):
       BACKGROUND: The extensive volume of biomedical scientific literature requires efficient methods for retrieving relevant documents based on semantic technologies and biomedical concepts. While embedding-based methods have shown improvements over traditional keyword-based methods, the integration of domain-specific terminologies like Medical Subject Headings (MeSH) into these models remain underexplored.
    METHODS: This study compares three hybrid methods that integrate MeSH-based annotations with document embeddings (called "pre-annotation", "post-annotation" and"post-reduction"). We benchmark these hybrid methods against the following traditional methods: TF-IDF, standard neural embeddings (Word2Vec, fastText, Doc2Vec) and publicly available transformer-based models (BioBERT, SciBERT, SPECTER, SapBERT), using cosine similarity and Word Mover's Distance (WMD) as evaluation metrics. The benchmark experiments are based on the RELISH corpus, a manually curated dataset of PubMed articles where experts have labeled pairs of documents with regards to their relevance to each other, providing a 2-class (relevant vs. non-relevant) as well as a 3-class (relevant, partially relevant, non-relevant) judgment.
    RESULTS: Transformer-based models, particularly fine-tuned BioBERT and SciBERT align best with the expert judgements after fine-tuning. Among non-transformer methods, Doc2Vec and MeSH-based hybrid methods also perform well, demonstrating the benefits from combining structured biomedical vocabularies with embedding methods. Our experiments deliver extensive results showing that the baseline performance of 76-78% precision at position 5 can be achieved through almost all approaches, improvements of 2-4% with MeSH concepts can be achieved, but performances up to 90% is left to the fine-tuned large-scale public models.
    CONCLUSION: The performance gains from the integration of concepts may be underwhelming, however the benefits lie in the successful integration and benchmarking of structured vocabularies with embedding methods, the applicability of these techniques to aligning literature with other data sources via a controlled vocabulary and the potential for stronger performance on tasks and corpora where concept-based resources are better suited. All experiments have been conserved as a Dockerized pipeline, making the full benchmarking workflow reproducible and supporting future research in biomedical document retrieval.
    Keywords:  Biomedical document retrieval; Biomedical document similarity; Document embeddings; Recommendation systems; Semantic similarity
    DOI:  https://doi.org/10.1186/s13326-026-00365-6
  5. Pediatr Transplant. 2026 Aug;30(8): e70409
       BACKGROUND: Pediatric kidney transplant recipients and their caregivers require comprehensive education to support self-management. Increasingly, families supplement clinic teaching with online platforms such as Google and YouTube, yet the quality, readability, and usefulness of these resources are uncertain.
    METHODS: Structured searches on Google and YouTube were conducted using patient-style questions and medically oriented key terms across four prioritized topics: healthy eating, exercise, travel, and mental health. Resources were screened for relevance and assessed for accuracy, completeness (defined as including all key facts needed to support recommendations), readability, and actionability (defined as providing clear, practical steps users can follow). Two pediatric transplant healthcare providers independently reviewed each resource. Two researchers assessed understandability and actionability using the Patient Education Materials Assessment Tool (PEMAT), and readability using multiple readability tools.
    RESULTS: Searches yielded 87 Google and 41 YouTube resources; most targeted adult transplant recipients. Healthcare providers assessed 65% of Google and 68% of YouTube resources to be accurate, but only 18% and 5% were factually complete. Actionable recommendations were identified in 42% (Google) and 29% (YouTube). Readability ranged from grade 5 to 17, with only 6 of 82 Google resources meeting the recommended grade 6 or lower. Mean understandability scores were 67% for both platforms. Actionability scores averaged 39% (Google) and 77% (YouTube). Resources suitable for pediatric audiences (n = 28) had similar readability and understandability scores but slightly higher actionability (Google: 43%, YouTube: 81%). Most resources lacked pediatric focus, co-creation, and practical guidance.
    CONCLUSIONS: Online searches generated limited high-quality resources tailored to pediatric needs. Healthcare providers should guide families to vetted resources and consider co-creating materials to improve relevance and impact.
    Keywords:  Google; YouTube; caregivers; internet; patient education; pediatrics
    DOI:  https://doi.org/10.1111/petr.70409
  6. Eur Spine J. 2026 Aug 14.
       PURPOSE: Lumbar spinal stenosis (LSS) is a common degenerative spinal condition and a leading cause of pain and disability in adults. With increasing use of artificial intelligence (AI) for medical information, this study evaluated and compared the accuracy and completeness of responses generated by large language models (LLMs) and Google Search to standardized patient-oriented questions about LSS, aiming to characterize AI performance in patient education.
    METHODS: Eight frequently asked questions regarding LSS were identified using NHS hospital websites, NASS clinical guidelines, Google "People Also Ask," Reddit discussions, and common ChatGPT prompts. Each question was submitted to Google Search, ChatGPT, Gemini, and Perplexity under standardized conditions. Orthopedic spine surgeons independently rated responses for accuracy and completeness using 5-point Likert scales. Parametric (one-way ANOVA) and non-parametric (Kruskal-Wallis) tests evaluated overall differences. A minimal clinically important difference (MCID) was established as a ≥ 1.0-point delta, with post-hoc pairwise comparisons evaluated using Tukey's Honestly Significant Difference.
    RESULTS: Gemini achieved the highest completeness (4.47 ± 0.67) and accuracy (4.34 ± 0.65), followed by Perplexity (completeness 4.03 ± 0.82; accuracy 4.00 ± 0.95). Google and ChatGPT demonstrated similar completeness (3.75 ± 1.04 vs. 3.78 ± 0.71), though Google showed slightly higher accuracy than ChatGPT (3.78 ± 0.91 vs. 3.63 ± 0.75). Global performance variations were highly significant across both omnibus parametric testing (completeness: p = 0.002; accuracy: p = 0.004) and non-parametric Kruskal-Wallis testing (completeness: p = 0.003; accuracy: p = 0.004). Post-hoc pairwise testing via Tukey's HSD confirmed that this statistical significance was driven solely by Gemini, which outperformed both ChatGPT (completeness p = 0.006; accuracy p = 0.004) and Google Search (completeness p = 0.004; accuracy p = 0.036). No other isolated pairwise combinations achieved statistical significance (p > 0.05). Despite these isolated statistical boundaries, no platform pairing cleared the predefined ≥ 1.0-point threshold for clinical significance, with maximum mean score deltas reaching only 0.72 for completeness and 0.71 for accuracy.
    CONCLUSIONS: Although newer-generation and search-augmented LLMs demonstrate clear statistical superiority over traditional web results in synthesizing structured spine information, these differences do not yet translate into clinically meaningful quality shifts under concise prompt constraints. Substantial inter-rater variability and the lack of threshold-clearing clinical significance reinforce that these emerging AI platforms should serve as supplements to, rather than substitutes for, physician-guided counseling in degenerative spine care.
    Keywords:  Artificial Intelligence; ChatGPT; Gemini; Google; Lumbar spinal stenosis; Perplexity
    DOI:  https://doi.org/10.1007/s00586-026-10249-z
  7. Digit Health. 2026 Jan-Dec;12:12 20552076261478722
       Objective: Acute cholecystitis (AC) is among the most frequently encountered conditions in emergency care settings. Recently, artificial intelligence (AI)-driven language models have emerged as innovative resources for accessing and synthesising medical information; however, their reliability and readability across different languages remain insufficiently established. Therefore, this study aimed to assess the reliability and readability of AC information provided by AI-based language models.
    Methods: Seven standardised questions were formulated in Spanish and English. Two independent reviewers subsequently evaluated the reliability of the responses via a validated assessment tool. Spanish readability was measured via the Flesch-Szigrist formula. In English, the Flesch Reading Easy score and the Flesch‒Kincaid readability score were employed. The results were then compared by tool and language.
    Results: Regarding reliability in Spanish, Perplexity® generated the highest percentage of complete responses (57.14%), followed by both ChatGPT® and Gemini® (42.86%). In English, Gemini® demonstrated the strongest performance with a score of 85.71%, followed by both Perplexity® and ChatGPT® with 71.43% each. The readability analysis revealed significant differences among the AI models in both Spanish and English (p < 0.05). In Spanish, Gemini® generated the most readable responses, whereas Perplexity® produced the least readable text. In English, ChatGPT® achieved the lowest Flesch-Kincaid grade level (highest readability), while Perplexity® generated the most complex responses.
    Conclusion: Despite the absence of statistically significant differences in reliability among the AI models, Gemini® generated the highest proportion of complete responses in English, whereas Perplexity® showed the highest proportion in Spanish. Moreover, significant differences in readability were observed among the three AI language models, with Perplexity® consistently generating the least readable content across both languages. There is a clear need for improvements to optimise the accuracy and accessibility of AI-generated medical information about AC.
    Keywords:  Artificial Intelligence; acute cholecystitis; generative Artificial Intelligence; large language models; readability
    DOI:  https://doi.org/10.1177/20552076261478722
  8. Front Public Health. 2026 ;14 1853895
       Background: Online video platforms such as YouTube and Bilibili have become important sources of health information for the public. However, the quality and reliability of online health education videos vary substantially. Artificial intelligence (AI), particularly large language models (LLMs), has recently shown potential for automated evaluation of health information, yet evidence regarding the agreement between AI-based and expert evaluations remains limited.
    Methods: A cross-sectional study was conducted to analyze asthma-related health education videos retrieved from YouTube and Bilibili. Two medical experts independently evaluated video quality using the Global Quality Scale (GQS) and modified DISCERN (mDISCERN). AI-assisted evaluation was performed using the GPT-4 large language model based on video transcripts under standardized prompts. Interrater agreement between experts was assessed using intraclass correlation coefficients (ICCs) and Spearman correlations. Agreement between AI and expert ratings was evaluated using Spearman correlation and Bland-Altman analysis. Multivariable linear regression was performed to identify factors associated with differences between AI and expert scores.
    Results: A total of 200 asthma-related health education videos were included, with 100 videos from each platform. AI ratings showed moderate correlations with expert ratings for both GQS (ρ = 0.55, p < 0.001) and mDISCERN (ρ = 0.49, p < 0.001). AI scores were slightly higher than expert scores, with mean differences of 0.24 for GQS and 0.60 for mDISCERN. Multivariable regression analysis identified platform (YouTube), source type (organizational source), and PEMAT actionability as significant factors associated with differences in mDISCERN scores, whereas no significant predictors were identified for GQS score differences.
    Conclusion: AI-generated evaluations demonstrated moderate positive correlations with expert assessments in evaluating the quality of asthma-related health education videos. AI may serve as a scalable tool for preliminary screening of online health information, although systematic differences remain across platforms and content characteristics. Expert evaluation remains essential to ensure the accuracy and reliability of medical information assessment.
    Keywords:  Global Quality Scale; artificial intelligence; asthma; health education video; mDISCERN; social media
    DOI:  https://doi.org/10.3389/fpubh.2026.1853895
  9. Front Public Health. 2026 ;14 1889768
       Background: The demonstrated protective effects of leisure activities on physical and mental health underscore the need for accessible guidance. Large Language Models (LLMs) like ChatGPT-5 offer a potential solution, yet their application in non-clinical leisure health advice requires rigorous evaluation. This study aims to conduct a multidimensional assessment of ChatGPT-5's performance in this context.
    Method: We generated responses from ChatGPT-5 to 34 common leisure-and-health questions, categorized into six thematic areas (e.g., mental, physical, social health). The responses were assessed using validated evaluation instruments, including the modified DISCERN tool (mDISCERN) to determine reliability, the Global Quality Scale (GQS) to assess overall quality, and a 7-point Likert scale to evaluate perceived usefulness. Readability was assessed using the Flesch Reading Ease (FRE) formula.
    Results: ChatGPT-5 demonstrated moderate reliability, good quality, and relatively high usefulness, with mean scores of 3.58/5 for reliability (mDISCERN), 4.11/5 for quality (GQS), and 5.79/7 for usefulness. However, performance varied thematically, with the highest scores in "Leisure and Social Health" and the lowest in personalized contexts like "Age- and Stage-Appropriate Planning." A critical finding was the low average FRE score of 39, indicating a "difficult" reading level equivalent to U. S. college grades 13-16, which poses a significant accessibility barrier.
    Conclusion: While ChatGPT-5 shows promise as a complementary tool for generating leisure health advice, its utility is constrained by suboptimal readability, inconsistent source transparency, and limitations in handling nuanced, personalized scenarios. For safe and effective integration, future developments must prioritize readability optimization, enhanced source citation, and emotional intelligence, all within a framework that emphasizes human oversight.
    Keywords:  health communication; large language models; leisure activities; leisure health; quality of health information
    DOI:  https://doi.org/10.3389/fpubh.2026.1889768
  10. JMIR Form Res. 2026 Aug 14. 10 e91572
       Background: Peer-reviewed medical literature consistently violates established health literacy readability targets, creating a gap that effectively excludes patients and caregivers from accessing evidence-based information.
    Objective: This study aimed to evaluate whether a large language model (LLM) can generate plain-language summaries of pediatric strabismus literature while preserving clinical fidelity and meeting established health literacy readability targets.
    Methods: This cross-sectional study analyzed 85 open access, peer-reviewed pediatric strabismus articles published between 2022 and 2025, stratified by strabismus subtype, surgical relevance, and publication type. Full-text articles were processed using DeepSeek-V3 (DeepSeek) via a structured prompt, which instructed the model to provide a simplified summary meeting the following requirements for each article: a seventh-grade or lower reading level, a maximum length of 800 words, and strict preservation of medically significant data. Primary outcomes were readability scores measured by the Flesch-Kincaid Grade Level (FKGL) and Simple Measure of Gobbledygook (SMOG) indices. Secondary outcomes included clinical fidelity, which was independently assessed by 2 fellowship-trained pediatric strabismus specialists.
    Results: Baseline articles demonstrated a mean FKGL score of 15.79 (SD 1.53) and a mean SMOG score of 14.41 (SD 1.09). Following LLM simplification, the mean FKGL score significantly decreased from 15.79 (SD 1.53) to 7.84 (SD 1.30), representing a mean difference of 7.95 (95% CI 7.52-8.38; P<.001). Similarly, the mean SMOG score decreased from 14.41 (SD 1.09) to 7.68 (SD 0.94), representing a mean difference of 6.73 (95% CI 6.42-7.04; P<.001). Postsimplification readability did not differ significantly by strabismus subtype or surgical relevance (all adjusted P>.05). However, case reports retained slightly higher FKGL scores (mean 8.35, SD 0.89) compared to reviews (mean 7.89, SD 1.37) and original research (mean 7.48, SD 1.40) (adjusted P=.003). Out of the 85 summaries, clinical fidelity was rated good in 81 (95.29%), moderate in 4 (4.71%; these were exclusively summaries of review articles), and poor in 0 (0%).
    Conclusions: DeepSeek-V3 effectively reduced the reading level of complex pediatric strabismus literature by approximately 8 grade levels, achieving National Institutes of Health-recommended eighth-grade or lower targets without compromising clinical accuracy. When integrated with clinician oversight, LLM-generated summaries offer a scalable, equitable tool to enhance health literacy and support shared decision-making for patients and caregivers.
    Keywords:  FKGL; Flesch-Kincaid Grade Level; LLM; SMOG; Simple Measure of Gobbledygook; clinical fidelity; large language model; pediatric strabismus; readability
    DOI:  https://doi.org/10.2196/91572
  11. Ear Nose Throat J. 2026 Aug 10. 1455613261477452
      IntroductionXerostomia is a common condition that can impair quality of life. Google and ChatGPT are increasingly used for patient education, though the readability and understandability of this information remain unclear. This study evaluated commonly asked questions about xerostomia and assessed the quality of online and AI-generated patient educational material (PEM).Methods"SEO minion" was used to identify Google People Also Ask (PAA) questions from four primary search terms. Questions were categorized using Rothwell's classification and entered into ChatGPT. Online sources and ChatGPT responses were evaluated using the Flesch Kincaid Grade Level (FKGL), where higher scores indicate higher reading grade level; Flesch Reading Ease (FRE), where lower scores indicate more difficult readability (0 (complex) to 100 (easy)); JAMA benchmark criteria, where higher scores indicate greater source quality (0-4); and Patient Education Materials Assessment Tool (PEMAT), where higher scores indicate greater understandability (0-100).ResultsA total of 117 unique PAA questions were identified, most of which were value-based (55.6%) or fact-based (42.7%). Online sources had a mean JAMA score of 1.09, mean FKGL of 9.87, mean FRE of 51.37, and mean PEMAT score of 79.92. ChatGPT responses performed worse, with a mean FKGL of 15.24, mean FRE of 30.14, and mean PEMAT score of 72.33.ConclusionsBoth online PEM and ChatGPT responses on xerostomia exceeded recommended reading levels, with ChatGPT generating substantially less readable content. Although many materials met understandability thresholds, their elevated reading level may still limit accessibility. More readable, reliable, and patient-centered educational materials are needed.
    Keywords:  chatGPT; health literacy; patient education; xerostomia
    DOI:  https://doi.org/10.1177/01455613261477452
  12. Cureus. 2026 Jul;18(7): e112387
       BACKGROUND: Pediatric pelvic floor dysfunction is a common cause of lower urinary tract symptoms and bowel dysfunction in children, requiring accurate and reliable educational resources for patients and caregivers. With the increasing use of online health information, both YouTube (Google LLC, Mountain View, CA, USA) and artificial intelligence (AI) chatbots have become frequently used sources of medical guidance. This study aimed to evaluate and compare the educational quality, reliability, and safety of YouTube videos and AI-generated responses related to pediatric pelvic floor dysfunction.
    MATERIALS AND METHODS: In this cross-sectional infodemiology study, YouTube was searched using five predefined pediatric pelvic floor-related search terms, and the first 50 videos for each term were screened. After duplicate removal and eligibility assessment, 25 videos were included for analysis. Three publicly available AI chatbots (ChatGPT (OpenAI, San Francisco, CA, USA), Gemini (Google LLC, Mountain View, CA, USA), and Copilot (Microsoft Corporation, Redmond, WA, USA)) were evaluated using a standardized set of 20 pediatric pelvic floor-related questions. Educational quality was independently assessed using the DISCERN instrument, the Global Quality Scale (GQS), and the Pediatric Pelvic Floor Education and Reliability Score (PPFERS), a pediatric-specific assessment tool developed for this study.
    RESULTS: The 25 included YouTube videos accumulated 447,398 total views. Physiotherapists or pediatric physical therapists were the most common content creators (44.0%), followed by healthcare institutions (28.0%). The mean DISCERN, GQS, and PPFERS scores for YouTube videos were 66.24 ± 8.64, 4.28 ± 0.74, and 16.88 ± 1.96, respectively. Among AI chatbots, Gemini achieved the highest educational quality scores (76.00 ± 1.92, 4.95 ± 0.22, and 19.35 ± 0.59), followed by ChatGPT (74.20 ± 2.40, 4.85 ± 0.37, and 18.60 ± 0.82) and Copilot (67.85 ± 2.03, 4.35 ± 0.49, and 17.10 ± 0.72). Overall, AI-generated responses demonstrated higher mean educational quality scores than the YouTube video cohort.
    CONCLUSIONS: Both YouTube videos and AI-generated responses provided generally useful educational information regarding pediatric pelvic floor dysfunction. However, AI chatbots, particularly Gemini and ChatGPT, demonstrated greater educational consistency and closer alignment with contemporary pediatric urotherapy principles. AI-assisted educational tools may serve as valuable adjuncts to patient and caregiver education but should complement, rather than replace, professional medical evaluation and individualized clinical guidance.
    Keywords:  ai-based chatbots; artificial intelligence in medicine; dysfunctional voiding; pediatrics and child health; pelvic floor dysfunction; pelvic floor exercise; youtube videos
    DOI:  https://doi.org/10.7759/cureus.112387
  13. Patient Educ Couns. 2026 Aug 09. pii: S0738-3991(26)00347-2. [Epub ahead of print]152 109814
       OBJECTIVES: As breast cancer incidence increases globally, the demand for information related to patients' health, well-being, exercise, and nutrition post-diagnosis also increases. Online platforms are now a source of information for patients with cancer and their relatives globally. However, the quality of the information provided on these platforms and the extent of involvement among different types of information providers remain debatable.
    METHOD: This mixed-method study analysed Facebook, X, Instagram, TikTok, YouTube, Pinterest, Google, and ChatGPT for cancer nutrition content and collected the top 10 results, corresponding to 5 search terms selected by the research team. Post-diagnosis cancer nutrition claims were extracted from the collected content; claims were labelled based on the type of content providers (health professionals, personal accounts, or groups and organisations) and evaluated against the current evidence, including nutrition guidelines for patients with cancer, systematic reviews, and published human studies by a panel of experts.
    RESULTS: One hundred and fifty-one cancer nutrition claims were extracted from 387 unique pieces of content, collected across the online platforms. Overall, less than 13% of the extracted nutrition claims were supported by reliable evidence. The least evidence-supported information about cancer nutrition was on X, TikTok, and Pinterest. While the most accurate platforms were Google (33.33%) and ChatGPT (30%), the quality of information was relatively poor. "Health professionals" contributed the smallest number of nutrition claims overall; however, the accuracy of shared information was not significantly different among the content creators.
    CONCLUSIONS: The majority of online nutrition information for breast cancer patients is low quality and not supported by current guidelines. Non-evidence-based nutrition claims have been observed across all platforms and shared by all types of content providers.
    PRACTICE IMPLICATIONS: To combat misleading health information, education campaigns are needed to improve media literacy and sufficient resources allocated to disseminate evidence-based information through online platforms.
    Keywords:  Breast cancer; Cancer nutrition; Digital health; Mixed-method analysis; Nutrition information; Online health information; Patient education
    DOI:  https://doi.org/10.1016/j.pec.2026.109814
  14. J Taibah Univ Med Sci. 2026 Aug;21(4): 844-855
       Objectives: Temporomandibular disorders (TMDs) are a common cause of chronic orofacial pain, and patients are increasingly using the Internet to locate health information. However, the quality of web-based information, especially in Arabic, has not been previously evaluated.
    Methodology: In this study, we analyzed 408 Arabic language websites, including those for medical centers, dental centers, non-profit organizations, and universities. The DISCERN instrument, Journal of the American Medical Association standards, and Health On the Net (HONcode) certification were used to assess quality and reliability. Adapted Flesch Reading Ease (FRE), Flesch-Kincaid Grade Level (FKGL), and Simplified Measure of Gobbledygook (SMOG) indices were used as measures of readability.
    Results: Quality and transparency varied significantly across website affiliations (p < 0.001). Only 11% of sites had a HONcode certification, 31.4% included source attributions, and 59.6% disclosed authorship. The overall mean DISCERN quality scores ranged between 3.33 and 3.87 (scale 1-5). The quality varied significantly between affiliated websites (p < 0.001). Video content was present on 6.9% of the websites and audio content on 0.03% of websites. No validated Arabic-language instruments exist for readability assessment, so the FRE, FKGL, and SMOG indices (all developed for English) produced unrealistic values (mean FRE = 174.41), and given the scale of 0-100, it was impossible to reach any conclusions.
    Conclusion: Significant shortcomings were found in terms of transparency, evidence attribution, and balance in the presentation of TMD information online in Arabic. We recommend the creation of evidence-based, methodological, and patient-centered Arabic digital health content on TMDs.
    Keywords:  Arabic health information; Arabic websites; Patient education; Quality assessment; Readability; Temporomandibular disorder
    DOI:  https://doi.org/10.1016/j.jtumed.2026.07.007
  15. J Am Acad Orthop Surg. 2026 Aug 11.
       INTRODUCTION: Pathology of the long head of the biceps tendon is a frequent source of shoulder pain. Biceps tenodesis is a common surgical intervention, and many patients seek information online. The readability and quality of online resources regarding biceps tenodesis remain unexamined. The purpose of this study was to evaluate the source type, readability, and quality of online patient-facing resources regarding biceps tenodesis.
    METHODS: Search Engine Optimization Minion extracted "People Also Ask" (PAA) questions from four terms: "biceps tenodesis," "biceps tenodesis indications," "biceps tenodesis complications," and "biceps tenodesis recovery." Searches were done in an incognito browser with data cleared. Three reviewers categorized sources and assessed credibility using the Journal of the American Medical Association (JAMA) benchmark criteria and readability using the Flesch-Kincaid Grade Level (FKGL), Flesch-Kincaid Reading Ease, and Gunning Fog Index against national patient-education standards.
    RESULTS: Four hundred queries and associated sources were analyzed. Journals were the most common (115, 28.8%), followed by academic sources (107, 26.8%) and medical practice websites (73, 18.3%). Most sources (317, 79.3%) exceeded recommended sixth-grade to eighth-grade levels: mean FKGL 9.8 (SD = 2.5), GFI 13.7 (SD = 2.7), and FKRE corresponding to grade level 12.5 (SD = 2.9). The mean JAMA score was 2.2 (SD = 0.9). Government websites were the easiest to read, FKGL M = 4.1 (SD = 0.0); GFI M = 7.8 (SD = 0.0), and moderately credible, JAMA M = 2.0 (SD = 0.0), but represented only 13 (3.3%). Journal sources had the highest credibility, JAMA M = 3.2 (SD = 0.5), yet the lowest readability, FKGL M = 11.5 (SD = 2.1); GFI M = 15.3 (SD = 2.0).
    CONCLUSION: The prevalence of difficult-to-read and low-quality information on biceps tenodesis highlights a gap between patient interest and the quality of accessible resources. Future efforts should explore simplifying medical content and improving the readability of online materials related to biceps tenodesis.
    DOI:  https://doi.org/10.5435/JAAOS-D-26-00217
  16. Digit J Ophthalmol. 2026 ;32(2): 32-39
       Purpose: To analyze accuracy and readability of answers to cataract surgery queries produced by artificial intelligence (AI) models and determine whether AI models can significantly improve readability.
    Methods: Google Gemini Advanced, ChatGPT 4.0, and Microsoft Copilot Pro were prompted to answer 25 questions about cataract surgery, followed by a request to re-answer questions at a 6th-grade level. Objective readability of answers were measured with five validated reading formulas and word count. Accuracy and readability of each answer were further graded by three ophthalmologists. Comparisons were performed between original and 6th-grade versions and among the three AI models.
    Results: After being prompted to answer at a 6th-grade reading level, Google Gemini Advanced and Microsoft Copilot Pro had lower average reading level than ChatGPT 4.0 (8.04 vs 8.19 vs 9.43 [P < 0.001]). Microsoft Copilot answers had higher Flesch reading ease score (75.40 vs 71.24 vs 69.46 [P < 0.007]) and lower word count (130.28 vs 180.24 vs 166.08 [P < 0.001]) among AI models. Microsoft Copilot Pro and ChatGPT 4.0 answers had greater change in reading level (-6.13 vs -5.75 vs -3.31 [P < 0.001]) and Flesch reading ease score (39.67 vs 35.98 vs 23.67 [P < 0.001]) compared with Google Gemini Advanced. Graders determined that there were no changes in accuracy before and after being prompted to answer at a 6th-grade reading level.
    Conclusions: AI models can simplify reading level of responses to common cataract surgery queries while maintaining accuracy.
    DOI:  https://doi.org/10.5693/djo.01.2026.03.001
  17. J Pediatr Soc North Am. 2026 Nov;17 100420
       Background: Individuals frequently consult the internet when seeking information about their health. However, patient education materials (PEMs) published online are often written above the recommended sixth-to eighth-grade reading level. The readability and quality of online resources for inherited conditions are poorly studied. Additionally, little is known about the reading level of responses from artificial intelligence (AI) models when prompted with healthcare inquiries. We aimed to assess the readability and quality of online health information encountered when searching for inherited pediatric orthopaedic conditions, including outputs from popular large language model (LLM) chatbots.
    Methods: Twenty-three inherited pediatric orthopaedic conditions were queried using Google in three common search formats. First-page search results, including Google AI Overviews, were analyzed. Searches were replicated in three chatbots: Microsoft Copilot, ChatGPT, and Google Gemini. Readability was assessed using Flesch Reading Ease (FRE) and Flesch-Kincaid Grade Level (FKGL).
    Results: Of 706 identified webpages, 355 met inclusion criteria. Most webpages were hosted by government agencies, academic hospitals, or nonprofit organizations. Mean readability levels were high (FRE 38.2; FKGL 12.1), substantially exceeding the recommended levels for patient-facing health information. Only 18% of webpages met eighth-grade readability standards. AI-generated outputs demonstrated similar readability (FRE 36.1; FKGL 12.3). Only 2% of AI responses met eighth-grade readability recommendations. However, when the input prompt included a request to improve readability, AI models lowered average output readability by nearly six grade levels (FRE 69.0; FKGL 6.8).
    Conclusions: Online information for inherited pediatric orthopaedic conditions is frequently written above recommended reading levels, and AI-generated content does not meaningfully improve accessibility unless specifically prompted to do so. Despite generally reputable website sources, limited readability and accountability highlight opportunities for orthopaedic surgeons and organizations to advocate for improved clarity, accessibility, and patient-centered communication.
    Key Concepts: (1)Online patient education materials for inherited pediatric orthopaedic conditions are written at a higher reading level than the recommended 8th grade level.(2)Only 18% of websites and 2% of AI-generated responses meet the ≤8th grade readability standard.(3)Most online resources come from reputable sources (government, academic hospitals, nonprofits) but still have poor readability.(4)AI-generated health information has similar readability to websites but is significantly shorter.(5)Improving readability, accessibility, and patient-centered communication is a key opportunity for orthopaedic providers and organizations.
    Level of Evidence: V.
    Keywords:  Artificial intelligence; Orthopaedics; Patient education; Pediatrics; Readability
    DOI:  https://doi.org/10.1016/j.jposna.2026.100420
  18. Int Dent J. 2026 Aug 13. pii: S0020-6539(26)00356-4. [Epub ahead of print]76(5): 109763
       INTRODUCTION AND AIMS: Video-sharing platforms are key health information sources, but poor oversight and variable quality may mislead patients. This study assesses the quality, reliability, and dissemination of TMD videos on major platforms, and examines whether user engagement aligns with information quality.
    METHODS: A cross-sectional content analysis was conducted of TMD-related videos retrieved from YouTube, TikTok, and Bilibili using platform-specific search strategies. The top 100 videos from each platform were screened, and 260 eligible videos were included. Two independent reviewers assessed information quality using the modified DISCERN (mDISCERN), Global Quality Scale (GQS), and Video Information and Quality Index (VIQI). Video characteristics, uploader type, content features, and engagement metrics were extracted. Between-platform comparisons were performed using nonparametric tests, and associations between quality scores and engagement metrics were examined using Spearman correlation analysis.
    RESULTS: Overall information quality was moderate and varied significantly across platforms. YouTube videos scored higher on mDISCERN, GQS, and VIQI than TikTok and Bilibili (all P < .001). Medical professional videos were of higher quality than those from patients or self-media. However, higher-quality videos did not correlate with greater dissemination; lower-quality videos tended to attract more views, likes, comments, and shares, with only weak correlations observed.
    CONCLUSIONS: TMD-related health information on video-based social media platforms shows substantial variability in quality and a clear mismatch between informational value and public reach. Platform-specific dissemination patterns may allow lower-quality content to achieve substantial reach, underscoring the need for strategies that improve the visibility of credible medical information.
    CLINICAL RELEVANCE: TMD-related video quality varies widely across major social media platforms, and high engagement does not equate to reliability. Joint efforts by clinicians, educators, and platform designers are needed to deliver trustworthy and accessible health information to those who need it most.
    Keywords:  Public health; Social media; Temporomandibular disorder; TikTok; Video quality; YouTube
    DOI:  https://doi.org/10.1016/j.identj.2026.109763
  19. Z Evid Fortbild Qual Gesundhwes. 2026 Aug 12. pii: S1865-9217(26)00176-5. [Epub ahead of print]
       BACKGROUND: YouTube is an important source of health information. The YouTube Health initiative aims to improve access to high-quality information and support informed decision-making. Informed decisions require objective, transparent, and comprehensive information about the benefits and risks of medical interventions. This study systematically evaluated whether German-language YouTube videos on knee osteoarthritis (OA) treatment facilitate informed decision-making.
    METHODS: We conducted a systematic search on YouTube. Eligible health information materials (HIMs) were evaluated using the validated MAPPinfo checklist, which operationalises the guideline Evidence-Based Health Information across four quality dimensions: definition, transparency, content, and presentation. Two raters assessed all HIMs independently. Additional structured content mapping identified the range and balance of treatment options and assessed the reporting of benefits, complications, and anaesthesia for total knee arthroplasty (TKA). Information quality was scored from 0% to 100%, and descriptive analyses examined mean scores by provider type, country, and YouTube Health certification.
    RESULTS: Seventy-nine videos representing 50 HIMs met the inclusion criteria. Most HIMs (n = 23) were provided by hospitals; 28 were certified by YouTube Health. The mean information quality was 15.4% (range 5-25, SD 4.5), with no HIM meeting all criteria. The information quality was similar for certified and non-certified HIMs (15.4% [SD 4.5] and 15.3% [SD 4.4], respectively). Among HIMs mentioning TKA (n = 37), 20 omitted complications and 16 omitted benefits.
    CONCLUSIONS: The quality of German-language YouTube videos on knee OA treatment was insufficient for facilitating informed decisions. The quality criteria for YouTube Health certification did not align with the stated aims.
    Keywords:  Certification; Consumer health information; Evidence-based health information; Gesundheitsinformation; Gesundheitskommunikation; Health communication; Informed choice; Informierte Entscheidung; Quality assessment; Qualität; Social media; Soziale Medien; Zertifizierung
    DOI:  https://doi.org/10.1016/j.zefq.2026.07.003
  20. J Cancer Educ. 2026 Aug 15.
      For medical students, choosing a specialty is a pivotal decision with major professional and personal implications. Career information from popular resources like YouTube may be particularly relevant for radiation oncology, where exposure in medical curricula is limited. This study aims to characterise YouTube videos providing career insights into the field. The first 50 results from six YouTube searches related to radiation oncology careers were programmatically extracted using a Python script, yielding an initial pool of 300 videos. Videos were ranked based on frequency across searches and position in the results. After applying predetermined inclusion criteria, the top 50 videos were analysed using a previously validated video assessment tool. Videos originated from the United States (36/50), India (9/50), Canada (3/50), and Australia (2/50). Most (80%) were released within four years of the search, including 32% in 2020. Half were published by healthcare facilities, with 12 of 25 promoting their respective residency programs. Overall, coverage varied across career needs, with frequent focus on the nature of work (86%), altruism (60%), and intellectual satisfaction (60%), while several aspects, including lifestyle (28%) and salary (18%), received less attention. Radiation oncology careers were often depicted as patient-oriented (74%), technology-driven (72%), compassionate (68%), and meaningful (62%). YouTube videos highlight values inherent to the specialty but underrepresent practical career considerations. These findings characterise one accessible source of radiation oncology career information and identify areas where future resources could provide a more comprehensive overview of the field.
    Keywords:  Career Choice; Medical Education; Medical Students; Radiation Oncology; Social Media; YouTube
    DOI:  https://doi.org/10.1007/s13187-026-02965-3
  21. J Indian Soc Periodontol. 2026 Apr-Jun;30(2):30(2): 269-276
       Background: Gingival recession affects approximately 81% of people worldwide and represents one of the most common oral health issues. Increasingly, more people are using digital platforms to find information about their health, with YouTube represented as a popular source for learning about oral health topics. However, studies have not been conducted on the quality or trustworthiness of gingival recession information that is available on YouTube.
    Objective: To systematically evaluate the content quality, usefulness, and reliability of YouTube™ videos concerning gingival recession and its management using validated assessment tools.
    Materials and Methods: A cross-sectional analysis was conducted using four search terms: "gingival recession," "gum recession treatment," "receding gums," and "gum graft." Twenty videos that met the inclusion criteria were independently evaluated by two calibrated dental evaluators using the Global Quality Scale (GQS), audiovisual quality assessment, and a customized usefulness index. Videos were classified by source (expert vs. non-expert) and content level (high vs. low). Statistical analysis examined the relationships between quality metrics and video characteristics.
    Results: Expert-created videos comprised 85% of the sample, significantly higher than previously reported for other dental topics. The mean GQS score was 3.4 ± 0.9, indicating moderate-to-good overall quality. High-content videos accounted for 70% of the sample and expert videos demonstrated significantly superior quality across all metrics (GQS: 3.5 vs. 2.7, P < 0.05; usefulness index: 8.2 vs. 5.3, P < 0.01). Content analysis revealed strong coverage of treatment procedures (80%) but significant gaps in prevention education (35%) and practical considerations (15%).
    Conclusions: YouTube videos concerning gum recession had better quality in terms of content than did videos pertaining to other dental subjects. Videos found on YouTube were at a high level in terms of professional input and contained a moderate-to-high amount of educational information. There was little or no information provided in YouTube videos related to prevention as well as full patient instruction. Therefore, it appears that while YouTube can be a viable source for providing patients with educational materials, there are some subject matter areas that need greater emphasis by professionals.
    Keywords:  Digital health information; YouTube; gingival recession; patient education; periodontics; social media
    DOI:  https://doi.org/10.4103/jisp.jisp_284_25
  22. Tob Induc Dis. 2026 ;24
       INTRODUCTION: Passive smoking is associated with substantial morbidity and mortality, yet public awareness of its health risks remains limited. As social media increasingly serves as a source of health information, this study evaluated the quality and reliability of passive smoking-related short videos on two major Chinese platforms, Bilibili and Douyin (TikTok in China).
    METHODS: In this cross-sectional study, the top 100 videos for each Chinese search term ('passive smoking' and 'secondhand smoking') were retrieved from Bilibili and Douyin on 10 June 2026. After excluding duplicates, irrelevant videos, promotional videos, and videos shorter than 30 seconds, eligible videos were analyzed for basic characteristics and engagement metrics. Two independent reviewers assessed video quality using the Global Quality Scale (GQS), modified DISCERN (mDISCERN), Journal of the American Medical Association (JAMA) benchmark criteria, and Video Information and Quality Index (VIQI). Comparisons were performed using non-parametric tests, and correlations were examined using Spearman analyses.
    RESULTS: A total of 157 videos were included (59 Bilibili, 98 Douyin). Bilibili videos were significantly longer than Douyin videos (median: 130.0 vs 73.0 s, Z= -3.93, p<0.001), whereas Douyin videos received more comments (median: 89.0 vs 27.0, Z= -2.48, p<0.05) and shares (median: 1829.5 vs 72.0, Z= -4.88, p<0.001). Overall, video quality and reliability were moderate. No significant differences were observed between platforms in GQS, mDISCERN, or JAMA. However, Douyin videos achieved significantly higher VIQI scores than Bilibili videos (median: 12.0 vs 11.0, Z= -3.37, p<0.001). Videos produced by science communicators and doctors consistently demonstrated higher quality than those uploaded by lay users. Engagement metrics were not associated with video quality or reliability.
    CONCLUSIONS: Passive smoking-related videos on Douyin and Bilibili provide moderate-quality health information. Videos from science communicators and doctors were of higher quality than those from lay users. Video popularity does not reflect information reliability.
    Keywords:  Bilibili; TikTok (Douyin); cross-sectional; passive smoking; secondhand smoking
    DOI:  https://doi.org/10.18332/tid/225461
  23. Medicine (Baltimore). 2026 Aug 14. 105(33): e50130
      Short-video platforms are widely used for medical information, but the quality of anal fistula content remains uncertain. We evaluated anal fistula-related videos on TikTok and Bilibili and compared platforms and uploader categories. The top 200 videos under each platform's default comprehensive ranking were screened. Eligible videos were assessed using the Global Quality Score (GQS), modified DISCERN (mDISCERN), Journal of the American Medical Association benchmark criteria (JAMA), and 7-indicator completeness score (7-ICS). Platform and uploader differences were evaluated using nonparametric tests. Spearman correlations examined relationships between engagement and quality, and multivariable ordinal logistic regression identified factors associated with higher scores. Of 400 screened videos, 333 were included (TikTok, 167; Bilibili, 166). Median GQS, mDISCERN, JAMA, and 7-ICS scores were each 2.00. Treatment (71.4%) and clinical manifestations (63.9%) were the most frequently covered domains, whereas epidemiology (7.2%) and prevention (12.6%) were least covered. TikTok videos were shorter and received greater engagement than Bilibili videos (all P < .001). Bilibili had higher unadjusted 7-ICS scores and JAMA score distributions, while GQS and mDISCERN did not differ. After adjustment, TikTok was associated with higher GQS (OR, 2.26; 95% CI: 1.27-4.02) and lower 7-ICS (OR, 0.26; 95% CI: 0.13-0.50), but not with JAMA or mDISCERN. Specialist-physician uploader status and longer video duration were associated with higher scores for all 4 outcomes. Engagement metrics were strongly intercorrelated but were generally unrelated to quality measures. Anal fistula-related videos on both platforms were generally incomplete and of suboptimal quality. Platform differences were mixed rather than uniformly favoring TikTok or Bilibili. Specialist-physician authorship and longer duration were consistently associated with better scores. These cross-sectional findings represent associations, not causal effects.
    Keywords:  Bilibili; TikTok; anal fistula; health communication; reliability; short videos; video quality
    DOI:  https://doi.org/10.1097/MD.0000000000050130
  24. Front Public Health. 2026 ;14 1851530
       Background: Radiation dermatitis (RD), a common adverse reaction to radiotherapy (RT), is widely discussed on video-sharing platforms. However, user-generated content about RD lacks systematic scientific validation. This study evaluates the reliability and quality of RD-related videos on two major social media platforms: Douyin and Bilibili.
    Methods: We conducted a cross-sectional study on Douyin and Bilibili, using keywords including "radiation dermatitis ()," "radiotherapy-induced dermatitis ()," "radiation skin injury ()," and "radiation dermatitis care ()". A total of 190 eligible short videos were included for analysis. We collected data on video characteristics, including title, uploader identity, upload date, video duration, content type, engagement metrics (likes, comments, shares, saves), presentation format, and quality scores. Video quality and reliability were assessed using the modified DISCERN (mDISCERN), the Journal of the American Medical Association (JAMA) benchmark criteria, and the Global Quality Scale (GQS) by two independent reviewers.
    Results: The analysis revealed that Douyin videos were significantly shorter in duration than those on Bilibili (p < 0.05), yet they demonstrated higher audience engagement. Douyin yielded significantly higher median mDISCERN and JAMA scores than Bilibili, while the between-group difference in GQS scores did not reach statistical significance. Videos created by professional healthcare providers demonstrated higher reliability and quality, with mDISCERN scores of 3 and GQS scores of 3. Correlation analysis revealed a strong significant positive correlation between video shares and saves. Engagement metrics showed positive correlations with all three quality scores but with gradient differences in strength: strong correlations with mDISCERN scores, moderate correlations with GQS scores, and weak correlations with JAMA scores.
    Conclusions: Social media platforms support the dissemination of RD-related health information to some extent, but the overall quality of videos remains suboptimal. We recommend that professional healthcare providers obtain official platform certification to promote the spread of high-quality RD-related content.
    Keywords:  health information dissemination; radiation dermatitis (RD); short-form video platforms; user interaction; video quality
    DOI:  https://doi.org/10.3389/fpubh.2026.1851530
  25. Medicine (Baltimore). 2026 Aug 14. 105(33): e50245
      Lipid-lowering therapy is crucial for cardiovascular disease prevention, yet adherence is suboptimal due to concerns about side effects. Short video platforms are key health information sources in China, but content quality regarding pharmacotherapy remains under-investigated. This study aimed to systematically evaluate the quality, reliability, and content of short videos regarding "lipid-lowering drugs" on TikTok and Bilibili. A cross-sectional study was conducted in December 2025. The top 150 videos retrieved via the keyword "lipid-lowering drugs" from each platform were screened. A total of 199 videos meeting inclusion criteria were analyzed. Two independent researchers evaluated uploader profiles, content coverage, and interaction metrics. Video quality was assessed using the Global Quality Scale (GQS) and reliability via the modified DISCERN (mDISCERN) tool. Medical professionals dominated the ecosystem (90%), reaching 100% on TikTok. While videos frequently covered mechanisms (63.8%) and indications (54.3%), critical safety information regarding drug interactions (11.6%) and contraindications (11.1%) was systematically omitted. TikTok videos demonstrated significantly higher GQS scores than Bilibili (median 3.00 vs 2.00, P = .007), yet evidentiary reliability (mDISCERN) remained moderate across both (median 2.00). Crucially, user interaction metrics correlated significantly with production quality (GQS) but showed no association with reliability (mDISCERN). Despite professional dominance, online lipid-lowering drug information suffers from a "safety silence," lacking comprehensive contraindication warnings. A "popularity paradox" exists where engagement reflects presentation skills rather than medical reliability. Healthcare providers must prioritize safety profiles in digital education, and patients should be cautioned against judging accuracy by popularity.
    Keywords:  health communication; infodemiology; lipid-lowering drugs; patient education; short videos; social media
    DOI:  https://doi.org/10.1097/MD.0000000000050245
  26. Medicine (Baltimore). 2026 Aug 14. 105(33): e50274
      Retinal detachment (RD) is a sight-threatening emergency where public awareness of early symptoms is critical. Bilibili and TikTok now serve as prominent health information channels, yet RD-related video content quality and reliability remain unassessed. This cross-sectional study systematically retrieved the top 150 videos using the keyword "retinal detachment" from Bilibili and TikTok in October 2025. Following exclusion criteria application, 224 videos underwent analysis. Two ophthalmologists independently assessed video quality and reliability using the Global Quality Scale (GQS), modified DISCERN, and Journal of the American Medical Association (JAMA) benchmark criteria. Analyses also covered content themes, uploader identity, and correlations with user engagement metrics. Among the 224 videos, overall quality was suboptimal (median scores: GQS = 2, modified DISCERN = 2, JAMA = 2). Treatment information dominated (67.86%), while content on symptoms (48.21%) and particularly prevention (32.59%) was scarce. Although overall scores were low, the distributions of GQS (P = .014) and JAMA (P = .002) scores differed significantly between self-identified professional and nonprofessional uploaders, despite identical group medians. A key finding was the negligible correlation between user engagement metrics and quality scores, indicating that popularity does not reflect content accuracy. Current RD information on short-form video platforms is characterized by generally low quality, significant content imbalance, and a disconnect between engagement and reliability. Urgent collaborative efforts among content creators, platform regulators, and healthcare institutions are needed to establish quality control standards, promote evidence-based content, and enhance public health literacy to leverage social media for improving RD outcomes.
    Keywords:  Bilibili; TikTok; information quality; retinal detachment; short video
    DOI:  https://doi.org/10.1097/MD.0000000000050274
  27. Medicine (Baltimore). 2026 Aug 14. 105(33): e50214
      Viral myocarditis (VM) affects over 1 million people yearly worldwide and may trigger arrhythmia, heart failure, or sudden death, bringing heavy clinical and economic burdens. TikTok and Bilibili are vital health information hubs, yet unregulated VM video content lacks quality evaluation on Chinese platforms. This cross-sectional study assessed VM video reliability and completeness on both platforms to support evidence-based health guidance. Data collection finished on February 15, 2026. We collected the top 100 default-ranked videos per platform via incognito new accounts searching "viral myocarditis." After screening out ads, irrelevant, recent and duplicate videos, 180 clips (TikTok = 91, Bilibili = 89) were analyzed. Two trained cardiologists scored videos with the 0 to 5 modified DISCERN (instrument for evaluating health information quality) scale and 1 to 5 Global Quality Score and graded 6 disease domains' completeness using a 3-point scale. Nonparametric tests and Spearman correlation were adopted, with P < .05 as the statistical significance threshold. TikTok videos had significantly higher median likes, comments, saves, and shares (all P < .001), while Bilibili videos were longer (139 vs 95 seconds, P = .002). Professional creators accounted for 71.43% of TikTok uploads versus 44.94% on Bilibili, which hosted more nonprofessional contributors. TikTok fully covered symptoms best (40.66%) but ignored epidemiology (83.52% uncovered); Bilibili offered more complete etiology, prevention, and symptom explanations yet lacked epidemiological content. Bilibili contained far more high-rating DISCERN (4-5) and 5-star Global Quality Score videos and no 1-point low-quality clips, unlike TikTok. Engagement metrics were strongly positively correlated on both platforms; video duration weakly negatively correlated with interaction on TikTok but moderately positively correlated on Bilibili. Inter-rater reliability was excellent (r = 0.81-0.96). Bilibili's VM videos outperform TikTok in information completeness and overall quality. TikTok's higher traffic often favors low-quality content, and preventive knowledge coverage is insufficient on both platforms. Such gaps arise from distinct user groups, content formats, and algorithms. Optimized algorithms, medical creator certification, and differentiated science communication can standardize VM online education and help patients access reliable disease management knowledge.
    Keywords:  Bilibili; TikTok; information quality; short video; viral myocarditis
    DOI:  https://doi.org/10.1097/MD.0000000000050214
  28. Front Cardiovasc Med. 2026 ;13 1878079
       Background: Short-form video platforms are increasingly used as sources of health information. However, the accuracy and trustworthiness of cardiovascular disease (CVD)-related dietary guidance remains unclear. This study aimed to evaluate the quality, reliability, and content characteristics of CVD-related dietary guidance videos on Douyin, the Chinese version of TikTok.
    Methods: A cross-sectional study was conducted on April 7, 2026, using Douyin as the data source. A total of 98 videos directly related to CVD dietary guidance were included after screening. Baseline characteristics, uploader types, and user engagement metrics were extracted. Video quality and reliability were independently assessed using the Global Quality Score (GQS), the Journal of the American Medical Association (JAMA) benchmark criteria, and the modified DISCERN (mDISCERN) instrument.
    Results: Among the 98 included videos, 59.0% were uploaded by professional individuals and 11.2% by professional institutions. Fruits and vegetables (98.0%) and grains and tubers (87.0%) were the most frequently covered dietary categories, whereas oil and salt restriction was less frequently addressed (47.0%). Videos uploaded by professional institutions and medical professionals achieved higher GQS, mDISCERN, and JAMA scores than those from nonprofessional creators (all P < 0.001). Video quality showed weak positive correlations with engagement metrics, including followers and likes (Spearman r = 0.24-0.31, all P < 0.05). GQS, mDISCERN, and JAMA scores were positively correlated (Spearman r = 0.35-0.40, all P < 0.001).
    Conclusion: Douyin videos provide accessible dietary guidance for CVD, but overall quality and comprehensiveness remain suboptimal. Professional involvement improves reliability and evidence-based content. Regulatory oversight, standardized quality criteria, and targeted education for content creators are needed to ensure accurate dietary information dissemination.
    Keywords:  Douyin; accuracy; cardiovascular disease; diet; nutrition; quality; short videos
    DOI:  https://doi.org/10.3389/fcvm.2026.1878079
  29. Support Care Cancer. 2026 Aug 14. pii: 852. [Epub ahead of print]34(9):
       PURPOSE: Cancer patients are increasingly turning to social media, an indispensable source of health information, for mental health support. However, the quality and reliability of widely available mental health videos remain largely unevaluated. This study aimed to systematically evaluate and compare the quality and reliability of mental health videos on social media platforms.
    METHODS: A cross-sectional study was conducted. We systematically searched and screened the top 100 videos for each of three topics (mindfulness, relaxation techniques, music therapy) on both YouTube and TikTok. Video characteristics, uploader identity, and content format were classified using a standardized framework. Two independent raters assessed video quality using the Global Quality Scale (GQS) and reliability using the modified DISCERN (mDISCERN) instrument. Statistical analyses included Kruskal-Wallis H tests with post-hoc comparisons and K-means clustering analysis.
    RESULTS: A total of 600 mental health-related videos screened. Ultimately, 450 eligible videos (187 from YouTube, 263 from TikTok) were included for analysis. YouTube's ecosystem was more institutional, while TikTok was dominated by general users. Videos from healthcare professionals (doctors and non-physician mental health professionals) demonstrated significantly higher reliability scores on both platforms. Regarding audience engagement, highly popular videos on YouTube exhibited significantly lower reliability, whereas engagement and quality were unrelated on TikTok. Content format performance was platform-specific; integrated guidance was reliable on YouTube but scored lowest on TikTok (GQS, p = 0.001 vs practical guidance). Music therapy videos were consistently the lowest-performing topic.
    CONCLUSION: While healthcare professionals are the most reliable sources of mental health information on both YouTube and TikTok, their content is not necessarily the most visible or engaging. The disconnect between popularity and informational quality highlights the need for critical media literacy among patients and calls for platform-level innovations to prioritise reliable health information.
    TRIAL REGISTRATION: Not applicable.
    Keywords:  Cancer patients; Education; Mental health; Public health; Social media
    DOI:  https://doi.org/10.1007/s00520-026-11078-y
  30. J Thorac Dis. 2026 Jul 31. 18(7): 773
       Background: Chronic obstructive pulmonary disease (COPD) poses a significant public health burden, particularly in China. This study mainly assessed the quality of clinical information and prevalent misinformation related to COPD-related short videos on TikTok, RedNote, and Bilibili.
    Methods: A cross-sectional observational study was conducted using videos collected on November 15, 2025. "Chronic obstructive pulmonary disease" was used as the search keyword on TikTok, RedNote, and Bilibili. After removing duplicate and irrelevant videos, 510 videos (167 from TikTok, 167 from RedNote, and 176 from Bilibili) were included in the analysis. Video characteristics, uploader types, and content categories were documented. The quality of information was assessed using the Global Quality Scale (GQS) and the modified Decision-Making Information Support Criteria for Evaluating the Reliability of Non-randomized Studies (mDISCERN) instrument.
    Results: TikTok videos exhibited significantly higher user engagement (likes, comments, shares, and saves) than RedNote and Bilibili videos (P<0.001). However, regarding information quality, Bilibili videos achieved significantly higher GQS and mDISCERN scores than TikTok and RedNote videos (P<0.001). Content produced by professional institutions and individuals received significantly higher scores than that produced by nonprofessional institutions. A strong positive correlation was found among the popularity metrics, but no significant correlation was found between popularity indicators and quality scores.
    Conclusions: COPD-related short videos displayed uneven clinical information quality despite broad public exposure. Viral popular content cannot guarantee medical authenticity, and widespread misinformation may interfere with regular patient care. Sufficient professional participation and strict platform content supervision are therefore required to standardize online COPD health education and reduce clinical risks.
    Keywords:  Chronic obstructive pulmonary disease (COPD); health information quality; medical science; short video platform
    DOI:  https://doi.org/10.21037/jtd-2026-0907
  31. Pregnancy (Hoboken). 2026 May;2(3): e70301
       Objective: Pregnant people are increasingly utilizing TikTok to understand medical information. However, the accuracy and reliability of this content remain largely unexamined. This study aimed to evaluate videos on prenatal aneuploidy testing by engagement, reliability, and quality, comparing results by content category and account type.
    Study design: Researchers conducted TikTok searches (May-July 2025) using 11 sets of keywords. The first 100 videos for each search were collected and screened with inclusion and exclusion criteria. Videos were manually categorized based on content type and account type (i.e., healthcare professionals, laypeople, or other). Objective data collected directly from TikTok included video length and engagement metrics (views, likes, and comments). Two independent reviewers used the modified DISCERN (mDISCERN) criteria to assess reliability and the Global Quality Scale (GQS) to assess quality. Engagement metrics, mDISCERN, and GQS were compared across video category and account type using Wilcoxon rank-sum test.
    Results: A total of 1100 videos were identified, and 541 videos met the inclusion criteria for full review. A total of 69.9% (n = 378) of the videos were posted by laypeople and 23.8% (n = 129) of the videos were posted by healthcare professionals. Videos posted by healthcare professionals had a significantly higher mDISCERN (median 8, interquartile range [IQR] 7-9) and GQS (median 3.5, IQR 3-4) compared to videos posted by laypeople (mDISCERN = 3, IQR 2-4; GQS = 2, IQR 1.5-2.5), but significantly fewer median views (28,900 vs. 32,150, respectively), p < 0.01. Fearmongering and crowdsourcing videos (posted mainly by laypeople) were among the highest engagement categories, but had low mDISCERN and GQS scores.
    Conclusion: TikTok provides unreliable and low-quality medical information on aneuploidy testing in pregnancy. Healthcare professionals' videos are higher in quality and reliability but have lower views than those by laypeople. These findings highlight the need for strategies to improve the reach of accurate, evidence-based prenatal information on social media platforms.
    Keywords:  TikTok; medical misinformation; patient education; prenatal aneuploidy screening; prenatal genetic testing; social media
    DOI:  https://doi.org/10.1002/pmf2.70301
  32. Cureus. 2026 Jul;18(7): e112677
       INTRODUCTION: Anterior cruciate ligament (ACL) and medial collateral ligament (MCL) injuries are commonly encountered injuries in orthopedic practice. Simultaneously, patients are utilizing online resources to better understand diagnoses and treatment options. As such, understanding public interest and the accessibility of educational materials can provide insight into information-seeking behaviors and gaps in health literacy. This study used Google Trends (Google Limited Liability Corporation (LLC), Mountain View, CA) data to evaluate temporal patterns in public interest regarding ACL and MCL injuries and surgeries and to assess the readability of commonly accessed online educational materials, with the hypothesis that ACL injury and surgery-related terms would generate greater search interest and that the readability of articles for both ACL surgery and MCL surgery will exceed the American Medical Association's (AMA)-recommended sixth-grade reading level.
    METHODS: Google Trends was used to extract relative search volume (RSV) data in the United States from January 2004 through December 2025 for ACL- and MCL-related injury and surgical search terms. Injury terms (e.g., "ACL tear," "MCL tear") and surgical terms (e.g., "ACL reconstruction," "MCL surgery") were queried. Monthly and seasonal variations in RSV were evaluated for ACL and MCL surgery terms. Readability analysis was performed on the first 25 text-based Google search results for "ACL surgery" and "MCL surgery" using Flesch Reading Ease and Flesch-Kincaid Grade Level metrics.
    RESULTS: Among injury-related terms, "ACL tear" demonstrated the highest mean RSV compared to the lowest searched term (p < 0.001). Both "ACL tear" and "MCL tear" showed significant increases in RSV from 2004 to 2025 (p < 0.001). Among surgical terms, "ACL surgery" had the highest mean RSV overall, while "MCL surgery" was the most frequently searched MCL-related surgical term. RSV for both ACL and MCL surgery increased significantly over time. No significant differences in RSV were measured across seasons or months for either ACL surgery or MCL surgery. On readability analysis, the mean Flesch-Kincaid Grade Level was 10.4 for ACL surgery materials and 10.1 for MCL surgery materials, exceeding AMA recommendations. Only 4% of ACL surgery articles and 0% of MCL surgery articles met the recommended sixth-grade reading level.
    CONCLUSION: Public interest in ACL and MCL injuries, tears, and surgical treatments has substantially increased in the past two decades, with ACL injury and surgery-related terms generating greater search interest than MCL injury and surgery-related terms, respectively. Patients preferentially searched for "tear" terminology rather than "injury" terminology, and "surgery" was the most common surgical search term. Despite growing public interest in ACL and MCL information, the readability of online educational materials remains significantly above recommended levels, possibly limiting patient comprehension. Future efforts should focus on improving the accessibility and readability of online orthopedic educational resources.
    Keywords:  anterior cruciate ligament (acl); health care literacy; medial collateral ligament injury; orthopedic sports medicine; patient education; readability score
    DOI:  https://doi.org/10.7759/cureus.112677