bims-librar Biomed News
on Biomedical librarianship
Issue of 2026–09–13
35 papers selected by
Thomas Krichel, Open Library Society



  1. F1000Res. 2026 ;15 43
      Recognizing and appreciating the diversity of people around the world is a necessary step in achieving equality for all. The opening of equitable chances for society as a whole is one of the ways in which libraries contribute to the reduction of the exclusion gap. Despite their expanding relevance, institutions like libraries demonstrate inconsistent patterns of integration and research. This literature analysis aimed to examine the methods employed by public libraries to foster social inclusion across various populations from 2014 to 2024. A comprehensive systematic literature review was performed, identifying n = 69 interventions that fulfilled the inclusion criteria according to PRISMA methodological principles. The findings indicate a necessity for collaboration between libraries and other organizations to enhance their effectiveness in promoting social inclusion. The findings underscore the necessity of adopting a comprehensive perspective on inclusion in libraries, which encompasses an examination of infrastructure and participation in physical spaces, as well as additional aspects and dimensions, integrating both physical and digital inclusion.
    Keywords:  Community; Education; Inclusivity; Library; knowledge
    DOI:  https://doi.org/10.12688/f1000research.175122.2
  2. Account Res. 2026 Sep 09. 2725029
       OBJECTIVE: Little attention has been paid to how publication and access models influence the evidence that AI systems retrieve, process, and summarize. This article introduces the concept of AI-induced evidence skew, proposes four pathways through which it may arise, and provides a preliminary empirical illustration.
    METHODS: Nine widely used large language models (LLMs) were prompted to generate literature-supported content. The accessibility status of the references provided was compared with the percentage of open-access records identified through searches of three scholarly databases.
    RESULTS: AI-induced evidence skew is a structural distortion in AI-assisted evidence synthesis whereby outputs disproportionately represent open-access or otherwise machine-accessible literature. Four proposed pathways are differential representation of literature in training data, retrieval-access constraints, human verification practices, and system- or user-imposed access restrictions. Among 78 verifiable references generated by LLMs, a mean of 95.38% had freely available full text (87.43% were open-access), compared with a mean open-access percentage of 57.83% across scholarly database search results.
    CONCLUSIONS: Although exploratory and non-generalizable, these findings support the hypothesis that AI-assisted search and synthesis may overrepresent publicly accessible literature. The concern is not the quality of open-access literature, but that accessibility may become an unintended determinant of included evidence. Mitigation requires greater awareness, human oversight, transparent reporting of AI use and retrieval limitations, verification of AI-generated references, and improved transparency regarding AI evidence coverage. Further empirical investigation across disciplines, platforms, and AI tools is warranted.
    Keywords:  Artificial intelligence in publishing; biomedical publishing; large language models; publication models; research integrity
    DOI:  https://doi.org/10.1080/08989621.2026.2725029
  3. F1000Res. 2025 ;pii: Chem Inf Sci-260. [Epub ahead of print]14
      Effective research depends on building on the knowledge found in the scientific literature. Designed to streamline literature tasks, the EPA's Abstract Sifter literature tool, now at version 8, has been continually extended and enhanced since its introduction in 2017[1]. Early enhancements to the tool have primarily focused on core tasks common to all researchers. For example, citation retrieval from PubMed has been made faster and the returned citation threshold increased to 10,000. Features that allow deeper examination of the literature have been introduced as well. A functionality called Term-mapping allows for fast, dynamic relevancy ranking of returned citations. MeSH substances, such as proteins, genes, and chemicals, can now be extracted from a retrieved corpus of citations, ranked by frequency and explored through the MeSHMine functionality. Features that facilitate user engagement with publications have also been improved: formatting and colorization ease reviewing of the abstract text and the tagging and noting citations functionality has been streamlined. Version 8 introduced multiple features that break new ground in working with chemical literature. For example, chemical entity extraction from scientific publications has been streamlined through download of PDFs and automated table extraction. Following entity extraction, the chemical names can be used as inputs to retrieve EPA's chemical identifiers, the DSSTox (Distributed Structure-Searchable Toxicity) chemical IDs (DTXSIDs). Once these identifiers have been retrieved, a wealth of chemical information is available through built-in functions accessing EPA's Computational Toxicology and Exposure application programming interface (CTX-APIs) [2]. This new functionality allows researchers to build on the EPA's efforts in chemical data assembly and curation. The Abstract Sifter version 8 is a valuable tool for researchers endeavoring to understand chemicals and their effects on the environment and biological systems.
    Keywords:  DSSTox; Literature mining; PubMed; drug discovery; knowledge mining; toxicology
    DOI:  https://doi.org/10.12688/f1000research.160617.2
  4. PLOS Glob Public Health. 2026 ;6(9): e0007257
      Discovering, selecting, and citing sources to craft claims is fundamental to public health research and writing. Sources influence our research questions, shape our methodological approaches, connect our work to the broader public health literature, and affect how successfully our writing persuades readers that our research is valuable and that our arguments are sound. Despite the importance of source and citation practices to successful scientific writing, many student researchers receive limited or no formal training on these practices, particularly on drawing from sources to build persuasive arguments. Without specific training, many student researchers struggle to use sources effectively and often encounter pitfalls such as citation bias, quotation errors, and academic plagiarism. In this article, we provide practical considerations for effective source use and citation practices for academic papers. We highlight best practices for source discovery, from strategies for conducting comprehensive literature searches to active reading and synthesis techniques that inform academic thinking and writing. Given the rise of AI, including generative AI and AI-powered research tools, we address both the potential benefits such tools can offer and the potential pitfalls arising from uncritical use. Finally, we discuss selecting and using sources to craft and position claims, including choosing sources to support different types of claims, building credible arguments, and framing citations to reflect a stance or judgment. In this article, we offer practical introductory advice for student researchers to strengthen their source use and citation practices. Using this guidance as a starting point, student researchers can improve the quality of their publications and refine their reading, critical thinking, and writing skills which will serve them throughout their research careers.
    DOI:  https://doi.org/10.1371/journal.pgph.0007257
  5. Surg Innov. 2026 Sep 09. 15533506261488588
      BackgroundPatients increasingly consult artificial intelligence (AI) tools for breast cancer information. While Large Language Models (LLMs) enhance information accessibility, their accuracy, reliability, and alignment with patient health literacy remain critical concerns. This study compared the quality, reliability, and readability of breast cancer-related responses generated by ChatGPT and Gemini.MethodsIn this cross-sectional study, conducted between March 20 and March 31, 2026, 40 questions spanning diagnosis, treatment, genetics, and follow-up were submitted to ChatGPT-5.3 and Gemini 3.0 Flash. Three independent surgeons evaluated the responses in a double-blinded manner using the modified DISCERN (mDISCERN) for reliability and Global Quality Score (GQS) for content quality assessment. Readability was assessed via Flesch Reading Ease (FRES), Flesch-Kincaid Grade Level (FKGL), Gunning Fog Index (GFI), and Simple Measure of Gobbledygook (SMOG) indices.ResultsGemini demonstrated statistically significant superiority over ChatGPT in both GQS (4.35 ± 0.30 vs 3.90 ± 0.29; P < .001) and mDISCERN (3.63 ± 0.58 vs 3.02 ± 0.48; P < .001) scores. In readability analysis, Gemini exhibited higher FRES (53.90 vs 45.13) and lower FKGL (9.31 vs 10.91) values, indicating enhanced patient accessibility (P < .05). For both models, the "Diagnosis" category yielded the highest readability, whereas "Treatment" scored the lowest. Inter-rater reliability for mDISCERN was moderate (ICC = 0.595).ConclusionsGemini significantly outperforms ChatGPT in response quality, reliability, and linguistic accessibility for breast cancer education. However, both models exceed the recommended sixth-grade reading level, indicating suboptimal optimization for general health literacy. While LLMs serve as promising auxiliary tools, expert supervision and cross-validation remain mandatory to ensure patient safety.
    Keywords:  ChatGPT; Gemini; breast cancer; health literacy; readability
    DOI:  https://doi.org/10.1177/15533506261488588
  6. Br J Oral Maxillofac Surg. 2026 Jul 11. pii: S0266-4356(26)00186-5. [Epub ahead of print]
      Clear and understandable patient information is crucial for both informed consent and postoperative recovery following elective orthognathic surgery. Despite National Health Service (NHS) guidance recommending suitability for the reading level of an 11-year-old, current patient information leaflets (PILs) do not meet this standard. This study was undertaken to assess the readability of NHS orthognathic surgery patient information and whether large language models (LLMs) could improve readability to meet the recommended levels. A cross-sectional analysis was conducted of online PILs on elective orthognathic surgery from United Kingdom NHS hospital websites. Readability was assessed using five validated scoring systems: Flesch Reading Ease Score (FRES), Flesch-Kincaid Grade Level (FKGL), Gunning Fog Index (GFI), Coleman-Liau Index (CLI), and Simple Measure of Gobbledygook (SMOG). Three LLMs, ChatGPT, Claude, and Gemini, were prompted to simplify PILs to the reading age of an 11-year-old. Of 111 UK Hospital Trusts with Maxillofacial Surgery Units, 73 (65.8%) offered elective orthognathic surgery, of which 21 (28.8%) provided online PILs that yielded 24 eligible texts. Original PILs fell short of the recommended NHS readability standards on 4/5 indices, achieving acceptable levels only on FKGL (6.23 ± 1.26). Following revision, all three LLMs significantly improved readability across all indices (p < 0.001). Gemini achieved the strongest performance, meeting the target thresholds on 3/5 indices (FRES 80.35 ± 4.31, FKGL 4.73 ± 0.79, CLI 6.77 ± 0.59). Despite significant improvements, no model consistently met all NHS standards. NHS orthognathic surgery PILs are commonly written above the recommended readability levels. LLMs can improve comprehension; however, unrestricted LLM-generated materials still fail to meet the recommended NHS literacy standards, highlighting the need for targeted frameworks to improve readability.
    Keywords:  ChatGPT; Claude; Elective orthognathic surgery; Health literacy; Large language models; Readability
    DOI:  https://doi.org/10.1016/j.bjoms.2026.06.007
  7. Front Public Health. 2026 ;14 1831988
       Purpose: This study aimed to evaluate and compare the quality, reliability, and readability of patient education materials on endometrial cancer surgery generated by ChatGPT (GPT-5) and DeepSeek (R1).
    Materials and methods: This cross-sectional study analyzed the responses generated by ChatGPT and DeepSeek to totally 41 questions covering four domains: surgical planning, preoperative evaluation, postoperative care, and long-term follow-up. Reliability was assessed through the DISCERN and EQIP instruments, quality was evaluated by the Global Quality Score (GQS), and readability was analyzed by the Flesch Reading Ease Score (FRES), Gunning Fog Index (GFI), and Flesch-Kincaid Grade Level (FKGL). Statistical comparisons were performed by using paired t-tests and Wilcoxon signed-rank tests.
    Results: The two large language models (LLMs) generated education materials of comparable quality, as reflected in GQS scores (median: DeepSeek vs. ChatGPT 5.00 vs. 4.67, p = 0.077). DeepSeek demonstrated statistically significantly higher reliability scores on both DISCERN and EQIP instruments (both p < 0.001). Readability scores (FRES, GFI) were similar between groups, while DeepSeek exhibited a higher FKGL (10.28 vs. 8.84, p < 0.001), indicating the greater text complexity. Subgroup analysis showed that DeepSeek performed better in terms of reliability in the postoperative care and long-term follow-up domains, while ChatGPT exhibited better readability in the surgical planning domain.
    Conclusion: Both DeepSeek and ChatGPT can generate patient education text drafts that are commendable in their structural coherence and linguistic clarity. DeepSeek demonstrates a significant advantage in information reliability, particularly excelling in postoperative and follow-up management content. ChatGPT shows a slight edge in the readability of surgical planning sections. However, the text readability of both models exceeds the general public's health literacy level. This indicates that large language models can only serve as auxiliary tools for generating patient education materials. Their outputs must undergo review by clinical experts and readability optimization to ensure both accuracy and comprehensibility of the information.
    Keywords:  ChatGPT; DeepSeek; artificial intelligence; endometrial cancer; health literacy; large language models (LLMs); patient education; readability
    DOI:  https://doi.org/10.3389/fpubh.2026.1831988
  8. Front Public Health. 2026 ;14 1935121
       Objective: To compare the safety, accuracy, empathy, reliability, information quality, and readability of five publicly accessible large language model chatbots when answering patient-facing lung cancer prognostic questions under standardized single-turn English prompting.
    Methods: In this Chatbot Health Advice Reporting Transparency-guided cross-sectional evaluation, 53 standardized English prompts were submitted once to ChatGPT, Gemini, Copilot, DeepSeek, and Doubao through official web interfaces during April 1-21, 2026. Five blinded raters assessed 265 responses for safety, accuracy, empathy, DISCERN, EQIP, JAMA benchmark criteria, Global Quality Scale, and readability. Paired repeated-measures analyses were used.
    Results: Inter-rater agreement was good to excellent. Safety differed significantly across models (Cochran's Q = 14.089, df = 4, p = 0.007). Gemini generated the highest proportion of safe responses (48/53, 90.6%), whereas DeepSeek generated the lowest (33/53, 62.3%). The only adjusted pairwise safety difference that remained significant was Gemini versus DeepSeek (adjusted p = 0.023). Accuracy, empathy, reliability, information quality, and readability also differed significantly across models (all p < 0.001). Gemini showed the most favorable descriptive profile for safety, accuracy, empathy, and reliability, while Copilot produced the most readable responses.
    Conclusion: Public-facing chatbots differed substantially in safety, reliability, communication quality, and readability. These findings are time-, interface-, and prompt-dependent. Chatbots may support general patient education but should not replace individualized clinician-led prognostic communication.
    Keywords:  empathy; generative AI chatbots; large language models; lung cancer; prognostic information; readability; reliability; safety
    DOI:  https://doi.org/10.3389/fpubh.2026.1935121
  9. Front Public Health. 2026 ;14 1913260
       Background: Search-enabled large language model interfaces are increasingly used by the public for health information, but their performance in mpox-related public health consultation remains unclear. This study evaluated their safety, accuracy, empathy, reliability/information quality, and readability.
    Methods: We conducted a single-query comparative cross-sectional evaluation using 52 predefined mpox-related public consultation questions. Each question was submitted once to each of six search-enabled LLM interfaces, yielding 312 first responses. Responses were assessed against a guideline-based reference framework. Safety was coded as a binary outcome, while accuracy and empathy were rated on 5-point scales. Reliability/information quality was evaluated using DISCERN, EQIP, JAMA benchmark criteria, and GQS. Readability was assessed using six established readability indices. Five trained raters independently evaluated the human-scored outcomes.
    Results: Unsafe responses were relatively infrequent but occurred in all six interfaces, with safe-response rates ranging from 86.5 to 92.3%. No pairwise difference in Safety remained statistically significant after Benjamini-Hochberg correction. Overall differences across interfaces were statistically significant for Accuracy, Empathy, all four reliability/information quality measures, and all six readability indices. Benjamini-Hochberg-adjusted post hoc analyses identified outcome-specific pairwise differences, although the pairwise patterns varied across measures.
    Conclusion: The evaluated search-enabled LLM interfaces showed heterogeneous performance across safety, accuracy, empathy, reliability/information quality, and readability. Although unsafe responses were relatively uncommon, potentially harmful outputs occurred in every interface. These findings support the need for guideline-based evaluation, source transparency, readability optimization, and robust safety safeguards when such interfaces are evaluated or considered for mpox-related public health consultation. The results represent a time- and configuration-specific interface-level snapshot; they should not be attributed to the underlying base models in isolation or interpreted as establishing reproducible performance or a stable hierarchy across sessions, versions, or settings.
    Keywords:  artificial intelligence; digital public health; health information quality; large language model; monkeypox; mpox; public health consultation; readability
    DOI:  https://doi.org/10.3389/fpubh.2026.1913260
  10. Cureus. 2026 Aug;18(8): e114235
       BACKGROUND: Limb lengthening requires patients to understand a complex procedure, a prolonged recovery, and potential complications. This study evaluated the quality, reliability, readability, and clinical completeness of ChatGPT-3.5 (OpenAI, San Francisco, CA) responses to common patient questions about limb lengthening.
    METHODOLOGY: In this descriptive, cross-sectional content analysis, a purposive sample of 10 commonly encountered and clinically relevant questions was identified from the frequently asked questions pages of 10 healthcare institutions. Each question was entered separately into ChatGPT-3.5 on January 18, 2024, without follow-up prompts. Two senior authors independently evaluated each response against published literature and scored it with the Quality Criteria for Consumer Health Information (DISCERN) instrument, Journal of the American Medical Association (JAMA) benchmark criteria, and Flesch-Kincaid grade level. Outcomes were summarized descriptively, and interrater agreement was assessed with the Cohen kappa coefficient.
    RESULTS: The mean DISCERN score was 35.5 (range: 21-42). One response was rated very poor, seven were rated poor, and two were rated fair. Every response received a JAMA benchmark score of 0. The mean Flesch-Kincaid grade level was 13.6 (range: 12.3-15.7), indicating postsecondary reading demand.
    CONCLUSIONS: ChatGPT-3.5 provided broad introductory information, but most responses lacked transparent sources, required clinical clarification, and were too complex for general patient education. It may be useful as a clinician-reviewed adjunct but should not replace individualized counseling by a qualified healthcare professional.
    Keywords:  chatgpt in healthcare; limb lengthening procedures; patient education; post-op pain management; post-procedure
    DOI:  https://doi.org/10.7759/cureus.114235
  11. Knee. 2026 Sep 11. pii: S0968-0160(26)00317-0. [Epub ahead of print]63 104635
       PURPOSE: To compare the information quality, accuracy, and readability of patient-directed responses generated by large language models (LLMs), including ChatGPT-o3, ChatGPT-5.2, Gemini 3, and DeepSeek, regarding robotic-assisted total knee arthroplasty (RA-TKA).
    METHODS: Thirty frequently asked patient questions were identified using LLM outputs and Google search queries. Responses were evaluated for information quality using the DISCERN and Quality Analysis of Medical Artificial Intelligence (QAMAI) instruments, for clinical accuracy using a 5-point ordinal rating scale, and for understandability and readability using the PEMAT Understandability and Flesch-Kincaid Reading Ease scores.
    RESULTS: Median DISCERN scores were 46.0 (range, 35.0-50.0) for ChatGPT-o3, 45.75 (28.5-51.0) for ChatGPT-5.2, 43.75 (32.0-47.5) for Gemini 3, and 42.0 (32.0-50.0) for DeepSeek, with a significant overall difference among models (p < 0.001). The 5-point clinical accuracy scores were similar across models (median 4.0; p = 0.636). Median QAMAI scores were 23.0 for all four models, without a significant between-model difference (p = 0.462). PEMAT Understandability scores differed significantly among models (p < 0.001), with median scores of 90.0 for ChatGPT-o3 and ChatGPT-5.2, 88.0 for Gemini 3, and 85.0 for DeepSeek. Flesch-Kincaid Reading Ease scores also differed significantly (p < 0.001); Gemini 3 demonstrated higher readability than both ChatGPT models, whereas DeepSeek demonstrated higher readability than ChatGPT-o3.
    CONCLUSION: The evaluated LLMs demonstrated generally acceptable clinical accuracy but differed across measures of written information quality, understandability, and readability. No significant difference was detected using QAMAI. Although Gemini 3 and DeepSeek demonstrated greater readability in selected comparisons, median responses across all models remained above recommended patient-education reading levels. LLM-generated responses should therefore be regarded as supplementary rather than standalone sources of patient information regarding RA-TKA.
    Keywords:  Arthroplasty; Artificial intelligence; Large language models; Patient information; Robotic knee arthroplasty
    DOI:  https://doi.org/10.1016/j.knee.2026.104635
  12. Musculoskeletal Care. 2026 Sep;24(3): e70260
       INTRODUCTION: Patient education handouts are used to enhance patient knowledge and support decision making, although many handouts do not meet recommended readability benchmarks. For people with hallux valgus, there is a dearth of co-designed, credible, evidence-based and readable resources to enhance health literacy. The aim of this study was to co-design an evidence-based, clinically accurate handout for hallux valgus targeting a readability level of grade seven or below.
    METHODS: Between October 2023 and August 2024, a descriptive qualitative design was used across two stages. Stage 1 involved focus groups with six health professionals and two individuals with hallux valgus to inform content development. Stage 2 presented a prototype to five health professionals and 10 individuals with hallux valgus for feedback on clarity, completeness and accuracy. Readability was assessed using Flesch Reading Ease and Flesch-Kincaid Grade Level tools.
    RESULTS: The handout included information about the characteristics of hallux valgus and a balance between non-surgical and surgical approaches, with transparency about limited evidence for some treatments. Participants emphasised the need for clarity, relevance, and accessibility and reinforced that treatment decisions, especially surgery, are often not urgent and can evolve over time. The handout is a starting point for discussion and not a standalone resource. The final version meets the readability standards for ages 9-10 years. The readability was further enhanced with the use of visuals and adherence to accessibility guidelines.
    CONCLUSION: A co-designed, evidence-based handout was produced that meets recommended readability and accessibility standards for consumer health information.
    Keywords:  hallux valgus; patient education handout; qualitative research
    DOI:  https://doi.org/10.1002/msc.70260
  13. Front Public Health. 2026 ;14 1942399
       Background/objectives: Patients increasingly use large language model (LLM) interfaces for health information, but their safety and quality for public questions about interstitial cystitis/bladder pain syndrome (IC/BPS) remain uncertain. This study evaluated the safety, accuracy, empathic communication, information quality, reliability, and readability of five publicly accessible LLM interfaces.
    Methods: In this CHART-guided cross-sectional comparative study, 58 public-facing IC/BPS questions were submitted once, in English, to ChatGPT, Gemini, Microsoft Copilot, DeepSeek, and Doubao using a standardized single-turn, zero-shot protocol. Three blinded senior urologists independently assessed safety, accuracy, empathic communication, DISCERN, EQIP, JAMA benchmark criteria, and Global Quality Score. Six readability indices were calculated.
    Results: Inter-rater agreement was significant for all manually assessed metrics. Fleiss' kappa for safety was 0.822, and ICC(2,1) values for other rater-assessed metrics ranged from 0.761 to 0.848. Unsafe responses occurred in all interfaces, ranging from 5.2% for ChatGPT to 8.6% for DeepSeek and Doubao, without a significant between-interface difference (Cochran Q = 1.000, p = 0.910). Accuracy and empathic communication differed significantly across interfaces (both p < 0.001). ChatGPT had the highest median accuracy score, whereas DeepSeek had the highest empathic communication score. Information-quality and reliability scores also differed significantly (all p < 0.001); ChatGPT achieved higher DISCERN, EQIP, and GQS scores, while Gemini achieved higher JAMA scores. Readability differed significantly across interfaces, but none met predefined patient-facing readability benchmarks.
    Conclusion: Publicly accessible LLM interfaces showed domain-specific differences when answering IC/BPS-related public questions. Unsafe responses were uncommon but present in all interfaces, and no interface consistently outperformed the others. LLM interfaces may support general IC/BPS education and question preparation but should not replace clinician-led evaluation or individualized medical advice.
    Keywords:  accuracy; bladder pain syndrome; empathy; interstitial cystitis; large language models; readability; reliability; safety
    DOI:  https://doi.org/10.3389/fpubh.2026.1942399
  14. Ann Transl Med. 2026 Aug 31. 14(4): 48
       Background: Patients increasingly use large language models (LLMs) to obtain medical information, including preoperative guidance. However, the quality of LLM-generated responses to questions regarding anterior cervical discectomy and fusion (ACDF) remains unclear. This study evaluated the performance of GPT-5 and GROK 4 in answering common preoperative ACDF questions. We hypothesized that both models would provide generally accurate and comprehensible responses while demonstrating limitations in completeness, clinical nuance, and personalization.
    Methods: As an expert-rated pilot study, eighteen frequently asked ACDF preoperative questions were independently entered into GPT-5 and GROK 4 without additional prompting. Responses were evaluated by three attending spine neurosurgeons and one neurosurgery resident using 5-point Likert scales for accuracy, completeness, and comprehensibility. Inter-rater reliability was assessed using intraclass correlation coefficients (ICCs) and ordinal Krippendorff's alpha. Readability was measured using the Flesch-Kincaid Grade Level (FKGL).
    Results: Mean accuracy scores were identical for GPT-5 and GROK 4 (4.52). GROK 4 demonstrated higher completeness (4.61 vs. 4.26), whereas GPT-5 showed higher comprehensibility (4.74 vs. 4.41). Inter-attending ICC values were low (0.12-0.21), while Krippendorff's alpha indicated very high ordinal consistency despite restricted score variance. Agreement between the attending mean score and resident ratings was moderate (ICC ≈0.51-0.52). FKGL scores were approximately 10 for both models.
    Conclusions: LLMs provide generally accurate and understandable responses to common ACDF preoperative questions but lack nuanced risk stratification and individualized clinical guidance. These findings support their use as adjuncts, rather than replacements, for physician-led counseling.
    Keywords:  Large language models (LLMs); anterior cervical discectomy and fusion (ACDF); artificial intelligence
    DOI:  https://doi.org/10.21037/atm-2026-0085
  15. Patient Educ Couns. 2026 Sep 04. pii: S0738-3991(26)00369-1. [Epub ahead of print]153 109836
      Low back pain (LBP) and associated radicular leg pain are among the most prevalent musculoskeletal conditions worldwide. Despite clinical guidelines promoting evidence-based management strategies, online health information remains of inconsistent quality, with potential implications for care outcomes.
    OBJECTIVE: To conduct a systematic evaluation of the credibility, information presentation, accuracy, and comprehensiveness of French-language, noncommercial websites providing treatment information on LBP and radicular pain.
    METHODS: We conducted in August 2022 systematic Google searches across country-specific domains in France, Belgium, Switzerland, and Canada. We included freely accessible, French-language webpages that provided at least one treatment recommendation for LBP or radicular pain. Commercial webpages and those requiring registration were excluded. Two independent reviewers assessed each webpage using the Journal of the American Medical Association benchmark criteria for credibility and a modified Netscoring framework, a tool designed to evaluate the quality and presentation of online health information. Treatment recommendations were coded against a synthesis of clinical guidelines to evaluate accuracy and comprehensiveness.
    RESULTS: Overall, 88 webpages were included, 41% having been updated since 2020. Comprehensiveness varied widely: only 41% of LBP recommendations and 26% of radicular pain recommendations were addressed. Absolute accuracy of webpage recommendations was 22%, with institutional webpages showing greater accuracy than mainstream sources (28% versus 15%, p < 0.001). Pharmacological therapies were often inappropriately promoted. Concerning credibility, only a small minority of the webpages (19%) disclosed potential conflicts of interest, and 60% presented bibliographic references.
    CONCLUSIONS: Most French-language noncommercial websites provide incomplete and frequently inaccurate treatment information on LBP and radicular pain. Institutional websites generally offer better quality content, but significant gaps remain - particularly for radicular pain.
    PRACTICE IMPLICATIONS: Health professionals should be aware of the poor quality of online French-language information on LBP and guide patients toward trustworthy resources. Developing validated, user-friendly platforms could help reduce misinformation and support evidence-based self-management.
    Keywords:  Internet-based intervention; Low back pain; Medical informatics; Patient education as topic; Radicular pain
    DOI:  https://doi.org/10.1016/j.pec.2026.109836
  16. Physiother Res Int. 2026 Oct;31(4): e70338
       BACKGROUND AND PURPOSE: For people with bronchiectasis, health information, including guidance on physiotherapy management, is readily accessible online. However, the content and quality of web-based resources, and the enablers and barriers to implementing this information into clinical care, remain unclear. This study aimed to: (1) identify web-based resources providing information on the physiotherapy management of bronchiectasis; (2) evaluate their quality; and (3) explore enablers and barriers to applying web-based information in clinical practice.
    METHODS: A three-stage study was undertaken. Stage 1 involved a search of the World Wide Web using common search engines, during which 3600 web-based resources were screened. Stage 2 evaluated eligible resources using the Health Information Website Evaluation Tool. Stage 3 comprised a nested scoping review of electronic databases to identify published studies evaluating the use or implementation of web-based resources for bronchiectasis in clinical practice.
    RESULTS: Nine web-based resources met the inclusion criteria and consistently included information relevant to the physiotherapy management of bronchiectasis, particularly airway clearance techniques and exercise. Six were rated as good quality; two were moderate and one was poor. Three studies met the inclusion criteria for the nested scoping review, including one conference abstract. Across these studies, enablers included comprehensive content, engaging resources, and ease of use, whereas barriers included access difficulties, technology limitations, and limited time for clinical implementation.
    DISCUSSION: The identified web-based resources provide physiotherapy-specific information that may assist clinicians in selecting resources to complement individualised physiotherapy care, reinforce airway clearance and exercise instruction, and support patient education, and self-management in bronchiectasis. Findings from the nested scoping review suggest that addressing access and technology limitations, together with allowing sufficient time for implementation, may facilitate the integration of these resources into clinical care.
    Keywords:  bronchiectasis; health education; internet; physical therapy modalities
    DOI:  https://doi.org/10.1002/pri.70338
  17. Laryngoscope. 2026 Sep 09.
       OBJECTIVE: To assess the readability and quality of online patient education materials regarding thyroid radiofrequency ablation.
    METHODS: Google searches were conducted using six search terms related to thyroid radiofrequency ablation, and the top 50 webpages from each search were reviewed. Readability was assessed using eight widely used readability formulas, and quality was independently assessed by two raters using the validated DISCERN instrument. Webpages were categorized by source type and geographic origin.
    RESULTS: A total of 63 webpages were included, comprising academic (48%), private (38%), and industry (14%) sources. The majority (79%) originated from the United States. The mean grade reading level was 11.89 ± 1.67, and no individual webpage was written at or below the recommended sixth-grade level. Academic webpages were significantly more readable than private and industry sources across multiple readability formulas. Regarding quality, the mean total DISCERN score was 41.29 ± 7.79 out of 80, indicating poor to fair quality. Mean DISCERN subscores were 15.4 out of 40 for reliability, compared with 23.1 out of 35 for quality.
    CONCLUSION: The majority of webpages were written well above the average reading level of US adults, creating significant barriers to comprehension. In addition, many webpages scored poorly on reliability, demonstrated bias, and inadequately addressed procedural risks. Deficits in either readability or quality may undermine informed decision-making and informed consent. Artificial intelligence and validated quality assessment tools can help clinicians develop online materials that meet readability and quality standards.
    LEVEL OF EVIDENCE: N/A.
    Keywords:  health literacy; patient education; quality assessment; readability assessment; thyroid radiofrequency ablation
    DOI:  https://doi.org/10.1002/lary.70876
  18. Cureus. 2026 Aug;18(8): e114029
      Introduction Families often seek medical information online; this study assesses the accessibility of online pediatric skull fracture materials by evaluating the quality and readability of easily obtained sources. Methods Six terms were queried using Google (Google LLC, Mountain View, California, United States), Bing (Microsoft Corporation, Redmond, Washington, United States), and Yahoo (Yahoo Inc., New York City, New York, United States). Twenty-four unique websites were categorized as commercial, academic, or medical practice. Readability was assessed using Flesch Reading Ease Score (FRES) and Flesch-Kincaid Grade Level (FKGL), while quality was measured using Quality Evaluation Scoring Tool (QUEST) and the DISCERN tool. Analysis of variance (ANOVA) compared readability and quality scores by authorship classification. Linear regression was performed for QUEST vs DISCERN scores and quality vs readability. Results Among 24 unique sites, 8 (33%) were commercial, 10 (42%) from academic institutions, and 6 (25%) from medical practices. Mean FRES and FKGL scores corresponded to an eighth-grade reading level. Mean DISCERN and QUEST scores were 48.8 ± 8.7 (Fair) and 15.0 ± 6.3 (Fair). ANOVA indicated no significant difference in readability scores by website type (p > 0.05). Linear regression revealed a significant correlation between QUEST and DISCERN scores (p = 0.011) but no significant correlation between readability and quality metrics (p = 0.31). Conclusions Online materials do not meet the American Medical Association's recommended sixth-grade level. Sites lacked proper source attribution and clear explanations of treatment options. QUEST and DISCERN both measure quality in comparable ways. Readability and quality were not significantly linked. Complex language, reflected in high readability scores, may hinder health literacy, especially among lower socioeconomic status populations. Efforts should focus on improving content quality while simplifying language to promote health equity.
    Keywords:  health-care equity; pediatric; quality; readability; skull fracture
    DOI:  https://doi.org/10.7759/cureus.114029
  19. Front Digit Health. 2026 ;8 1858554
       Background: Social media is a key channel for public health information, but its open nature leads to mixed quality of information regarding Ischemic Stroke (IS) treatment, which may mislead patient decisions. This study aimed to systematically compare IS treatment information published by professional and lay sources on two major Chinese social media platforms, Weibo and REDnote, in terms of content, engagement, and quality.
    Methods: This study employed computational social science and Natural Language Processing (NLP) techniques to analyze posts from four sources (Weibo-Pro, Weibo-Lay, REDnote-Pro, REDnote-Lay). We used topic modeling and co-occurrence networks to analyze content features and developed an automated scoring system based on a Large Language Model (LLM) to quantitatively evaluate information quality on two dimensions: "linguistic features" and "evidence-based medical content".
    Results: Lay-sourced content exceeded professional-sourced content in both volume and user engagement. Content and quality patterns differed across the two platforms: on Weibo, the evidence-based quality of professional content was significantly higher than that of lay content (p < 0.001), whereas no significant professional-lay difference was detected on REDnote. We refer to this observed convergence as a "quality paradox," while recognizing that the cross-sectional design does not establish a platform effect. Content themes also differed: Weibo-Lay content centered on Traditional Chinese Medicine, whereas REDnote-Lay content focused on rehabilitation experiences and family support.
    Conclusion: Distinct platform- and source-related patterns were observed across four "discursive communities." On REDnote, professional- and lay-sourced posts showed similar evidence-based quality in this sample. These findings support further investigation of platform-specific communication environments and may inform cautious, context-sensitive approaches for patients, clinicians, and platform managers.
    Keywords:  REDnote; Weibo; health information; ischemic stroke; large language models (LLMs)
    DOI:  https://doi.org/10.3389/fdgth.2026.1858554
  20. Inquiry. 2026 Jan-Dec;63:63 469580261488930
      IntroductionCreator-side analytics can characterize viewing behavior around surgical videos but cannot establish learning or educational quality. We compared advanced procedural and patient-facing videos from one orthopedic YouTube channel.MethodsWe performed a retrospective cross-sectional analysis of cumulative video-level YouTube Studio metrics extracted on February 5, 2026, for videos published from June 2018 through February 4, 2026. One author classified videos using predefined title and metadata criteria. Co-primary outcomes were average view duration and video-attributed net subscriber change per 1,000 engaged views. Mann-Whitney U tests and rank-biserial correlations were used.ResultsAmong 262 included videos (175 advanced procedural; 87 patient-facing), advanced procedural videos were longer (median, 357 vs 60 seconds), had longer average view duration (107 vs 32 seconds), and had higher net subscriber change per 1,000 engaged views (8.04 vs 2.57; both p< .001). Patient-facing videos had higher average percentage viewed (61.16% vs 30.30%; p< .001). Post hoc sensitivity analyses preserved the direction and statistical significance of the comparative findings.ConclusionWithin this single channel, content type corresponded to distinct and consistent within-channel engagement patterns: advanced procedural videos sustained longer absolute viewing and greater net subscriber change per 1,000 engaged views, whereas patient-facing videos were watched more completely. Because absolute and proportional engagement diverged, these dimensions should be interpreted separately. These creator-side metrics show how different orthopedic video formats are watched and how video-attributed net subscriber change varies relative to engagement across formats, providing a single-channel behavioral reference rather than a measure of educational effectiveness or learning.
    Keywords:  YouTube; education; internet; learning analytics; medical; orthopedics; social media; video recording
    DOI:  https://doi.org/10.1177/00469580261488930
  21. Comput Inform Nurs. 2026 Sep 10.
      This study aimed to assess the content and quality of existing YouTube videos on the urethral catheterization procedure. The study used a descriptive research design. Forty-one (41) videos were examined on YouTube. Urethral Catheterization Skill Checklist, DISCERN, Global Quality Standard (GQS), and JAMA were used to collect the study data. Among the videos included in the study, 51.2% were for female patients and 48.8% were for male patients. Content analysis revealed that critical procedural steps were frequently omitted in a substantial proportion of the videos: hand hygiene was not addressed in 22% of the videos, preprocedure information was not provided in 10.1%, latex allergy screening was absent in 82.9%, patient privacy was not ensured in 12.2%, and adherence to aseptic principles was not demonstrated in 7.3%. No statistically significant relationship was found between DISCERN, GQS, and JAMA scores and popularity variables such as video duration, year of upload to YouTube, views, and likes ( P >.05). The urethral catheterization videos evaluated in this study were generally of moderate quality; however, they exhibited notable shortcomings with respect to scientific reliability and adherence to standard procedural steps. Accordingly, online instructional videos should be developed in a more reliable, evidence-based, and professionally standardized manner.
    Keywords:  YouTube; social media; urethral catheterization; urethral catheterization insertion; urinary catheterization
    DOI:  https://doi.org/10.1097/CIN.0000000000001614
  22. Cureus. 2026 Aug;18(8): e114066
       INTRODUCTION: Peyronie's disease (PD) is a penile connective tissue disorder characterized by fibrous plaque formation within the tunica albuginea, resulting in penile curvature, pain, deformity, and impaired sexual function. Due to embarrassment and limited awareness, many patients seek health information online, with YouTube becoming a widely used source of medical education. However, the quality and reliability of PD-related YouTube content remain unclear. This study aimed to evaluate the quality, reliability, and educational value of YouTube videos related to PD and assess their suitability as a patient information resource.
    METHODS: A YouTube search was conducted on May 20, 2026, using five predefined search terms related to PD. The first 20 videos from each search were screened. Videos were included if they were English-language, patient-oriented educational videos lasting between 1 and 20 minutes. Video characteristics were recorded, and content quality was assessed using the DISCERN instrument, Global Quality Scale (GQS), and Journal of the American Medical Association (JAMA) Benchmark Criteria. Statistical analysis included one-sample t-tests and Spearman's rank correlation analysis.
    RESULTS: Of 100 screened videos, 12 met the inclusion criteria. The mean video duration was 4:03 ± 3:08 minutes, with a mean of 212,275 ± 440,000 views. The mean total DISCERN score was 51.42 ± 9.59, indicating lower-range good-quality information, with no significant difference from the score of 51, which is the good-quality threshold (p = 0.883). Individual DISCERN domains, compared to the midpoint score of 3, demonstrated statistically significantly higher scores for clarity of aims (3.92 ± 1.08, p = 0.0137), achievement of aims (3.67 ± 0.98, p = 0.0388), and relevance to patients (3.83 ± 0.94, p = 0.0105). The mean GQS score of 2.92 ± 1.38 was consistent with moderate educational quality and did not differ significantly from the reference score of 3 for moderate-quality educational content (p = 0.838). The mean JAMA Benchmark Criteria score was 1.50 ± 1.23, indicating low reliability. No significant correlations were identified between DISCERN, GQS, and JAMA scores.
    CONCLUSION: YouTube videos regarding PD provide generally acceptable but inconsistent educational value and reliability. While many videos effectively address patient concerns, important deficiencies remain in treatment-related information, transparency, and evidence-based content. Patients should be encouraged to critically evaluate online information, and clinicians should guide individuals toward reliable educational resources. Further development of transparent, comprehensive, and patient-centered digital resources is required to improve online education for men with PD.
    Keywords:  erectile dysfunction; peyronie’s; youtube study; youtube videos; youtube© video analysis
    DOI:  https://doi.org/10.7759/cureus.114066
  23. Medicine (Baltimore). 2026 Sep 04. 105(36): e50566
      Short-video platforms such as TikTok and Bilibili have become widely used sources of health information. This study assessed the informational content, quality, and reliability of videos related to cervical spondylotic myelopathy (CSM) on these platforms. A cross-sectional search was conducted on October 1, 2025, using the Chinese-language equivalent of "cervical spondylotic myelopathy" as the sole search term. For each eligible video, data on duration, user interaction metrics, uploader identity, and content themes were collected. Video quality and reliability were evaluated using the Global Quality Scale (GQS) and the modified DISCERN (mDISCERN) instrument. Group differences were analyzed using the Mann-Whitney U and Kruskal-Wallis tests, and correlations were examined using Spearman analysis. In total, 274 videos met the inclusion criteria. Most videos addressed diagnosis (76.28%) and treatment (74.82%), while prognostic information was scarce (16.06%). The median GQS score was 3.00 (interquartile range [IQR]: 2.00-3.00), and the median mDISCERN score was 4.00 (IQR: 3.00-4.00). TikTok videos achieved significantly higher GQS and mDISCERN scores than those on Bilibili (P < .001). Videos produced by specialized healthcare professionals (SHPs) obtained the best evaluation scores (P < .05). Engagement indicators showed no significant correlation with GQS or mDISCERN scores (P > .05). The overall educational value of CSM-related videos was suboptimal. Content created by SHPs demonstrated superior quality and reliability. Video engagement metrics were not associated with video quality or reliability. Greater professional involvement and preferential presentation of evidence-based content warrant further investigation as potential approaches to improving the dissemination of high-quality CSM information on these platforms.
    Keywords:  Bilibili; Cervical spondylotic myelopathy; TikTok; reliability assessment; video quality
    DOI:  https://doi.org/10.1097/MD.0000000000050566
  24. BMC Public Health. 2026 08 28. pii: 2537. [Epub ahead of print]26(1):
       BACKGROUND: Osteoporosis (OP) affects approximately 145.86 million people in China, imposing a substantial disease burden and broader societal costs. Effective management requires evidence-based patient education. Short-video platforms have become increasingly important channels for health communication, yet few studies have provided systematic comparisons of OP content across Chinese platforms using concurrent measures of overall quality, reliability, and transparency.
    METHODS: This cross-sectional study analyzed 300 OP-related short videos (100 each from TikTok, RedNote, and Bilibili). Two trained raters independently assessed all videos using the Global Quality Scale (GQS), modified DISCERN (mDISCERN), and Journal of the American Medical Association (JAMA) benchmarks; a third reviewer adjudicated disagreements. The primary outcomes evaluated in this study were GQS, mDISCERN, and JAMA scores. Scores were summarized primarily as median (interquartile range [IQR]) and compared using Kruskal-Wallis test, Dunn-Bonferroni pairwise comparisons, and Holm adjustment across the three omnibus tests. Spearman correlations were used to explore associations between quality scores and engagement metrics.
    RESULTS: TikTok videos had the highest descriptive engagement metrics, whereas Bilibili videos had the longest duration. Platform differences were observed for GQS, H(2) = 39.878, P = 2.19 × 10⁻⁹, Holm-adjusted P = 6.57 × 10⁻⁹, η2H = 0.128; mDISCERN, H(2) = 20.453, P = 3.62 × 10⁻5, Holm-adjusted P = 7.24× 10⁻5, η 2H = 0.062; and JAMA, H(2) = 9.651, P = 0.008, Holm-adjusted P = 0.008, η2H = 0.026. Bilibili had the highest mDISCERN and JAMA distributions, whereas RedNote had the lowest GQS distribution. Institutions had higher JAMA scores than all other uploader groups; their GQS advantage was limited to comparison with ordinary users, and mDISCERN did not differ significantly by uploader identity. In the pooled analysis, no statistically significant association was detected between engagement indicators and any quality score (all |ρ|< 0.10).
    CONCLUSION: Overall, the quality, reliability, and transparency of OP information on major short-video platforms were suboptimal, varied markedly across platforms, and showed a striking disconnect between popularity and credibility.
    Keywords:  Cross‑platform analysis; Health information quality; Osteoporosis; Reliability; Short‑video
    DOI:  https://doi.org/10.1186/s12889-026-29150-x
  25. Medicine (Baltimore). 2026 Sep 04. 105(36): e50582
      Short-video platforms increasingly disseminate health information. This study evaluated the content, quality, reliability, and transparency of acute leukemia (AL)-related videos on Douyin and Bilibili. On August 27, 2025, algorithm-ranked search results were sequentially screened, and the first 100 eligible videos from each platform were included, reflecting prominently exposed eligible content rather than the full AL-related video population. Video characteristics, engagement metrics, uploader identity, and content categories were recorded. Quality, reliability, and transparency were assessed using the Global Quality Score (GQS), modified DISCERN, and Journal of the American Medical Association (JAMA) benchmark criteria. Group differences were analyzed using nonparametric tests, and associations between engagement and quality were examined using Spearman correlation and multivariable ordinal logistic regression. Inter-rater reliability was assessed using weighted and unweighted Cohen's kappa. A total of 200 videos were analyzed. Median scores were 3.00 for GQS, 1.50 for modified DISCERN, and 1.50 for Journal of the American Medical Association. Clinical manifestations and treatment were commonly covered, whereas prevention and etiology were underrepresented. Videos uploaded by blood disease-related experts had significantly higher quality and reliability scores. On Douyin, engagement metrics were negatively associated with GQS in correlation analyses, but this association was no longer significant after adjustment for uploader type, video duration, and days since upload (adjusted odds ratio = 0.97, 95% confidence interval: 0.92-1.02, P = .18). Inter-rater reliability was substantial to almost perfect. Algorithm-exposed AL-related videos on Douyin and Bilibili showed suboptimal quality and reliability. Uploader identity, rather than engagement, was more closely associated with information quality. Findings are limited to prominently exposed videos retrieved on one date.
    Keywords:  Bilibili; Douyin; acute leukemia; health communication; information quality; social media
    DOI:  https://doi.org/10.1097/MD.0000000000050582
  26. Obes Surg. 2026 Sep 11.
       INTRODUCTION: Social media platforms have become major sources of health information, including content related to metabolic bariatric surgery (MBS). While these platforms offer wide accessibility and high user engagement, concerns remain regarding the quality, reliability, and completeness of the information presented. Comparative evaluations across platforms integrating both quantitative quality assessment and qualitative insights are limited.
    METHODS: A cross-sectional mixed-methods content analysis was conducted on 45 publicly available videos (15 per platform) from TikTok, YouTube, and Facebook, identified using standardized search terms. Data extraction included video metadata, engagement metrics, uploader characteristics, and content features. Video quality and reliability were assessed using the Video Information and Quality Index (VIQI), Global Quality Score (GQS), and DISCERN instrument. Inter-rater reliability was evaluated using intraclass correlation coefficients. Additionally, thematic analysis of rater comments was performed to explore qualitative content patterns.
    RESULTS: Significant differences were observed across platforms in video duration, content characteristics, and engagement (p < 0.001). TikTok videos were shortest and demonstrated the highest median views and engagement rates, while YouTube videos were longest and Facebook videos showed moderate duration and engagement. Quality assessment revealed significantly lower VIQI, GQS, and DISCERN scores for YouTube compared with TikTok and Facebook (p < 0.001), with no significant difference between TikTok and Facebook. High-quality content (GQS ≥ 4) was absent on YouTube but present in 60-66.7% of TikTok and Facebook videos. Inter-rater reliability was good to excellent across all tools (ICC range: 0.81-0.88). Qualitative analysis demonstrated recurring themes of incomplete clinical information, limited discussion of risks, and a predominance of positively framed narratives, especially in short-form content.
    CONCLUSION: MBS-related content on social media varies significantly across platforms in terms of quality, reliability, and engagement. While TikTok and Facebook demonstrated higher overall quality scores than YouTube, all platforms exhibited important deficiencies in critical patient-centered information. High engagement does not equate to high informational quality. These findings highlight the need for improved content regulation, increased involvement of healthcare professionals, and development of standardized guidelines to enhance the educational value of bariatric surgery information on social media.
    Keywords:  DISCERN; Global quality score; Metabolic bariatric surgery; Social media; Video information and quality index; Video quality assessment
    DOI:  https://doi.org/10.1007/s11695-026-08922-9
  27. J Clin Sleep Med. 2026 09 09. pii: 162. [Epub ahead of print]22(1):
      The reliability and educational quality of TikTok videos addressing insomnia management remain uncertain. A total of 127 TikTok videos were assessed with Global Quality Score (GQS) for educational value and with Viewing and Engagement Index for popularity. Among these, 66 videos by general users, 20 by physicians, 32 by other healthcare professionals, and 9 by brand-affiliated accounts showed overall poor quality: only 20 videos (15.7%) had high educational value, while another 11 (8.7%) had moderate educational value; only 5 videos received the highest score for consistency with guidelines, only 5 videos recommended consulting a physician, and only 2 included scientific references. Approximately 55% of physicians' videos had GQS ≥ 3, compared with 25%, 18%, and 0% of other healthcare professionals, general users, and brand accounts respectively. GQS was not correlated with engagement metrics. Sponsored videos were associated with lower-quality content. In conclusion, TikTok videos related to insomnia management generally do not provide reliable, guidelines-based information, raising concerns regarding the use of TikTok as a source of insomnia management information.
    Keywords:  Global Quality Score; Insomnia; Social media; TikTok; Viewing and Engagement Index
    DOI:  https://doi.org/10.1007/s44470-026-00163-y
  28. Sports Health. 2026 Sep 11. 19417381261479699
       BACKGROUND: "Cycle syncing" is a method of tailoring exercise to different menstrual cycle phases and has gained popularity across social media, with claims of increasing energy and optimizing athletic performance. This study aimed to assess the quality and accuracy of TikTok video content regarding cycle syncing.
    HYPOTHESIS: The quality and accuracy of TikTok content related to cycle syncing will be low.
    STUDY DESIGN: Cross-sectional study.
    LEVEL OF EVIDENCE: Level 4.
    METHODS: TikTok was queried on April 18, 2024, using the terms "cycle syncing workout," "cycle syncing exercise," "ovulation phase workout," "luteal phase workouts," and "follicular phase workout." Engagement metrics, video duration, and poster characteristics were collected. Two independent reviewers assessed video content across 10 categories. In addition, 2 independent experts in women's sports medicine and physiology assessed the educational quality of each video, using the modified DISCERN (scored 0-5) and Global Quality Score (GQS, scored 1-5).
    RESULTS: A total of 160 videos with a collective 18.9 million views were included. Only 14 (9%) videos were posted by healthcare providers. Most videos expressed a positive attitude towards cycle syncing (98%). The most popular topic involved exercise composition (83%). Overall, the videos were of poor educational quality (modified DISCERN, 0.8 ± 0.5; GQS, 1.8 ± 0.6). Furthermore, 80% of videos promoted a product through affiliate marketing; 75% of these were related to cycle syncing.
    CONCLUSION: These findings demonstrated the low educational value of videos focused on cycle syncing on TikTok and an overwhelmingly positive perception of this topic. Affiliate marketing was often the priority. Healthcare providers and social media consumers should be aware of the growing use of social media for this trending topic.
    CLINICAL RELEVANCE: It is important for female athletes interested in cycle syncing to understand the potential inaccuracies shared on social media, before putting recommendations into practice. Medical providers with expertise should consider sharing evidence-based information through social media.
    Keywords:  cycle syncing; female athlete; menstrual cycle
    DOI:  https://doi.org/10.1177/19417381261479699
  29. J Thorac Dis. 2026 Aug 31. 18(8): 890
       Background: Public attention to myocarditis has increased markedly since the coronavirus disease 2019 (COVID-19) pandemic and vaccine-related discussions. Because misleading online content may distort symptom recognition, treatment-seeking behavior, and vaccine attitudes, the quality and clinical completeness of myocarditis-related social media information have important public health implications. This study evaluated the characteristics, clinical content coverage, user engagement, and information quality of myocarditis-related videos on TikTok.
    Methods: This was a cross-sectional content analysis of Chinese-language myocarditis-related TikTok videos retrieved using the keyword "myocarditis" in the mainland Chinese internet environment. Videos published before April 16, 2026 were screened for eligibility. The search was conducted using a newly registered account. Video characteristics, uploader type, user engagement metrics, and five clinical content domains, including etiology, symptoms, diagnosis, treatment, and follow-up recommendations, were extracted. Two cardiologists independently assessed video quality and reliability using the Global Quality Score (GQS) and modified DISCERN (mDISCERN). Discrepancies were adjudicated by a senior cardiovascular physician. All statistical analyses were performed using R version 4.3.3, including Mann-Whitney U tests, Spearman correlation analyses, and linear regression analyses.
    Results: A total of 150 videos were included. Health professionals were the most common uploaders (73%), followed by news organizations (19%) and general users (8%). Clinical content coverage was uneven: symptoms were discussed most frequently (64%), followed by etiology (35%) and diagnosis (30%), whereas follow-up recommendations (19%) and treatment (17%) were less commonly addressed. Videos posted by health professionals had significantly higher GQS and mDISCERN scores than those from non-professional sources, but non-professional videos generated greater audience engagement. Longer video duration was associated with higher GQS scores [β=0.004; 95% confidence interval (CI): 0.003-0.006; P<0.001] and mDISCERN scores (β=0.002; 95% CI: 0.001-0.004; P<0.001), whereas engagement metrics were not associated with information quality.
    Conclusions: Myocarditis-related videos on TikTok showed a mismatch between content credibility and popularity. Longer video duration was associated with higher information quality, but clinical content coverage remained incomplete. Future digital health communication should include more expert-led, accessible, and clinically actionable content that clearly explains key clinical information. Platform-level support, including clearer verification of medical creators and increased visibility of trustworthy health information during periods of high public concern, may help improve online health literacy and reduce the potential harms of misinformation.
    Keywords:  Global Quality Score (GQS); Myocarditis; TikTok; digital health communication; modified DISCERN (mDISCERN)
    DOI:  https://doi.org/10.21037/jtd-2026-1227
  30. J Nephrol. 2026 Sep 09. pii: aajaf059. [Epub ahead of print]
       BACKGROUND: Simple renal cysts are common and usually asymptomatic. Nowadays, short videos on social media have become a major means for obtaining health information. However, there is still insufficient information to evaluate their quality.
    METHODS: In this study, 376 short videos on this topic (ie. simple renal cysts) were selected for a cross-sectional analysis using internationally recognized evaluation tools. The assessment covered content quality, reliability, understandability, and trustworthiness, combined with an analysis of user engagement metrics. The present study was conducted based on data derived from publicly accessible videos on TikTok, Bilibili, Rednote, and Kwai.
    RESULTS: Videos created by western-medicine specialists showed significantly higher scores for quality and trustworthiness. In contrast, videos made by traditional Chinese medicine practitioners attracted larger audiences but received lower quality ratings. In addition, content with mention of the Bosniak classification tended to be of higher quality, yet with less user interaction.
    CONCLUSIONS: Our study reveals a "quality-engagement paradox", in that more accurate and professional content receives less user interaction, while lower-quality videos gain more engagement, thus presenting a challenge for patients to acquire reliable information. Therefore, formulation of platform-specific strategies to improve the visibility and accessibility of trustworthy health content is needed.
    Keywords:  patient education; public health; simple renal cysts; social media
    DOI:  https://doi.org/10.1093/joneph/aajaf059
  31. J Thorac Dis. 2026 Aug 31. 18(8): 911
       Background: Pulmonary tuberculosis (PTB) remains a major public health concern worldwide despite being preventable and curable. China remains one of the 30 high tuberculosis (TB) burden countries, and short-video platforms such as Douyin (the Chinese mainland version of TikTok) have become important sources of health information, yet the quality of PTB-related content remains unclear. This study aimed to evaluate the content completeness, reliability, quality, and engagement of PTB-related videos on Douyin.
    Methods: We conducted a cross-sectional content analysis on October 23, 2025. Using the hashtag "#pulmonary tuberculosis" under the default "comprehensive ranking", we initially retrieved 220 videos. Videos ranked 201-220 were reserved as a priori for pilot reviewer calibration, and the remaining top 200 videos underwent formal screening, of which 191 met the inclusion criteria. We extracted video characteristics, uploader type, and engagement metrics. A PTB-specific checklist covering six dimensions (definition, symptoms, risk factors, evaluation, management, outcomes) was scored, and reliability/quality were assessed using Journal of the American Medical Association (JAMA) benchmarks, Global Quality Score (GQS), modified DISCERN instrument (mDISCERN), and Medical Quality Video Evaluation Tool (MQ-VET). Spearman correlation and multivariable linear regression were performed.
    Results: Of 191 videos, 150 (78.5%) were posted by health professionals. Most were uploaded in 2024-2025 (86.4%); median duration was 74 s [interquartile range (IQR) 48-110]. Content completeness was limited (median 3.0, IQR 2.0-4.5), with risk factors and outcomes rarely addressed; social-context coverage was also limited. Health professionals and science communicators achieved higher content scores than general users and news agencies (P<0.001). Overall quality was modest (median MQ-VET 46, IQR 42-48), lowest among general users and highest among science communicators (P<0.001). Engagement indicators were strongly intercorrelated but weakly or negatively associated with quality, and high-quality videos were not concentrated in top-ranking positions. In regression models, higher comment counts were associated with lower quality, whereas collections were associated with higher quality.
    Conclusions: PTB-related videos on Douyin showed substantial heterogeneity in content completeness and overall quality. Higher-quality videos were not consistently prioritized by platform ranking or user engagement, highlighting the need to promote trustworthy and comprehensive PTB information.
    Keywords:  Douyin; Tuberculosis (TB); cross-sectional studies; health communication; social media
    DOI:  https://doi.org/10.21037/jtd-2026-1289
  32. Soc Sci Med. 2026 Sep 08. pii: S0277-9536(26)00888-9. [Epub ahead of print]408 119810
      With recent changes to the health information landscape, understanding how sources of health information are associated with preventive health behaviors (PHBs) remains an important area for research. In particular, understanding where structurally marginalized populations like Hispanic immigrants receive trusted health information- and how this information is associated with PHBs-is critical, given their more limited access to formal information sources. Further, how gender patterns the association between trust in information sources and PHBs is often overlooked in research on structurally marginalized groups. In this study, we examine how Hispanic immigrants' trusted sources of information are associated with PHBs protecting against COVID-19, focusing on how this association varies by gender. We leverage data from the VidaSana Study, a survey of Hispanic immigrants in Indiana, and its extension, the Hispanic COVID-19 Rapid Response Study. We first identify characteristics associated with trusting sources on COVID-19 before predicting the number of PHBs as a function of trusted information sources and gender. Results indicate that participants who trust the radio as a source of information about COVID-19 are more likely to engage in PHBs. However, interaction models show that men who do not trust radio or TV about COVID-19 are driving this association, as they are less likely to engage in PHBs. We suggest that traditional health communications sources focus on fostering trust among Hispanic men as a way to disseminate accurate and timely information about health concerns.
    Keywords:  COVID-19; Gender; Health behaviors; Health information; Hispanic immigrants; Trust
    DOI:  https://doi.org/10.1016/j.socscimed.2026.119810
  33. J Multidiscip Healthc. 2026 ;19 633607
       Objective: Autoimmune rheumatic diseases (ARDs) require long-term self-management, making health information-seeking behavior (HISB) essential. Prior studies have largely relied on surveys that identified isolated correlates of HISB without clarifying the underlying mechanisms. This exploratory study used the Stimulus-Organism-Response (S-O-R) framework to qualitatively investigate HISB among patients with ARDs, providing theory-informed process-oriented insights into information-seeking behavior in this population.
    Methods: We conducted a descriptive qualitative interview study. Seventeen patients with ARDs were recruited using purposive sampling from a tertiary hospital in Wuhan, China. Data were collected through face-to-face, semi-structured, in-depth interviews between September and November 2025 and analyzed using directed content analysis guided by the S-O-R theory.
    Results: Three main themes emerged: (1) Stimuli included disease-related events, digital platform information, and social support; (2) Organism comprised personal characteristics, disease cognition and information literacy, and emotional responses; and (3) Response yielded active seeking, passive acceptance, selective acceptance, and distress-driven avoidance. Beyond the conventional linear S-O-R pathway, participants described recurrent disease-related stimuli that appeared to repeatedly reactivate information-seeking processes across the illness trajectory.
    Conclusion: This study provides qualitative insights into HISB mechanisms among patients with ARDs. Findings suggest that HISB may involve recurrent stimulus-response patterns, which require validation through longitudinal research.
    Keywords:  autoimmune rheumatic diseases; health information-seeking behavior; patient education; qualitative research; self-management; stimulus-organism-response framework
    DOI:  https://doi.org/10.2147/JMDH.S633607
  34. Int J Med Inform. 2026 Sep 04. pii: S1386-5056(26)00448-X. [Epub ahead of print]222 106708
       BACKGROUND: The rapid growth of digital health technologies is changing how chronic illness patients access & view their health information. During the COVID-19 pandemic there was an increase in mobile health (mHealth) service usage and an increasing number of people exposing themselves to false information via social media. To develop safe and effective health communication strategies, it is important to understand how diabetic patients navigate this new digital world.
    AIM: to examine the health information seeking behaviour of diabetic patients in Saudi Arabia, including the use of digital health tools; interactions with the MOH; the use of social media; and how individuals verify information they find online.
    METHODS: A descriptive cross-sectional survey was conducted among 419 adults with diabetes. Participants were recruited through stratified random sampling to ensure variation in age, education, and digital health familiarity. Data were collected using a structured online questionnaire comprising demographic items, Likert-scale measures of digital health engagement, and perceptions of MOH-provided AI services (e.g., chatbots and self-assessment tools). Analyses were performed using SPSS and R. Structural equation modelling (SEM) examined relationships between latent constructs, while UpSet plots visualised behavioural intersections. Reliability was assessed using Cronbach's alpha.
    RESULTS: Participants demonstrated multidimensional information-seeking patterns, frequently combining physician consultation (n = 245), peer-reviewed sources (n = 213), and government-affiliated platforms (n = 187). MOH digital services, particularly the 937 consultation hotline and Sehhaty virtual clinics, were widely utilised. Trust in AI-based tools remained moderate, with most respondents viewing them as supportive rather than decision-making substitutes. SEM revealed a strong positive association between eHealth engagement and trust in AI tools (β = 0.69, p < 0.001). UpSet plots indicated common co-occurrence of cautious behaviours, including cross-source verification. Cronbach's alpha values ranged from 0.73 to 0.91, confirming acceptable to high internal consistency.
    CONCLUSION: Diabetes patients in Saudi Arabia appear digitally engaged yet critically evaluative, favouring authoritative sources and cautiously adopting AI-enabled services. These findings emphasise the need to strengthen digital health literacy, integrate trusted information within widely used platforms, and enhance transparency and credibility in AI-driven health systems.
    Keywords:  Chatbots in health communication; Diabetes self-management; Digital health literacy; Health information seeking behaviour; Post-pandemic digital engagement; Social media and health misinformation; eHealth trust and verification
    DOI:  https://doi.org/10.1016/j.ijmedinf.2026.106708