bims-librar Biomed News
on Biomedical librarianship
Issue of 2026–07–26
twenty-six papers selected by
Thomas Krichel, Open Library Society



  1. J Chiropr Educ. 2026 Jun 26. pii: eJCE-26-2. [Epub ahead of print]40(1):
       Objective: The purpose of this study was to investigate the services and resources health sciences libraries and librarians can provide to support chiropractic faculty in conducting research and increase institutional research capacity.
    Methods: A one-time, anonymous, online survey was administered to a convenience sample of 120 faculty at a North American chiropractic college in December 2024. The 11-item questionnaire was adapted from prior research in academic health sciences libraries and included both quantitative and qualitative questions about respondents' research interests, experience, needs and library use. Descriptive data analysis was completed.
    Results: The survey was completed by 39 faculty (response rate = 32.5%). Seventy-four percent (n = 29) of respondents had sought research assistance from the library. The areas of greatest interest in library research support were literature searching (56%, n = 20), citation management (53%, n = 19), data analysis (53%, n = 19), research design (39%, n = 14), and journal selection (39%, n = 14). Fifty-six percent (n = 22) of respondents were interested in library research workshops, particularly asynchronous video training. Respondents reported the following research interests: educational research (62%, n = 24), case reports (46%, n = 18), clinical research (44%, n = 17), evidence synthesis (38%, n = 15), historical research (23%, n = 9), and other (5%, n = 2).
    Conclusion: Most respondents had sought research assistance from the library in the past, and more than half were interested in future research training from the library which suggests that the library supports research capacity building within the organization.
    Keywords:  Chiropractic; Education; Faculty; Library Services; Research
    DOI:  https://doi.org/10.7899/JCE-26-2
  2. Med Ref Serv Q. 2026 Jul 21. 1-9
      The ACRL Framework for Information Literacy provides a lens to examine learners' behavior, attitudes, and beliefs. In this follow up to a recent research study, we examine how students transitioning to clinical education understood how they learned to use information as new clinicians. Literature often describes learners with a deficit narrative, where students are expected to follow a prescribed scholarly path to gain knowledge. Participants described a variety of experiences from throughout their lives as valuable which contradicts this narrative. This highlights that librarians' engagement with health sciences students should center their humanity and agency as we guide them toward effective information practices in service of patient care.
    Keywords:  Libraries + Education; health sciences students; information literacy
    DOI:  https://doi.org/10.1080/02763869.2026.2704012
  3. Med Ref Serv Q. 2026 Jul 20. 1-10
      Research in the healthcare environment is required to facilitate improvements in clinical practice and patient care. Health and hospital libraries are in a unique position to support capacity building of health professionals wanting to undertake research. This paper presents a case report from a north Australian health service where Library Services initiated and delivered a Research Support Week. Objectives of the week were to raise the profile of Library Services as a source for support with research, highlight new ways to discover published research and to promote a research community. The program included lightning talks, guest speakers, a research poster competition and a Showcase event. Fifteen sessions were delivered to 150 participants, laying the foundation for future opportunities.
    Keywords:  Collaboration; health professionals; health research; librarians; research support
    DOI:  https://doi.org/10.1080/02763869.2026.2703534
  4. J Robot Surg. 2026 Jul 22. pii: 737. [Epub ahead of print]20(1):
      Robot-assisted radical cystectomy (RARC) is a complex procedure that requires patients to understand surgical indications, urinary diversion, perioperative treatment, complications, recovery, and long-term functional outcomes. Although artificial intelligence (AI) chatbots are increasingly used to obtain medical information, their suitability for RARC patient education remains unclear. We conducted a cross-sectional comparative evaluation of four contemporary AI chatbots: ChatGPT-5, DeepSeek-V4, Claude Sonnet 4.6, and Gemini 3.5 Pro. A set of 20 core patient-education questions on RARC was developed by three senior urologic experts. Chatbot responses were assessed using DISCERN, the Ensuring Quality Information for Patients tool, the Global Quality Scale, and JAMA benchmark criteria. Readability was evaluated using the Automated Readability Index, Coleman-Liau Index, Flesch-Kincaid Grade Level, Flesch Reading Ease, Gunning Fog Index, and SMOG. Reliability scores differed significantly across models for DISCERN, EQIP, and GQS, while JAMA benchmark criteria were summarized descriptively as transparency signals. DeepSeek-V4 achieved the highest mean scores for DISCERN, EQIP, and GQS, while ChatGPT-5 and DeepSeek-V4 showed the strongest JAMA benchmark performance. Gemini 3.5 Pro generally had the lowest reliability and transparency scores. Readability also varied across models. DeepSeek-V4 produced the most readable responses overall, whereas Gemini 3.5 Pro generated the most complex text. However, all models exceeded the recommended sixth-grade reading level, and FRES scores remained below the recommended threshold. Contemporary AI chatbots generated responses with variable presentation quality, transparency, and readability for common RARC patient-education questions. Because factual accuracy was not directly assessed, these tools should not be interpreted as validated sources of clinical guidance and should not replace individualized counseling by urologists.
    Keywords:  Artificial intelligence; Chatbot; Information quality; Patient education; Readability; Robot-assisted radical cystectomy
    DOI:  https://doi.org/10.1007/s11701-026-03698-7
  5. Int Emerg Nurs. 2026 Jul 24. pii: S1755-599X(26)00144-8. [Epub ahead of print]88 101885
       PURPOSE: Childhood accidents are among the leading causes of injury during early childhood. This study aimed to evaluate and compare the accuracy, clarity, and comprehensiveness of pediatric first aid information generated by LLMs.
    METHODS: A cross-sectional comparative evaluation design was employed. Twenty standardized pediatric first aid questions were developed based on international guidelines and expert consensus. Responses generated by ChatGPT, Claude, Gemini, and Copilot were independently evaluated by a pediatric nurse and a physician using a 5-point Likert scale. Inter-rater reliability was assessed using Cohen's kappa coefficient, and differences among models were analyzed using one-way analysis of variance.
    RESULTS: Moderate inter-rater agreement was observed across all evaluation domains. Statistically significant differences were identified among the four LLMs. Claude demonstrated the highest overall performance across all evaluation domains. Gemini demonstrated relatively high accuracy but lower clarity and comprehensiveness scores. Copilot performed well in clarity but showed limited depth of clinical content. ChatGPT received the lowest scores across all assessed domains.
    CONCLUSIONS: The findings reveal considerable variability in the quality of pediatric first aid information generated by LLMs. While certain models may serve as supportive educational tools, none should be considered a substitute for professional medical assessment or emergency care.
    Keywords:  Artificial intelligence; Child; First aid; Natural language processing
    DOI:  https://doi.org/10.1016/j.ienj.2026.101885
  6. BMC Public Health. 2026 Jul 21.
       BACKGROUND: Vaccine hesitancy is a major public health problem that threatens global immunisation programmes. The use of artificial intelligence-based chatbots to access health information is increasing. This study aimed to evaluate the responses of ChatGPT-4o to frequently asked questions about vaccine hesitancy in terms of health information quality, based on expert opinion.
    METHODS: This cross-sectional expert evaluation study included 20 questions on vaccine hesitancy. These questions were submitted to ChatGPT-4o, and the responses were evaluated by nine experts from the fields of public health, paediatrics, and infectious diseases. The evaluation was based on six criteria synthesised from HONcode, DISCERN, JAMA Benchmarks, CRAAP, GQS, and QUEST: scientific accuracy, comprehensiveness, understandability, correction of misinformation, source attribution, and actionability. Each criterion was scored on a scale from 1 to 10. Inter-rater reliability was assessed using the intraclass correlation coefficient (ICC) and, as a prevalence-robust sensitivity analysis, Gwet's AC2 with quadratic weights.
    RESULTS: The highest mean scores were found for understandability (9.4 ± 1.1) and scientific accuracy (8.9 ± 1.3), while the lowest mean score was found for source attribution (4.5 ± 3.4). At the question level, the highest scores were obtained for Question 4, on aluminium in vaccines, and Question 14, on claims that pharmaceutical companies endanger children's health by producing and promoting vaccines. The lowest scores were obtained for Question 5, on live vaccines during pregnancy, and Question 6, on natural immunity. ICC analysis showed moderate agreement only for source attribution (ICC = 0.582; p < 0.001); the low ICC values for the other criteria were attributable to a ceiling effect, as confirmed by Gwet's AC2, which indicated substantially higher agreement for the high-scoring criteria.
    CONCLUSIONS: ChatGPT-4o can generate scientifically accurate and understandable responses to questions about vaccine hesitancy. However, it shows a systematic shortcoming in source attribution. Although the model has potential as a supportive tool in public health communication, its outputs should be reviewed by experts and supported with verifiable sources.
    Keywords:  Artificial intelligence; ChatGPT; Expert evaluation; Health information quality; Large language models; Vaccine hesitancy
    DOI:  https://doi.org/10.1186/s12889-026-28626-0
  7. J Vitreoretin Dis. 2026 Jul 23. 24741264261460669
       Purpose: To evaluate the readability, quality, and misinformation of patient education materials generated by large language models, including ChatGPT-4o (OpenAI), Gemini 1.5 Pro (Google), and Copilot Pro (Microsoft), compared with American Society of Retina Specialists (ASRS) brochures for retinal diseases.
    Methods: A cross-sectional comparative analysis was performed by generating patient education materials on 3 retinal conditions: retinal detachment, diabetic retinopathy, and age-related macular degeneration. Materials were created using a general prompt (prompt A) and a prompt specifying a sixth-grade readability level (prompt B). Readability was evaluated using 6 validated metrics. Quality was assessed through DISCERN and the Patient Education Materials Assessment Tool. Misinformation was graded using a 5-point Likert scale. Assessments were performed independently by 2 masked retina specialists.
    Results: Average readability of Gemini (11.65; P = .005) and Copilot (11.23; P = .003) materials was significantly better than that of ASRS materials (14.17), whereas ChatGPT showed no significant difference (12.85; P = .06). ChatGPT's average readability was significantly lower compared with Gemini (12.85 vs 11.65; P = .01) and Copilot (12.85 vs 11.23; P < .001). Prompt B significantly improved readability across all large language models relative to ASRS but still exceeded the sixth-grade readability level. DISCERN scores were comparable across groups. ASRS materials had an understandability score of 74.37%, which was significantly lower than ChatGPT (94.44%; P = .02) and Gemini 1.5 (95.83%, P = .02) scores. No significant differences were observed for actionability or misinformation. Readability showed no significant correlation with quality or misinformation (P > .05).
    Conclusions: Large language models, when appropriately prompted, can generate retina-related patient education material with superior readability compared with existing ASRS brochures, while maintaining comparable quality and accuracy. Large language models represent a promising approach for addressing literacy barriers, though expert oversight remains essential.
    Keywords:  ChatGPT; Copilot; Gemini; artificial intelligence; consumer health information; large language models; readability; retina
    DOI:  https://doi.org/10.1177/24741264261460669
  8. Clin Exp Dent Res. 2026 Aug;12(4): e70416
       OBJECTIVES: Traumatic dental injuries (TDIs) are frequent in clinical practice and require rapid, guideline-based decisions, yet accessing accurate and reliable information may be challenging. Large language models (LLMs) such as ChatGPT, Gemini, DeepSeek, and Qwen are increasingly used as quick online information tools; however, evidence regarding their accuracy, consistency, and the influence of different user interfaces is limited. This study aimed to evaluate the performance of several LLMs in answering TDI-related questions through both web-based interfaces and mobile phone applications.
    MATERIAL AND METHODS: Twenty questions were prepared according to the 2020 International Association of Dental Traumatology (IADT) guidelines, including 10 open-ended and 10 yes-no items. Four LLMs (ChatGPT-4o, DeepSeek-V3, Gemini 2.0 Flash, Qwen2.5-Max) were queried simultaneously via web and mobile interfaces over five consecutive days, generating 800 responses. Open-ended answers were assessed using the Global Quality Score (GQS) and modified DISCERN (mDISCERN), while yes-no responses were compared with a predetermined answer key. Statistical analyses were performed using IBM SPSS v23.0, with significance set at p < 0.05.
    RESULTS: Qwen2.5-Max demonstrated comparatively higher GQS and mDISCERN scores across both interfaces. Accuracy for yes-no questions ranged from 86% to 91% without significant differences among models. Interface comparisons showed that ChatGPT-4o generated comparatively higher-quality responses on the web, whereas Qwen2.5-Max performed better on mobile. Over the 5-day period, Qwen2.5-Max showed relatively higher temporal consistency, while DeepSeek-V3 exhibited notable day-to-day variation.
    CONCLUSIONS: LLMs may serve as useful supplementary tools for providing guideline-based information on TDIs, especially for straightforward, closed-ended clinical questions. However, their performance varies by model, interface, and question type. Qwen2.5-Max demonstrated comparatively higher performance across several evaluated measures. Despite these results, LLM-generated information should be interpreted cautiously and verified by dental professionals before being used in clinical decision-making.
    Keywords:  artificial intelligence; dentistry; endodontics; information reliability; natural language processing; traumatic dental injuries
    DOI:  https://doi.org/10.1002/cre2.70416
  9. Health Informatics J. 2026 Jul-Sep;32(3):32(3): 14604582261470624
      PurposeTo assess the accuracy, comprehensiveness, and reliability of Large Language Model chatbots in answering frequently asked questions about refractive errors.MethodsForty-four questions about refractive errors were posed to four chatbots, including Copilot, Perplexity, Gemini, and ChatGPT. Responses to each question were independently evaluated by three experts using a three-point accuracy scale. The readability of the chatbots' responses was evaluated using several indices. Similarity was assessed using Sentence-Bidirectional Encoder Representations from Transformers (SBERT). Inter-rater agreement among the graders was evaluated using the Gwet Agreement Coefficient 1 (AC1) statistic.ResultsThe overall agreement among the graders for all chatbot responses was almost perfect (Gwet AC1: 0.87). All chatbots received scores above 85% in every category, and there was no significant difference in chatbot accuracy scores (P= 0.168). The highest mean comprehensiveness score was observed for Perplexity (8.13 ± 0.95, P = 0.028). The similarity scores of the chatbots were very close to each other. All readability scores showed significant differences (P<0.001) between the chatbots.ConclusionAll four chatbots showed comparable accuracy in answering questions about refractive errors. Readability levels of all chatbots exceeded public health recommended thresholds; however, ChatGPT produced relatively more accessible output.
    Keywords:  ChatGPT; chatbots; copilot; gemini; perplexity; readability; refractive errors
    DOI:  https://doi.org/10.1177/14604582261470624
  10. World J Otorhinolaryngol Head Neck Surg. 2026 Jul 06.
       Objectives: Tongue base obstruction is a cause of persistent sleep apnea after adenotonsillectomy in pediatric patients. The internet is the most used tool by patients for gathering healthcare information. This study aims to assess the quality and readability of websites directed to patients regarding tongue base procedures for pediatric obstructive sleep apnea.
    Methods: Queries for the terms "lingual tonsillectomy," "posterior midline glossectomy," "tongue reduction," "tongue base suspension," and "hypoglossal nerve stimulator" were entered into Google on September 11, 2023. Thirty sites were pulled for each, and up to 10 sites per query meeting inclusion criteria were included for analysis. The DISCERN tool was used to evaluate quality. Flesch Reading Ease Score (FRES) and Flesch-Kincaid Grade Level (FKGL) were used to measure readability. Sites were analyzed for mention of pediatrics.
    Results: Thirty-five unique sites were analyzed. The mean DISCERN score was 44.7 (range 28.5-64.0) or "fair." The mean FRES was 46.3 (range 9.6-66.6) and FKGL 11.5 (range 7.4-19.2), indicating college and 12th-grade reading levels. Only 8/35 sites mentioned pediatrics. The average number per query of relevant sites directed towards patients on the first search engine page was 3.6. Despite tongue base procedures being relevant to pediatrics, only one in five sites mentioned children.
    Conclusion: Websites directed at patients regarding tongue base procedures are of "fair" quality and at readability grade levels far above what is recommended. The majority of sites available for tongue base procedures are geared to professionals rather than patients. This study demonstrates a need for high-quality information for patients about these procedures.
    Keywords:  education; glossectomy; obstructive sleep apnea; patient; readability; tonsillectomy
    DOI:  https://doi.org/10.1002/wjo2.70120
  11. J Foot Ankle Res. 2026 Sep;19(3): e70180
       BACKGROUND: Tarsal tunnel syndrome (TTS) is a relatively uncommon compressive neuropathy of the posterior tibial nerve that is often underdiagnosed due to variable presentation and overlap with other conditions. As patients increasingly turn to YouTube for health information, concerns have arisen about the reliability and comprehensibility of such content.
    METHODS: On December 23, 2024, YouTube was queried using the terms "tarsal tunnel syndrome," "tarsal tunnel release," and "tarsal tunnel injection." After excluding duplicates, unrelated, non-English, and short (< 30 seconds) videos, 88 were reviewed and the 50 most-viewed were analyzed. Each was evaluated using three tools: the Journal of the American Medical Association (JAMA) benchmark criteria (0-4) for reliability, a 4-point Likert scale for comprehensibility, and a 19-point Tarsal Tunnel Syndrome-Specific Score (TTS-SS) for educational content. Video source, content type, and video power index (VPI) were also recorded. The Shapiro-Wilk test assessed normality, and Kruskal-Wallis and Wilcoxon rank sum tests compared group with Holm adjustment.
    RESULTS: The median TTS-SS educational content score was 5.2 (IQR 3.0-8.0). Mean JAMA and comprehensibility scores were 2.0 and 2.6, respectively. Only 10% of videos met all four JAMA criteria, and fewer than 20% were rated as highly comprehensible. Disease-specific videos had the highest TTS-SS (mean = 6.8), while non-surgical management scored lowest (mean = 2.8, p = 0.025). Videos by trainers and physical therapists had higher VPI (mean = 338.4) than those by physicians (mean = 0.26, p < 0.001), despite lower educational value. Surgical technique videos were least comprehensible (mean = 1.4) compared with other types (p < 0.05).
    CONCLUSION: Among the most-viewed English-language YouTube videos sampled at a single time point, content related to TTS demonstrated generally low educational completeness, reliability, and comprehensibility. While healthcare professionals produce more accurate content, their videos attract less engagement than those by non-physicians. Improving visibility of evidence-based videos is necessary to ensure patients receive accurate information for this underrecognized condition.
    Keywords:  YouTube; compression neuropathy; educational quality; tarsal tunnel syndrome; video analysis
    DOI:  https://doi.org/10.1002/jfa2.70180
  12. Am J Emerg Med. 2026 Jul 20. pii: S0735-6757(26)00356-6. [Epub ahead of print]109 204-208
       STUDY OBJECTIVE: Erotic asphyxia (sexual choking or breath-control play) is a high-risk practice associated with syncope, hypoxic brain injury, and death, yet increasingly normalized in online sexual content. We sought to characterize how erotic asphyxia is portrayed on YouTube, focusing on video quality, safety messaging, and misinformation.
    METHODS: We performed a retrospective content analysis of YouTube videos identified using 16 search terms related to erotic asphyxia. Eligible videos were English-language, had primary erotic asphyxia content, and were viewed between April and November 2025. Trained coders used a standardized abstraction form to collect data on video characteristics, participant demographics, choking techniques, safety messages, and misinformation. Content quality was assessed using the Global Quality Score (GQS), reliability using a modified DISCERN tool, and interrater reliability using Cohen's kappa.
    RESULTS: A total of 103 videos met inclusion criteria, with a mean duration of 3.9 ± 2.8 min and a cumulative 3.4 million views (mean 34,702 per video). Content depicted consensual partner choking in 67.0%, nonconsensual choking in 11.7%, and autoerotic asphyxia in 21.3%. Most identifiable participants appeared female (63.8%), Caucasian (59.4%), and aged 20-25 years (51.1%). Thirteen distinct choking or suffocation techniques were described. The median GQS was 2 (interquartile range [IQR] 2-3), and the median modified DISCERN reliability score was 1 (IQR 1-2), indicating low quality and very poor reliability. Only 4 videos (3.9%) included trigger warnings, and 74 (71.9%) contained clear misinformation. Interrater reliability was substantial (κ = 0.79).
    CONCLUSIONS: Widely viewed erotic asphyxia content on YouTube is low quality, unreliable, and often misleading, highlighting the need for emergency physicians to address sexual choking directly in risk-reduction counseling and for platform-level harm-reduction measures to prevent serious injury and death.
    Keywords:  Autoerotic asphyxiation; Breath-control play; Content analysis; Erotic asphyxia; Sexual choking; Social media; Strangulation; YouTube
    DOI:  https://doi.org/10.1016/j.ajem.2026.07.033
  13. Sci Rep. 2026 Jul 19.
      Diabetic foot represents one of the most serious complications of diabetes mellitus. Short videos on diabetic foot are increasingly disseminated through TikTok and Bilibili apps, which have emerged as dominant sources of health-related content. However, the credibility and quality of information in these short videos have not been systematically evaluated. Our study aims to assess the quality and reliability of Chinese short videos on diabetic foot shared on TikTok and Bilibili. A cross-sectional study design was adopted, and a total of 244 short videos related to diabetic foot were collected from two platforms: TikTok and Bilibili. On December 15, 2025, information quality and reliability assessment was conducted using three validated evaluation tools: GQS for quality assessment, and the mDISCERN and JAMA benchmarks for reliability assessment. Meanwhile, user interaction indicators and video characteristics were extracted. Nonparametric tests were applied to compare differences across platforms and uploader types, and Spearman correlation was used to examine relationships among video characteristics, engagement metrics, and quality scores. Compared to Bilibili, TikTok demonstrated significantly higher engagement metrics (all P < 0.001). The quality of short videos was suboptimal on both platforms, with median GQS of 2.00 (1.00,3.00), mDISCERN of 2.00 (1.00,2.00), and JAMA score of 2.00 (2.00,2.00). TikTok videos achieved significantly higher GQS and JAMA scores than Bilibili (both P < 0.001). By uploader type, non-specialists achieved significantly higher GQS scores than specialists (median 3.00 vs. 2.00, P < 0.001). Content distribution was imbalanced: clinical manifestations (48.36%) and treatment (44.26%) were prevalent, whereas epidemiology (7.79%) and diagnosis (14.34%) were underrepresented; 51.23% of videos covered a single theme and 48.77% addressed multiple themes. Video duration was positively correlated with GQS (r = 0.59, P < 0.001) and mDISCERN (r = 0.26, P < 0.001) but not with JAMA, while engagement metrics showed only weak or non-significant correlations with all quality scores. Our study shows that the quality of short videos on diabetic foot is poor on TikTok and Bilibili, with significant content imbalances. Videos uploaded by professionals demonstrated better reliability than those from individual users. Engagement metrics were not reliable indicators of information quality. Thus, medical information short videos on these platforms must be carefully evaluated for scientific soundness, and platforms should strengthen content regulation to ensure accurate diabetic foot education.
    Keywords:  Diabetic foot; GQS; Health information quality; JAMA; Short videos; mDISCERN
    DOI:  https://doi.org/10.1038/s41598-026-62665-2
  14. Health Informatics J. 2026 Jul-Sep;32(3):32(3): 14604582261471307
      BackgroundInflammatory Bowel Disease (IBD) is a chronic condition with increasing global prevalence. Short-video platforms like TikTok and Bilibili are popular health information sources, but their IBD-related content quality remains uncertain.ObjectiveTo assess the quality and reliability of IBD-related videos on TikTok and Bilibili.MethodsWe analyzed the top 100 IBD-related videos from each platform (200 total). Quality was evaluated using Global Quality Scale (GQS), modified DISCERN (mDISCERN), and JAMA benchmarks. Engagement metrics and quality scores were correlated.ResultsTikTok videos had higher engagement (P < .001) and better median scores (GQS: 3, mDISCERN: 3, JAMA: 3) than Bilibili (GQS: 3, mDISCERN: 3, JAMA: 3; P < .05). Gastroenterologist-created videos scored highest. GQS correlated positively with engagement (r = 0.16-0.26, P < .05), while mDISCERN correlated only with likes, shares, and saves (r = 0.16-0.18, P < .05). JAMA scores negatively correlated with duration (r = -0.24, P < .001).ConclusionsIBD-related short videos showed moderate quality, with TikTok outperforming Bilibili. Gastroenterologist-produced content was most reliable. Viewers should critically evaluate such health information to avoid misinformation.
    Keywords:  Bilibili; TikTok; health information; inflammatory bowel disease; short videos
    DOI:  https://doi.org/10.1177/14604582261471307
  15. Medicine (Baltimore). 2026 Jul 24. 105(30): e49924
      Lupus nephritis (LN) is an important cause of end-stage renal disease. Recently, more and more patients have been accessing LN-related health information through TikTok and Bilibili, but the quality of the content is uneven. This cross-sectional study aims to systematically evaluate the content quality of LN-related short videos on these 2 major platforms. On December 13, 2025, we searched for "LN" in Chinese on TikTok and Bilibili and obtained the top 100 results from each platform. After screening, 159 videos were analyzed. The Global Quality Score (GQS) and modified DISCERN (mDISCERN) scales were used for multidimensional quality evaluation, and the correlation between video characteristics and interaction metrics was analyzed. The results showed that Bilibili's videos had significantly longer durations (median: 486 seconds vs TikTok's 82.5 seconds) and higher GQS scores (median: 3.00 vs 2.00, P <.05). After adjusting for confounding factors, there was no significant difference in GQS (odds ratio = 1.40, 95% confidence interval: 0.43-4.64) or mDISCERN (odds ratio = 0.78, 95% confidence interval: 0.28-2.16) between TikTok and Bilibili. TikTok videos were more interactive, receiving more likes, shares, and comments. The content on both platforms mainly focused on "treatment" and "symptoms" while rarely covering topics such as "prevention" and "epidemiology." A positive correlation was found between content quality and the professional background of the creator: the median GQS and mDISCERN scores for videos released by professional clinicians were 3.00, whereas those for individual users were only 1.00. In addition, audience interaction was negatively correlated with content quality. In summary, videos on both Bilibili and TikTok demonstrated modest overall quality.
    Keywords:  Bilibili; TikTok; lupus nephritis; short videos; video quality
    DOI:  https://doi.org/10.1097/MD.0000000000049924
  16. BMC Oral Health. 2026 Jul 20.
       BACKGROUND: Social media has become an important source of orthodontic information, yet patient experience content varies in completeness and may shape treatment expectations. This study evaluated information completeness, comment sentiment, and engagement patterns in orthodontic experience content on three major Chinese social media platforms.
    METHODS: A cross-sectional sample of 180 eligible posts was collected from Bilibili, Douyin, and Xiaohongshu using two search terms, "orthodontic experience" and "braces diary." Information completeness was assessed using a modified 6-point Information Completeness Score (ICS). Comment sentiment and title sentiment were quantified using sentiment analysis scores (SAS). Engagement metrics, including likes, favorites, shares, and comments, were recorded. Platform differences and subgroup differences by appliance type, treatment population, video duration, and video age were analyzed.
    RESULTS: Inter-rater reliability for total ICS was excellent (ICC (3,1) = 0.954). The automated SAS classifications showed strong agreement with the human consensus gold standard, with an overall accuracy of 91.6% (κ = 0.848). ICS differed significantly across platforms (H = 40.70, p < 0.001), with Bilibili showing the highest score (3.53 ± 1.46) and Xiaohongshu the lowest (1.78 ± 1.24). Treatment details and treatment duration were the most frequently mentioned dimensions, whereas clinician or institution credentials and cost were least frequently covered. Title sentiment showed a statistically significant but weak positive correlation with comment sentiment in the full sample (rho = 0.247, p < 0.001), with platform-specific variation. Engagement indicators were strongly intercorrelated (rho = 0.69-0.87), but ICS showed no meaningful correlation with engagement. Longer videos, clear aligner content, adult orthodontic content, and older videos had higher ICS, whereas SAS varied only slightly across groups.
    CONCLUSIONS: Orthodontic patient experience content on Chinese social media shows substantial variation in information completeness. High engagement does not necessarily indicate high information quality. Professional orthodontic content should provide balanced information on treatment benefits, risks, costs, duration, and professional sources to support informed patient decision-making.
    Keywords:  Information quality; Orthodontics; Patient testimonials; Sentiment analysis; Social media
    DOI:  https://doi.org/10.1186/s12903-026-09282-7
  17. Prev Med Rep. 2026 Aug;68 103565
       Objective: This study examined the association of cancer-preventive lifestyle behaviors with cancer information-seeking experiences among United States adults without a history of cancer.
    Methods: We analyzed data from the Health Information National Trends Survey (HINTS) 7 from 2412 participants between March 2024 and September 2024. A healthy behavior index was constructed based on lifestyle behaviors and categorized into quintiles (Q1-Q5), where Q5 reflects a higher healthy lifestyle score. The Information Seeking Experience scale was computed from four questions that capture participants' experiences for cancer-related information (i.e., search required a lot of effort, was frustrating, raised concerns about information quality, or was hard to understand) and was categorized into quartiles (Q1-Q4, where Q4 reflects the most positive experiences).
    Results: Compared with the lowest healthy behavior index quintile, the quintile was associated with higher relative risk ratios (RRRs) for Information Seeking Experience Q3 (RRR = 2.3; 95% CI: 1.2,4.5) and Q4 (RRR = 2.7; 95% CI: 1.5,5.0). Maintaining a healthy weight (RRR Information Seeking Experience Q4 = 1.9, 95% CI = 1.2,3.1) and not consuming alcohol (RRR Information Seeking Experience Q3 = 2.4; 95% CI = 1.1,5.3) were also associated with more positive experiences.
    Conclusions: Adults with lower healthy behavior profiles reported more negative cancer information-seeking experiences.
    Keywords:  Cancer prevention; Healthy behaviors; Information seeking experiences
    DOI:  https://doi.org/10.1016/j.pmedr.2026.103565
  18. Otol Neurotol. 2026 Jul 20.
       OBJECTIVE: To compare the quality of tinnitus-related information generated by multiple generative artificial intelligence (GenAI) systems and web search using expert evaluation.
    STUDY DESIGN: Cross-sectional comparative study.
    SETTING: Digital platforms evaluated in their native public interfaces.
    PATIENTS: Not applicable. Thirty commonly searched tinnitus-related questions derived from United States Google Trends data (2020-2025).
    INTERVENTIONS: Questions were submitted to 6 GenAI systems (OpenEvidence, Claude, DeepSeek, GPT-5, Gemini, GPT-4) and Google Search (first organic result). Responses were independently rated by 6 experts using the QAMAI framework.
    MAIN OUTCOME MEASURES: Mean expert-rated quality scores across 5 domains (accuracy, clarity, relevance, completeness, and usefulness).
    RESULTS: Overall quality differed significantly across systems (Friedman P<0.001; Kendall W=0.34). OpenEvidence achieved the highest mean score (4.45±0.72; 95% CI: 4.40-4.49), followed by Claude (4.00±1.02), DeepSeek (3.92±1.13), GPT-4 (3.89±0.84), Gemini (3.62±0.98), and GPT-5 (3.30±1.11). Google Search scored lowest (2.27±1.12; 95% CI: 2.20-2.35). Completeness was the lowest-performing domain across systems (range: 1.70-4.41). Pairwise comparisons showed significant differences between OpenEvidence and all other systems (effect size r=0.49-0.86). Inter-rater reliability was high (ICC=0.82). Readability demonstrated an inverse pattern relative to expert-rated quality. OpenEvidence demonstrated the lowest readability (Flesch-Kincaid Grade Level 17.5; Flesch Reading Ease 2.2), corresponding to a postgraduate reading level, whereas general-purpose LLMs produced more accessible responses at a sixth-seventh grade reading level.
    CONCLUSIONS: The quality of tinnitus information varies substantially across digital platforms. While GenAI systems generally outperform web search, deficiencies in completeness persist. Readability analysis revealed an inverse relationship between expert-rated quality and response accessibility, suggesting that clinician and patient assessments of informational value may not always align. These findings highlight the need for continued evaluation and clinician oversight to ensure safe, comprehensive, and accessible patient-facing information.
    Keywords:  Generative artificial intelligence; Health information quality; Large language models; Patient education; Tinnitus
    DOI:  https://doi.org/10.1097/MAO.0000000000005015
  19. Front Public Health. 2026 ;14 1837642
       Background: Hypertension requires sustained self-management beyond routine clinical encounters, yet evidence on how patients engage with digital health information and what shapes their trust in AI-based physician chatbots remains limited, particularly in Saudi Arabia.
    Aim: This study examined digital health information-seeking behaviors, information verification practices, and perceptions of AI-based physician chatbots among adults diagnosed with hypertension in Saudi Arabia and tested a theoretically grounded model linking six constructs to chatbot safety and trust.
    Methods: A cross-sectional online survey was conducted among 322 adults with hypertension recruited through WhatsApp, Telegram, and Facebook patient groups. A structured questionnaire measured six constructs Information Validation Sources (IVS), Digital Self-Diagnosis Tools (DSDT), Online Information-Seeking Behavior (OISB), Familiarity with MOH Digital Services (FMOHDS), Perceived Reliability of Self-Assessment (PRISA), and Chatbot Safety and Trust (CST) adapted from past studies. Internal consistency, descriptive statistics, confirmatory factor analysis, bootstrapping and Kruskal-Wallis group comparisons were performed.
    Results: All constructs scored above the scale midpoint (M = 3.32-3.78). WhatsApp was the dominant platform for health information seeking (43.5%), verification (51.9%), and communication with health services (41.9%). Measurement model and CFA results shows that scales are reliable and valid, factor loadings, alpha values, AVE and composite reliability and VIF values met threshold. Discriminant validity using Fornell-Larcker is also reported. Bootstrapping results have revealed that OISB has positive significant effect on CST, while PRISA has negative but significant effect on CST, remaining IVS, DSDT and FMOHDS effects on CST were not significant. Moreover, direct effects of predictors on PRISA were also investigated, findings have revealed that FMOHDS is most dominant predictor of PRISA, while OISB negatively but significantly predicted PRISA, however, IVS, DSDT effects on CST were not significant. Regarding indirect mediating effects only PRISA significantly mediates between FMOHDS and CST, while PRISA mediating effects were insignificant between IVS, DSDT OISB and CST.
    Conclusion: Trust in AI-based physician chatbots among hypertensive patients is positively shaped by institutional familiarity with MOH digital services and perceived reliability of self-assessment tools and negatively influenced by active independent digital health engagement. The mediating effect of PRISA on FMOHDS and CST suggests that building institutional trust in existing digital infrastructure is a prerequisite for effective AI chatbot adoption. These findings have direct implications for the design of AI-enabled hypertension management strategies within Saudi Arabia's Vision 2030 digital health agenda.
    Keywords:  AI-based physician chatbots; digital health; health information seeking; healthcare; hypertension; information verification
    DOI:  https://doi.org/10.3389/fpubh.2026.1837642
  20. Digit Health. 2026 Jan-Dec;12:12 20552076261456579
       Objective: This research aims to characterize the English language tobacco and nicotine information environment on YouTube by examining the content to which those who actively seek tobacco and nicotine content are exposed.
    Methods: We used Google Trends, YouTube's Application Programming Interface (API), and a custom python script to sample the most common videos populated by the most common search terms, and the most common videos populated by YouTube's recommendation feed. After categorizing search terms as "Health Information Seeking" (HIS), "Pro-tobacco/Nicotine" (PtN), or Ambiguous, a content analysis of N = 606 relevant videos identified pro-versus anti-tobacco videos by tobacco manufacturer, broadcast media, health institution, and organic sources. Analyses describe how often each kind of search leads to each kind of source and to pro-tobacco videos in both the primary search results and the recommended videos.
    Results: In sum 394(65%) videos were pro-tobacco including 83% of PtN and 60% of ambiguous searches. However, just 5 of 118 videos returned by HIS searches or recommendations were pro-tobacco. Tobacco manufacturer videos were primarily found in PtN searches, though two were recommended in ambiguous searches. The most common primary search results for HIS were health institutions comprising 39% of combined primary (45%) and recommended (36%) videos.
    Conclusion: Although possible, tobacco and nicotine related HIS on YouTube is unlikely to drive exposure to pro-tobacco content and relatively likely to lead to reputable sources such as health institutions. Conversely, ambiguous or non-specific searches (e.g. vape, nic pouch) both lead to a pro-tobacco information environment. Potential individual and platform-level interventions are discussed.
    Keywords:  addiction; digital health; health communications; quantitative; social media
    DOI:  https://doi.org/10.1177/20552076261456579
  21. J Med Internet Res. 2026 Jul 20. 28 e90567
       Background: Patients with prostate cancer undergoing androgen deprivation therapy (ADT) must manage complex treatment side effects over extended periods outside the hospital, making online health information-seeking a key approach to self-management. However, individual differences in motivation, digital literacy, and psychosocial context significantly influence how patients seek and use online health information. The patient persona approach, which synthesizes individuals with similar behavioral patterns and needs into representative profiles, offers a practical method for capturing this heterogeneity and informing the design of tailored information support.
    Objective: This study aimed to explore the differences in online health information-seeking behavior among patients with prostate cancer undergoing ADT and construct patient personas to characterize distinct patterns of online health information-seeking experiences and related information support needs.
    Methods: A qualitative descriptive study was conducted in a tertiary hospital in Zhejiang Province, China, from July to October 2025. Purposive sampling was used to recruit patients receiving ADT. Semistructured interviews were conducted, and data were analyzed using inductive content analysis to derive codes, subcategories, and 4 core categories. Participants were systematically compared across these categories and grouped into personas based on recurring patterns. The personas were then presented with illustrative portraits.
    Results: A total of 20 participants were included. Four core categories were identified for persona construction: motivation, information access preferences, barriers, and needs. The personas were categorized as follows: patients demonstrating proactive information management and high health literacy, patients exhibiting avoidant information-seeking driven by latent anxiety, patients engaging in family-doctor-mediated information-seeking, and patients displaying dependent and passive information-seeking behavior.
    Conclusions: Online health information-seeking behavior among patients with prostate cancer receiving ADT is highly heterogeneous and shaped by individual capability, emotional responses, and sociocultural factors. The identified personas may provide a preliminary and practice-oriented framework for developing more differentiated, culturally sensitive, and patient-centered online health information support strategies for patients receiving long-term ADT.
    Keywords:  digital health; information-seeking behavior; patient persona; prostate cancer; qualitative research
    DOI:  https://doi.org/10.2196/90567
  22. Women Health. 2026 Jul 21. 1-13
      This study examined the impact of women's online health information seeking and verification behaviors and eHealth literacy on ovarian cancer awareness. This descriptive, cross-sectional study included 440 women aged 20-65 years attending a public Community Health Center in Türkiye between August 2025 and January 2026. Data were collected using the Online Health Information Seeking and Verification Behavior Scale (OHISVBS), eHealth Literacy Scale (eHEALS), and Ovarian Cancer Awareness Scale (OCAS). Pearson correlation and hierarchical regression analyses were performed. eHealth literacy was positively associated with symptom recognition (r = 0.305, p < .001) and risk factor recognition (r = 0.303, p < .001). All OHISVBS subdimensions were significantly correlated with both awareness subdimensions (r = 0.230-0.341; p < .001). Hierarchical regression showed that regular gynecological examination and eHealth literacy were independently associated with higher symptom recognition (β = .271 and .168, respectively) and risk factor recognition (β = .311 and .267, respectively). The final models explained 37.1 percent and 35.4 percent of the variance. Strengthening women's digital health literacy and promoting online health information seeking and verification behaviors may improve ovarian cancer awareness and timely healthcare seeking.
    Keywords:  EHealth literacy; information verification; online health information seeking; ovarian cancer awareness; women’s health
    DOI:  https://doi.org/10.1080/03630242.2026.2704068
  23. Health Care Sci. 2026 May 19.
       Background: Searching for health information online helps cancer patients better understand their health conditions and the way to cope with the adverse effects of treatment. Although eHealth literacy, self-efficacy and information utility are recognized as key factors that are associated with online health information seeking behavior (OHISB), the underlying mechanisms linking these factors in cancer patients remain understudied. This study aims to examine the mediating roles of self-efficacy and information utility in the relationship between eHealth literacy and OHISB.
    Methods: In this cross-sectional study, 502 cancer patients were selected from a tertiary hospital in Shanghai from June to December 2024. Demographic and clinical characteristics were obtained from electronic health records. eHealth literacy, self-efficacy, information utility and OHISB were assessed using self-report questionnaires. Mediation analysis was performed using structural equation modeling, with indirect effects evaluated through bootstrapping.
    Results: The prevalence of OHISB among cancer patients was 61.4%. There was a significant correlation between eHealth literacy, self-efficacy, information utility and OHISB (r = 0.357, 0.499, 0.469, p < 0.01). eHealth literacy demonstrated both direct (β = 0.190, standard error [SE] = 0.050, 95% confidence interval [CI] [0.099, 0.295]) and indirect (β = 0.311, SE = 0.045, 95% CI [0.227, 0.404]) impacts on OHISB, mediated by self-efficacy (β = 0.223, SE = 0.038, 95% CI [0.155, 0.304]) and information utility (β = 0.056, SE = 0.025, 95% CI [0.013, 0.110]). The findings indicated that self-efficacy and information utility served as a sequential mediating factor in the association between eHealth literacy and OHISB, contributing 6.6% to the overall effect (β = 0.033, SE = 0.012, 95% CI [0.010, 0.057]).
    Conclusions: eHealth literacy enhances the OHISB of cancer patients through both direct effects and a chain mediation pathway involving self-efficacy and information utility. Targeted interventions should focus on improving eHealth literacy, strengthening self-efficacy, and developing reliable information platforms to optimize patient-centered health information services.
    Keywords:  chain mediating effect; cross‐sectional study; health information seeking behavior; health literacy; information utility; self‐efficacy
    DOI:  https://doi.org/10.1002/hcs2.70079