bims-arines Biomed News
on AI in evidence synthesis
Issue of 2026–09–06
nineteen papers selected by
Farhad Shokraneh, Systematic Review Consultants LTD



  1. JMIR Form Res. 2026 Aug 31. 10 e86647
       Background: Systematic literature reviews (SLRs) are essential for evidence synthesis in health research but remain labor-intensive, especially at the screening stage. Manual review of titles and abstracts requires substantial human effort, while existing automation tools still have limited adoption in health technology assessment. The EQ-5D questionnaire, a widely used patient-reported outcome measure for health-related quality of life, provides data that frequently underpin reimbursement and policy decisions.
    Objective: This pilot study evaluated whether recent large language models (LLMs) can support the identification of publications reporting EQ-5D data in PubMed records, using only publicly available metadata (title, abstract, and keywords).
    Methods: A total of 200 publications retrieved through the EuroQol PubMed filter were manually labeled by experts as reporting or not reporting EQ-5D data. The dataset was split into stratified training, validation, and test subsets. Several machine learning approaches were compared, including a Naïve Bayes baseline using bag-of-words features, a decision-tree model based on full-text keyword occurrence, and transformer-based LLMs (Bidirectional Encoder Representations from Transformers [BERT], Biomedical BERT [BioBERT], Scientific BERT [SciBERT], and Biomedical Language Understanding Evaluation BERT [BlueBERT]). Both classifier-only and fine-tuned configurations were tested across multiple learning rates. Model performance was assessed using accuracy, precision, recall, and F1-score.
    Results: Baseline approaches achieved near-random test performance (accuracy around 0.53). Classifier-only LLMs modestly improved results (accuracy up to 0.64 with SciBERT). Fine-tuned models substantially outperformed these baselines, with BERT and BioBERT achieving the best performance (accuracy=0.70; F1-score=0.68). In screening-oriented evaluation, this configuration achieved 90.0% sensitivity, 40.0% specificity, and 6 false negatives on the held-out test set. The models reproduced human screening tendencies despite the small dataset size, demonstrating the technical feasibility of LLM-assisted article selection.
    Conclusions: This study provides the first demonstration of LLM-assisted identification of EQ-5D data in biomedical literature. The findings support technical feasibility but do not establish a reliable stand-alone automated screening tool. Although limited by dataset size, the proposed workflow is reproducible and adaptable to other patient-reported outcome measures. Because validation was based on a single small train-validation-test split, the results should be interpreted as preliminary; future work will scale data collection, include statistical testing, and explore semisupervised learning to further reduce manual screening workload.
    Keywords:  BERT; Bidirectional Encoder Representations from Transformers; BioBERT; Biomedical Bidirectional Encoder Representations from Transformers; EQ-5D; PubMed; health informatics; large language model; quality of life; systematic literature review
    DOI:  https://doi.org/10.2196/86647
  2. Artif Intell Rev. 2026 ;59(10): 216
      Full-text screening is the major bottleneck of systematic reviews (SRs), particularly in domains such as population health modelling of noncommunicable diseases (NCDs), where decisive eligibility information is scattered across long and heterogeneous full texts. In this methodological study, we introduce a scalable and auditable pipeline that reframes inclusion and exclusion as a fuzzy decision process. We evaluate the approach within the Population Health Modelling Consensus Reporting Network (POPCORN) and benchmark it against statistical and crisp baselines. Articles are parsed into overlapping chunks and embedded with a domain-adapted model; for each criterion (Population, Intervention, Outcome, Study Approach), we compute contrastive similarity (inclusion-exclusion cosine) and a vagueness margin, which a Mamdani fuzzy controller maps into graded inclusion degrees with dynamic thresholds in a multi-label classification setting. A large language model (LLM) judge adjudicates highlighted spans with tertiary labels, confidence scores, and criterion-referenced rationales; when evidence is insufficient, fuzzy membership is attenuated rather than hard-excluded. In a pilot on an all-positive gold set ([Formula: see text] full texts; [Formula: see text] chunks), the fuzzy system achieved document-level recall of 81.25% (Population), 87.50% (Intervention), 87.50% (Outcome), and 75.00% (Study Approach) with 95% Wilson confidence intervals, exceeding statistical baselines (recall 56.25-75.00%) and crisp baselines (recall 43.75-81.25%). Strict "All Criteria" inclusion was reached for 50.00% of articles, compared to 25.00% and 12.50% under the baselines. Cross-model agreement on justifications was 98.27%, and human-machine agreement was 96.07%. A pilot review showed 91% inter-rater agreement ([Formula: see text]) with screening time reduced from ∼20 minutes to under 1 minute per article at a significantly lower cost. These results suggest that fuzzy logic, combined with contrastive highlighting and explainable LLM adjudication, delivers high recall, stable rationales, and end-to-end traceability. We emphasize that the reported recall figures were computed on an all-positive gold set (no excluded documents), so specificity, ROC/PR curves, and AUC are not defined here; the complete operating-characteristic analysis on a mixed-label corpus will be reported in a forthcoming study.
    Keywords:  Biomedical NLP; Contrastive semantic similarity; Evidence synthesis; Explainable AI; Full-text screening; Fuzzy logic (Mamdani); Large Language Models (LLMs); Multi-label text classification; Systematic review automation
    DOI:  https://doi.org/10.1007/s10462-026-11599-2
  3. Glob Epidemiol. 2026 Dec;12 100282
       Introduction: Artificial intelligence (AI) is being rapidly integrated into systematic review workflows, yet its impact on methodological rigor, transparency, and reporting quality remains poorly understood. This work examines the current use of AI assistance in systematic reviews and identifies gaps in existing appraisal frameworks. We aim to propose a conceptual methodological and illustrative framework that maps AI-assisted processes in the systematic review workflow.
    Methods: We conducted a conceptual methodological analysis informed by a targeted, non-systematic review of recent literature on AI-assisted systematic review workflows, mapped AI use across review stages, and evaluated alignment with existing appraisal and reporting frameworks (AMSTAR-2, PRISMA-2020, PRISMA-S, and ROBIS).
    Results: We identified a misalignment between AI-assisted systematic review workflows and existing methodological standards, which were developed for human-led systematic review workflows. We propose a conceptual framework that maps AI use across the systematic review process and delineates three core domains of methodological evaluation: transparency, reproducibility, and validity. Within this framework, we define key sources of methodological risk, such as prompt dependency, algorithmic reproducibility, and epistemic opacity, and illustrate how these risks may not be fully captured by current appraisal and reporting instruments such as AMSTAR-2, PRISMA, and ROBIS.
    Discussion: AI has the potential to support efficient systematic reviews, but credibility depends on transparent reporting, reproducible processes, and rigorous human verification. In our targeted evidence scan, empirical evaluations primarily addressed isolated AI-assisted tasks rather than complete systematic review workflows. Further methodological work is needed to evaluate whether, when, and under what conditions AI-assisted systematic reviews preserve the standards required for evidence-based decision-making; the proposed framework is intended to guide such work rather than serve as a validated appraisal instrument.
    Keywords:  Artificial intelligence; Methodological quality; Reproducibility; Systematic reviews; Transparency
    DOI:  https://doi.org/10.1016/j.gloepi.2026.100282
  4. Nurse Educ. 2026 Sep 04.
       BACKGROUND: Generative artificial intelligence (AI) may support systematic review learning, but general-purpose chatbots and workflow tools do not explicitly teach methodological reasoning.
    PURPOSE: To develop a stage-structured retrieval-augmented generation learning agent and assess its perceived pedagogical fit and technical performance.
    METHODS: The tool combined planning, query rewriting, iterative retrieval, and context-sufficiency checking across 6 stages: topic selection, review question framing, search strategy, screening and management, critical appraisal and data extraction, and synthesis and reporting. Six nurse educators and 4 undergraduate nursing students completed author-developed questionnaires assessing perceived pedagogical fit and AI performance.
    RESULTS: The tool generated stage-aligned guidance across the review process. Participants reported moderate perceived pedagogical fit (M = 3.90, standard deviation = 0.45) and high perceived AI performance (M = 4.00, standard deviation = 0.25); personalization required improvement.
    CONCLUSIONS: The tool was feasible as an instructional scaffold, but findings reflect perceptions from a small sample. Larger studies should assess objective learning outcomes, methodological accuracy, and responsible AI use.
    Keywords:  digital pedagogy; evidence synthesis; evidence-based practice; large language models; research literacy; systematic review learning agent
    DOI:  https://doi.org/10.1097/NNE.0000000000002330
  5. J Glob Health. 2026 Sep 04. 16 03029
      Traditional secondary meta-analysis workflows are highly labour-intensive, time-consuming, and difficult to update in real time. Currently, there is a lack of comprehensive artificial intelligence frameworks capable of automating the entire meta-analysis workflow, including literature screening, data extraction, and quality assessment. Furthermore, a large-scale structured database for systematically analysing the global landscape of published meta-analyses remains unavailable. In this viewpoint, we aimed to evaluate the feasibility of large language models in automating meta-analysis workflows and develop the Meta-Analysis Screening, Transformation and Evaluation Review Agent (MASTER) agent; establish a large-scale Unified Meta-Analysis Repository (UMAR) and perform an exploratory panoramic analysis of the current evidence ecosystem; and develop an Agent-based Secondary Meta-analysis Platform (ASAP), integrating these capabilities. We subsequently applied the agent to process 311,751 meta-analysis records to establish the UMAR database. Building upon these resources, we developed the ASAP platform to support multimodal, automated meta-analysis workflows. In benchmark evaluations, the MASTER agent demonstrated high accuracy and stability in performing core automated meta-analysis tasks. The ASAP platform enabled automated literature retrieval, quality assessment, data extraction, and visualisation generation through predefined workflows. Here, we provide an initial exploration of the technical feasibility and scalability of artificial intelligence-driven automated meta-analysis.
    Keywords:  evidence synthesis; global health; large language models; meta-analysis; reporting bias; research integrity
    DOI:  https://doi.org/10.7189/jogh.16.03029
  6. Front Cell Infect Microbiol. 2026 ;16 1876326
       Background: Confirmed oncogenic microbes contribute significantly to cancer burden. Identifying and confirming novel microbial oncogenicity could yield strategies and tools that will reduce disease burdens. However, relevant evidence may be dispersed across a vast biomedical literature that is infeasible for humans to comprehensively synthesize. Large Language Models (LLMs) may enable scalable, expert-level systematic evidence synthesis to identify high priority microbe-cancer pairs; however, such capabilities have not yet been demonstrated.
    Methods: Domain experts were recruited to create a human-validated test dataset to benchmark the performance of LLMs (Gemini 2.5 Pro, Gemini 2.5 Flash, GPT-5, and GPT-5 Nano) on 24 original research papers using Mouse Mammary Tumor Virus-Like Virus and breast cancer as a case study. We devised a structured template for evidence extraction and appraisal of papers, consisting of multiple choice, Likert-scale, multi-select, and free-text question types (77 question items across 24 papers). Agreement between (1) experts, and (2) experts and each LLM, was determined per question instance using novel scoring metrics. LLMs were assessed by comparing inter-expert and expert-LLM agreement score distributions to determine whether LLMs behaved as additional experts by either increasing or maintaining inter-expert agreement. Free-text responses were further evaluated qualitatively.
    Results: Across all question types, LLM responses aligned closely with expert assessments, with two models (GPT-5, GPT-5 Nano) achieving score distributions statistically indistinguishable from those of experts. Gemini models behaved similarly for most tasks but were significantly more lenient in applying microbial oncogenesis criteria, often over-attributing criteria fulfillment. Hallucinations were rare, although more frequent in smaller models (Gemini 2.5 Flash, GPT-5 Nano). Methodological appraisal and identification of contradictions within full-text papers were the most persistent areas of LLM vulnerability, however, the error rate could not be directly compared with experts.
    Conclusions: Two LLMs (GPT-5, GPT-5 Nano) were indistinguishable from domain experts on structured domain research paper evaluation tasks. This evidence supports use of LLMs for automated systematic evidence synthesis. However, methodological appraisal tasks and contradiction identification in full-text papers remain weaknesses requiring further investigation, strengthening, and possibly multi-model strategies.
    Keywords:  artificial intelligence; biomedical literature; critical appraisal; evidence evaluation; evidence extraction; evidence synthesis; large language models; microbial oncogenesis
    DOI:  https://doi.org/10.3389/fcimb.2026.1876326
  7. Front Cell Infect Microbiol. 2026 ;16 1964702
      [This corrects the article DOI: 10.3389/fcimb.2026.1876326.].
    Keywords:  artificial intelligence; biomedical literature; critical appraisal; evidence evaluation; evidence extraction; evidence synthesis; large language models; microbial oncogenesis
    DOI:  https://doi.org/10.3389/fcimb.2026.1964702
  8. Front Drug Saf Regul. 2026 ;6 1846339
       Introduction: Advances in generative artificial intelligence (AI), particularly large language models (LLMs), have sparked discussions in automating pharmacovigilance (PV) workflows. It remains unclear whether these technological advancements fundamentally change the prior conclusions that full automation of Individual Case Safety Report (ICSR) processing is not feasible.
    Methods: This perspective examines recent developments in AI for PV and introduces a conceptual framework of "computable PV," in which tasks are evaluated based on their computational tractability and suitability for automation.
    Results: Routine, well-defined PV tasks, including completeness checks, detection of duplicated ICSRs, and structured information extraction, are increasingly amenable to automation. In contrast, complex activities such as case-level causality assessment remain difficult to formalize and continue to rely on expert judgment. The emergence of LLMs enables broader, cross-task capabilities compared to traditional task-specific, "small" models, but introduces challenges related to reliability, auditability, and governance. As a result, hybrid architecture combining large models, small models, and rule-based components is increasingly necessary.
    Conclusion: Generative AI, as of today, does not signal full automation of PV but rather shifts toward hybrid human-AI systems. While AI can augment efficiency and support evidence synthesis, final decisions must remain under human oversight. Future PV systems should prioritize transparency, validation, and the integration of AI outputs into expert-driven decision-making.
    Keywords:  AI; computability; human-AI systems; large langauge models; pharmacovigilance (MeSH)
    DOI:  https://doi.org/10.3389/fdsfr.2026.1846339
  9. Expert Rev Med Devices. 2026 Aug 31. 1-5
      
    Keywords:  AI medical device; European Union medical device regulations; artificial intelligence; automation
    DOI:  https://doi.org/10.1080/17434440.2026.2727060
  10. J Med Internet Res. 2026 Aug 31. 28 e98551
       Background: Generative AI (GAI) is rapidly transforming research practices, including qualitative methods in health research. While these tools offer efficiency in processing large volumes of textual data, concerns remain regarding their methodological rigor, interpretive capacity, equity, and ethical implications.
    Objective: This rapid review aimed to synthesize the current evidence on the use of GAI in health-related qualitative research, focusing on its applications, performance relative to human analysis, and implications for rigor, ethics, and equity.
    Methods: We conducted a rapid review following Joanna Briggs Institute and PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines. Peer-reviewed studies published between 2022 and December 2025 were identified through searches in PubMed, Web of Science, and Scopus. Eligible studies included qualitative or mixed methods research that used GAI tools (eg, ChatGPT, Gemini, and Claude) during qualitative analysis, including studies that compared GAI-generated outputs with human researchers, coders, or traditional qualitative analytic approaches. Data were extracted using a structured template and synthesized descriptively. Study quality was assessed using the Critical Appraisal Skills Programme (CASP) checklist. This rapid review was registered with the International Prospective Register of Systematic Reviews (PROSPERO; CRD420261280832). The review adhered to the registered PROSPERO protocol; no deviations occurred.
    Results: A total of 42 studies met the inclusion criteria; 71.4% (n=30) were published in 2025, and 81% (n=34) used qualitative designs. Thematic analysis (n=20, 40%) and content analysis (n=10, 20%) were the most common qualitative approaches. GAI was most applied during data familiarization, coding, and theme development, with ChatGPT being the most frequently reported GAI, accounting for nearly two-thirds of all model occurrences (n=37, 62.7%). Among studies evaluating GAI performance relative to human qualitative analysis, performance was strongest in inductive thematic and content analyses, with agreement often exceeding 80% for descriptive themes but dropping to approximately 30% for culturally nuanced themes. Several studies reported time to complete analyses up to 97% faster than human-led analyses. However, performance declined for reflexive and theory-driven analyses, particularly when interpreting culturally nuanced or emotionally complex data. Across studies, GAI improved efficiency but frequently produced superficial interpretations, misapplied theoretical frameworks, and generated occasional inaccuracies, including fabricated quotes. Human oversight was consistently identified as essential to ensure validity, contextual accuracy, and ethical integrity. Concerns related to bias, transparency, and data privacy were widely reported.
    Conclusions: GAI can effectively support early-stage qualitative analysis and enhance efficiency in health research; however, it cannot replace the interpretive and reflexive functions central to qualitative inquiry. A hybrid human-AI approach is recommended, in which GAI assists with data processing while researchers retain responsibility for interpretation, contextualization, and ethical oversight. Future research should prioritize developing guidelines that address equity, transparency, and responsible integration of GAI into qualitative methodologies.
    Keywords:  evidence synthesis; generative AI; health; large language models; qualitative methods
    DOI:  https://doi.org/10.2196/98551
  11. Eur J Transl Myol. 2026 Sep 01.
      Brain-Computer Interfaces (BCIs) are increasingly used in neurorehabilitation, but the rapid expansion of scientific literature complicates the identification of clinically relevant studies. This study investigated whether expert-defined relevance within BCI rehabilitation literature emerges as a structural property of semantic networks through the integration of graph theory and Principal Component Analysis (PCA). A Lexical Network Analysis Based on Graph Theory (LENGTH) was applied to randomized controlled trials indexed in PubMed over the last decade using the query "brain computer interface" AND rehabilitation. Titles and abstracts were analyzed to construct a semantic network linking articles and lexical terms. Multiple graph-theoretical metrics were calculated and residualized against weighted degree to minimize document-size bias. PCA was subsequently applied to the residualized metrics. Forty-eight studies were included. The network showed a compact and highly interconnected structure, centered on motor and functional recovery concepts. PCA identified two principal components explaining of total variance. Relevant articles tended to occupy regions characterized by higher semantic integration and lower hierarchical influence. Although no clear categorical separation emerged, a consistent positional tendency was observed. These findings suggest that relevance may be represented as a topological property within a multidimensional semantic landscape, supporting the use of semantic-network approaches for literature screening and evidence synthesis.
    Keywords:  Brain-computer interface; graph theory; network analysis; neurorehabilitation; principal component analysis
    DOI:  https://doi.org/10.4081/ejtm.2026.15674
  12. JMIR Form Res. 2026 Sep 02. 10 e100148
       BACKGROUND: Recent advances in AI, particularly large language models, have generated growing interest in their application to medical education and examination preparation. However, the accuracy, reasoning quality, and adherence to clinical guidelines of these tools in postgraduate urology assessments remain unclear.
    OBJECTIVE: This study aimed to evaluate the performance of 3 AI tools, ChatGPT (GPT-4.0), Claude (version 4.5), and AMBOSS, on European Board of Urology (EBU)-style multiple-choice questions, with a particular focus on accuracy, insight, concordance, and adherence to European Association of Urology (EAU) guidelines.
    METHODS: A total of 200 single-best-answer questions from the EBU In-Service Assessment workbook (2021-2022) were input into each AI model. Models were prompted to select an answer and provide an explanation. Two urologists with post-Fellowship of the Royal College of Surgeons (FRCS) training independently assessed the outputs. Accuracy was defined as correct answer selection. Concordance was defined as the logical alignment between the answer and its explanation. Insight was evaluated across 3 domains-nonobvious deduction, discriminative reasoning, and clinical validity-and was graded as low, moderate, or high.
    RESULTS: ChatGPT demonstrated the highest accuracy (171/200, 85.5%), compared to Claude and AMBOSS (both 159/200, 79.5%; P=.14). Concordance was also significantly higher for ChatGPT (190/200, 95%) than for Claude (176/200, 88%) and AMBOSS (152/200, 76%; P<.001). Nonobvious deduction was predominantly low to moderate across all models, reflecting the recall-based nature of many questions. ChatGPT and Claude showed stronger discriminative reasoning, while AMBOSS demonstrated limited exclusion of alternative options. Clinical validity was high overall, with ChatGPT showing the greatest consistency with EAU guidelines. There was substantial agreement between the 2 reviewers (weighted κ coefficient >0.61).
    CONCLUSIONS: AI tools can achieve high accuracy on EBU-style assessments; however, differences in reasoning quality and guideline adherence are evident. ChatGPT demonstrated superior performance across all evaluated domains, supporting its role as a potential adjunct in postgraduate urology education.
    Keywords:  AI; AMBOSS; ChatGPT; Claude; EAU guidelines; EBU; European Association of Urology guidelines; European Board of Urology; artificial intelligence
    DOI:  https://doi.org/10.2196/100148
  13. Spine J. 2026 Aug 31. pii: S1529-9430(26)00631-5. [Epub ahead of print]
       BACKGROUND CONTEXT: Generative artificial intelligence (AI) is increasingly used in spine care; however, concerns remain regarding citation hallucinations and reliability. ChatGPT may generate inaccurate or fabricated references, whereas OpenEvidence (OE) prioritizes verified, peer-reviewed literature. This is the first study comparing OE and ChatGPT using cervical spine clinical guideline (CSCG) queries.
    PURPOSE: To compare guideline alignment, citation validity, sourcing, and prompt-engineering effects between OE and ChatGPT using CSCGs.
    STUDY DESIGN/SETTING: Cross-sectional comparative analysis.
    PATIENT SAMPLE: No patient population was included.
    OUTCOME MEASURES: Primary outcomes were guideline alignment score and citation validity (fully correct, partially hallucinated, or fully hallucinated). Secondary outcomes included source type, publication year, proportion published after CSCG release, and prompt-engineering effects.
    METHODS: A total of 110 evidence-based clinical questions derived from 10 CSCGs authored by 6 academic societies were submitted to OE (v2.0) and ChatGPT-4o from June 1 to July15, 2025. A subset of prompts was repeated to evaluate prompt-engineering effects.
    RESULTS: OE generated 999 citations with 100% accuracy, whereas only 184/393 (46.8%) ChatGPT citations met accuracy criteria (p<0.001). OE demonstrated higher guideline alignment than ChatGPT (4.6 ± 0.8 vs 4.1 ± 0.7; p = 0.03), with almost perfect interrater agreement (weighted Cohen's κ = 0.88; 95% CI, 0.80-0.97). . ChatGPT produced 88 partially hallucinated citations (22.4%), most commonly due to incorrect hyperlinks, author names, or publication years. OE cited more peer-reviewed literature (79.8% vs 60.1%; p<0.001) and more recent studies (2018±5.6 vs 2012±7.4; p<0.001). Prompt-engineering analysis showed OE maintained higher citation validity and fewer hallucinations, although accuracy declined when outputs were reformatted through ChatGPT, suggesting cross-model contamination.
    CONCLUSION: OE demonstrated superior CSCG concordance and substantially lower susceptibility to hallucinations than ChatGPT. Nonetheless, physician oversight remains essential for safe AI integration into clinical practice.
    Keywords:  Artificial Intelligence; Cervical Spine; ChatGPT; Citation Accuracy; Evidence-Based Medicine; OpenEvidence; Prompt Engineering; Temporal Analysis
    DOI:  https://doi.org/10.1016/j.spinee.2026.08.002
  14. Cureus. 2026 Jul;18(7): e113654
       BACKGROUND: Large language models (LLMs) have demonstrated strong performance in generating medical responses; however, concerns persist regarding the accuracy of citations for these responses. Prior work has shown that earlier models frequently fabricate or misattribute references when queried on American Academy of Orthopaedic Surgeons (AAOS) Clinical Practice Guidelines (CPGs). With rapid advancements in model development, newer-generation LLMs may demonstrate improved citation accuracy and reliability. This study evaluates the citation accuracy of ChatGPT-5 in response to AAOS CPG-based queries.
    METHODS: A computer-based observational study evaluating the citation accuracy of ChatGPT-5 (December 2025) was conducted from February to March 2026. Seventy recommendations from four AAOS CPGs were converted into standardized prompts, and ChatGPT was asked to generate a list of references to support its claims in response to these prompts. Five independent graders evaluated responses for citation accuracy. Citation elements assessed included title, authorship, journal, year, volume, pages, and PubMed identifier (PMID). Hallucinated references were defined as nonexistent or non-indexed studies.
    RESULTS: A total of 350 queries generated 2,736 references, of which 7.13% were fabricated. Citation inaccuracies were most frequent for PMID (37.28%), title (30.70%), and pages (27.49%). Overall, 49.34% of citations were bibliographically accurate, and complete accuracy within a query occurred in 8.00% of cases.
    CONCLUSIONS: Citation errors and fabricated references persist in ChatGPT-5. Independent verification of generated references therefore remains necessary.
    Keywords:  artificial intelligence; citation accuracy; citations; clinical practice guidelines; large language models
    DOI:  https://doi.org/10.7759/cureus.113654
  15. J Am Board Fam Med. 2026 Sep;pii: 167736. [Epub ahead of print]39(1):
       INTRODUCTION: Clinicians require concise, accurate summaries of new research to inform practice. Patient-Oriented Evidence that Matters (POEMs), published in American Family Physician, are a benchmark for summarizing primary literature in family medicine, while large language models (LLMs) offer scalable summarization but require rigorous evaluation. The objective of this study was to evaluate the accuracy and quality of summaries generated by large language models compared with expert-authored POEMs.
    METHODS: In this study, we compared LLM-generated summaries (Microsoft Copilot, GPT-4o class) with 24 recent matched POEMs using a standardized prompt. Two trained raters independently scored each summary with a 13-item tool (score range 0-13), cataloged errors, recorded word counts, and indicated preferences on a 5-point scale.
    RESULTS: LLM summaries outperformed POEMs in total score (mean 12.1 vs 10.6; mean difference 1.5, 95% CI 1.1-2.0; P < 0.001), with similar lengths (328 vs 353 words; P = 0.23). Errors occurred in fewer LLM-DOCSs (2/24) than POEMs (9/24), with a mean error score difference of 20% (95% CI 7% -33%; P < 0.001). POEMs most often missed in the categories Contextual Background and Limitations; both approaches frequently missed in Clinical Applicability. Reviewer preference favored LLM-DOCS (mean 2.44 on a 1-5 scale; 95% CI 2.1-2.8).
    CONCLUSIONS: An enterprise LLM, prompted in POEM style, produced accurate, low-error clinical summaries that matched or exceeded expert-edited POEMs and were generally preferred by reviewers, though further research is needed to assess broader applicability and impact. Findings support pragmatic LLM-assisted summarization and highlight the need for standardized evaluation tools and explicit prompts for clinical applicability.
    Keywords:  Clinical Decision Support; Clinical Decision-Making; Evidence-Based Medicine; Family Medicine; Large Language Models; Natural Language Processing
    DOI:  https://doi.org/10.3122/jabfm.2025.250401R1
  16. J Vis Commun Med. 2026 Sep 01. 1-9
      Graphical abstracts are increasingly used to enhance scientific communication, yet their quality remains variable. Generative artificial intelligence (AI) tools can produce graphical abstracts, but their scientific reliability has not been systematically evaluated. To assess the scientific quality of AI-generated graphical abstracts in respiratory medicine. This multicentre study included a pilot phase (5 abstracts; 15 graphical abstracts) and a main phase (20 abstracts; 60 graphical abstracts). For each text abstract, three graphical abstracts were generated via ChatGPT (version 5.2), Claude Sonnet (version 4.5), and Gemini (version 3), using a standardised prompt. Six experts evaluated each graphical abstract using a 7-item scoring grid (total score/28). Inter-rater reliability and comparisons between models were assessed.The median total score was 21 [16-25]. Single-measure intraclass correlation coefficients (ICCs) indicated moderate agreement for the total score (ICC = 0.606), while average-measure ICC showed excellent reliability (ICC = 0.902). Significant differences were observed between models (p < 0.001), with a consistent ranking of Claude > ChatGPT > Gemini. Differences were observed across all evaluation criteria, with large effect sizes (Kendall's W up to 0.93). AI-generated graphical abstracts demonstrate moderate-to-high quality but remain heterogeneous across models. While promising as assistive tools, their use requires expert validation to ensure scientific accuracy and interpretative safety.
    Keywords:  Data visualisation; observer variation; pulmonary medicine; reproducibility of results; scientific communication
    DOI:  https://doi.org/10.1080/17453054.2026.2725547