bims-arines Biomed News
on AI in evidence synthesis
Issue of 2026–09–27
24 papers selected by
Farhad Shokraneh, Systematic Review Consultants LTD



  1. Value Health. 2026 Sep 19. pii: S1098-3015(26)02642-2. [Epub ahead of print]
    GenAI for Systematic Literature Reviews ISPOR Task Force
       OBJECTIVES: Systematic literature reviews (SLRs) are foundational to evidence-based medicine, including health technology assessment (HTA) and health economics and outcomes research (HEOR). Generative artificial intelligence (GenAI) tools are increasingly used in SLR workflows, yet no good practice guidance exists. This ISPOR Task Force report provides evidence-informed recommendations for responsible GenAI use across core SLR tasks.
    METHODS: A PRISMA-adapted rapid evidence assessment identified 115 empirical studies evaluating GenAI in SLR tasks published between November 2022 and July 2025. Findings were synthesized qualitatively across seven tasks. A structured task-level assessment framework spanning eight domains informed good practice recommendations, derived through Task Force consensus among experts in SLR methodology, HTA, AI development, bioethics, and regulatory science.
    RESULTS: Evidence supported GenAI use for high-recall title/abstract screening and structured first-pass data extraction within human-in-the-loop workflows with explicit oversight. Autonomous deployment was not supported. Evidence for full-text screening, qualitative synthesis, and report writing was more conditional, depending on workflow design and oversight, while risk of bias assessment showed the lowest readiness. End-to-end autonomous SLR generation was not recommended. Performance was most reliable within clearly defined workflows, with pre-specified rules for flagging records and explicit human review and conflict-resolution processes.
    CONCLUSIONS: Based on current evidence, GenAI can augment, but not replace, human expertise in SLRs. Responsible use requires evaluating GenAI suitability for each review task, retaining human accountability at all decision points, and documenting AI use as a core methodological component. Because GenAI evolves rapidly, these recommendations reflect evidence through July 2025 and warrant periodic updating.
    Keywords:  Evidence Synthesis; Generative Artificial Intelligence; Health Technology Assessment; Large Language Models; Scoping Review; Systematic Literature Review
    DOI:  https://doi.org/10.1016/j.jval.2026.08.007
  2. J Clin Epidemiol. 2026 Sep 25. pii: S0895-4356(26)00396-3. [Epub ahead of print] 112520
    Evidence Synthesis Infrastructure Collaboration Working Group
       AIM: This paper aims to identify and prioritise actionable, system-level solutions for the safe and responsible use of artificial intelligence (AI) in evidence synthesis.
    BACKGROUND: The increasing integration of AI Digital Evidence Synthesis Tools (AI-DEST) presents both opportunities and challenges. Concerns regarding ethical implications, biases in AI algorithms, and disparities in access to high-quality evidence necessitate a structured approach to ensure that AI is used responsibly and effectively.
    METHODS: The group employed a systematic approach, guided by the SHOW ME the Evidence Consensus framework. The framework that emphasizes transparency, stakeholder engagement, and iterative validation across the evidence lifecycle through structured scoping exercise comprising literature review, stakeholder survey, expert interviews, and iterative group discussions. This was a two-phase process: the first phase focused on problem identification through literature reviews, surveys, and stakeholder interviews, while the second phase concentrated on identifying and prioritizing potential solutions. A total of 50 solutions were initially proposed, which were then refined through a structured prioritization process based on criteria such as innovation, feasibility, and potential impact.
    RESULTS: The group identified several key solutions, including the development of an open-source Evidence Synthesis Studio (ESS) that integrates various Digital Evidence Synthesis Tools (DESTs) and a Comprehensive Evidence Synthesis Plug-In Architecture (CESPIA)- an interoperable framework for the validation, benchmarking, and integration of AI tools. These solutions aim to enhance the efficiency and transparency of evidence synthesis processes while ensuring equitable access for diverse user groups, particularly in the Global South.
    DISCUSSION: The proposed solutions emphasize the importance of engagement with key interest holders, including evidence synthesis practitioners, policymakers, and tool developers, alongside broader transparency-oriented public engagement. By addressing the unique challenges faced by underrepresented communities, the group aims to mitigate biases and promote ethical AI use. The integration of citizen feedback and iterative design processes is crucial for developing tools that meet the needs of diverse interest holders.
    CONCLUSIONS: The recommendations from Working Group 3 provide a comprehensive framework for the responsible use of AI-DEST. By fostering collaboration, transparency, and inclusivity, these solutions aim to strengthen evidence synthesis as a public good, ensuring that high-quality evidence is accessible to all, regardless of background or resources.
    DOI:  https://doi.org/10.1016/j.jclinepi.2026.112520
  3. J Clin Epidemiol. 2026 Sep 21. pii: S0895-4356(26)00390-2. [Epub ahead of print] 112514
       OBJECTIVE: Machine learning tools for literature screening have existed for a decade, yet adoption remains limited: active learning needs hundreds of decisions before prioritizing well, and commercial large language model (LLM) chatbots are used without task-specific training. We aimed to develop TITAN-SR (Training Infrastructure for Automated Nomination in Systematic Reviews), a screening tool built on biomedical Bidirectional Encoder Representations from Transformers (BERT) models, and to compare it with the active-learning tool ASReview and commercial LLM chatbots.
    STUDY DESIGN AND SETTING: We assembled 762,934 citation-label pairs from 19,787 completed reviews and trained TITAN-SR on a prespecified review-level split, testing on held-out reviews. The tool combines biomedical BERT models (PubMedBERT and BioLinkBERT) that read a review's eligibility criteria with each citation's title and abstract; training penalized false negatives 10 times more heavily than false positives. We compared TITAN-SR with four chatbots used zero-shot on a stratified 500-review benchmark (Claude Sonnet 4, GPT-4o, DeepSeek-V3, Gemini 2.0 Flash), with GPT-4o-mini on all 3,434 evaluable test reviews, and with ASReview on 22 temporally independent Cochrane reviews (142,504 records); ASReview received a warm-up of 20% of known includes.
    RESULTS: Across 3,434 test reviews, TITAN-SR achieved a median specificity at 99% recall of 0.902 and median area under the receiver operating characteristic curve (AUC) of 0.974. It outperformed all four chatbots (Bonferroni-corrected paired Wilcoxon p ≤ 5 × 10-31) and GPT-4o-mini. The chatbots showed calibration collapse: 82-99% of confidence scores were ≥ 0.9 regardless of accuracy. On the 22 external reviews performance was virtually identical (median AUC 0.976; specificity at 95% recall 0.897), and at a threshold fixed on the internal validation split TITAN-SR met the pre-registered primary criterion, retaining 99.8% of included studies (95% confidence interval [CI] 99.5 to 100.0%). It outperformed ASReview on 18 of 22 reviews (median paired difference in specificity at 95% recall +0.145, 95% CI +0.101 to +0.201).
    CONCLUSION: A purpose-built screening tool based on biomedical BERT models, developed from nearly 20,000 completed reviews, outperformed ASReview and commercial LLM chatbots evaluated without task-specific training. Temporal validation on 22 held-out Cochrane reviews confirmed generalization to reviews published after the model-development period.
    Keywords:  evidence synthesis; large language models; machine learning; natural language processing; screening automation; systematic review
    DOI:  https://doi.org/10.1016/j.jclinepi.2026.112514
  4. Res Synth Methods. 2026 Sep 21. 1-20
      Automation-assisted title-abstract screening is routinely evaluated using metrics such as recall, precision, and workload reduction, and results are increasingly summarized across studies. Synthesis presupposes that evaluation results refer to comparable target quantities. This article examines whether screening evaluation results produced under diverse automation-assisted workflows can meaningfully support cumulative inference, and under what conditions cross-study comparison is warranted. This conceptual article treats screening evaluation targets as workflow-induced rather than task-inherent. Evaluation design is decomposed into four workflow dimensions (representation, inference, governance, and evaluation) plus the reference-decision set produced or specified. Recurring choice patterns across these elements define evaluation configurations. Two study-level configurations and one synthesis-level pattern are illustrated: endogenous adjudication, inherited adjudication, and aggregation of heterogeneous targets. Across configurations, evaluation targets are induced by workflow-specific reference-decision sets and, in some designs, model-based target domains rather than by a common reference set. Nominally similar metrics can therefore estimate different quantities across studies, and apparent cumulativeness can arise from pooling estimates of different quantities. Standard heterogeneity statistics cannot distinguish variation in method effectiveness from variation in what is being estimated. Comparison is warranted only when targets align before pooling. Differences across screening evaluations may reflect differences in what is being estimated rather than statistical heterogeneity around a common estimand. When evaluation targets are not aligned across the framework, aggregate summaries can describe reported results but cannot support cumulative inference. Improved aggregation requires both better study-level specification of evaluation targets and explicit attention to configuration variation at synthesis .
    Keywords:  abstract screening; automation; evaluation methodology; research synthesis; systematic reviews
    DOI:  https://doi.org/10.1017/rsm.2026.10116
  5. JAMIA Open. 2026 Oct;9(5): ooag179
       Objectives: Large language models (LLMs) offer significant potential for automating clinical trial classification by eligibility criteria. However, the optimal input data remain unclear: while abstracts provide a condensed signal, full-text articles contain substantially more information. Whether this additional signal improves performance or whether accompanying noise negatively affects the model's reasoning capabilities remains unclear.
    Materials and Methods: GPT-5 was applied to classify 200 randomized controlled oncology trials, labelling whether patients with localized and/or metastatic disease were eligible. Each trial was classified twice-using the abstract and full text-and outputs were compared with manually annotated ground-truth labels. Performance was assessed using accuracy, precision, recall, and F1 score, and statistical significance using the McNemar test.
    Results: For identifying trials including patients with localized disease, GPT-5 achieved an accuracy of 86% (95% CI, 81%-91%; F1 = 0.90) using abstracts and 92% (95% CI, 88%-95%; F1 = 0.94) using full texts (P = .027). Performance for detecting trials, which include patients with metastatic disease, was comparably high (99% vs 98% accuracy; F1 = 1.00-0.99). Overall accuracy for assigning combined labels increased from 86% (95% CI, 81%-91%) using abstracts to 92% (95% CI, 88%-95%) using full texts (P = .027).
    Discussion and Conclusion: Providing full-text articles to GPT-5 significantly improved the classification of oncology trials by eligibility criteria in this dataset. Full-text analysis appears particularly valuable for extracting eligibility criteria in oncology that are frequently omitted or not explicitly described within the abstract.
    Keywords:  eligibility criteria; large language models; oncology; text mining
    DOI:  https://doi.org/10.1093/jamiaopen/ooag179
  6. Life (Basel). 2026 Sep 18. pii: 1566. [Epub ahead of print]16(9):
       BACKGROUND: Radiotherapy-related complications can impair long-term outcomes in nasopharyngeal carcinoma (NPC), while high-volume literature screening remains a bottleneck in evidence synthesis.
    METHODS: Four databases were searched through January 2026. After deduplication, 4916 records underwent two-stage manual screening, yielding 38 studies. Three Bidirectional Encoder Representations from Transformers (BERT) models, BERT-base, BioBERT, and PubMedBERT, were fine-tuned using stratified five-fold cross-validation, and screening performance was assessed using the area under the receiver operating characteristic curve (ROC-AUC), the area under the precision-recall curve (PR-AUC), and work saved over sampling (WSS) at the 100% and 85% recall thresholds (WSS@100% and WSS@85%).The clinical evidence was evaluated using the Prediction Model Risk of Bias Assessment Tool (PROBAST) and synthesized narratively according to endpoint, discrimination metric, time horizon, unit of analysis, validation level, and uncertainty reporting.
    RESULTS: At complete recall, BioBERT and PubMedBERT achieved mean WSS values of 97.8% and 97.5%, respectively, compared with 89.7% for BERT-base. Separately, the clinical evidence review identified 39 discrimination estimates: 33 conventional binary ROC-AUCs, three Harrell's C-indices, and three time-dependent AUCs. Only two studies contributed external-validation estimates and one contributed a temporal-validation estimate. Sixteen estimates lacked a usable measure of uncertainty, and PROBAST rated all studies as having a high overall risk of bias, driven by the analysis domain.
    CONCLUSIONS: Within the retrieved and manually annotated corpus, biomedical pretrained language models showed the potential to reduce the title-and-abstract screening burden while maintaining complete recall. Independently, the clinical evidence map identified promising discrimination but insufficiently robust and transportable evidence for routine clinical implementation.
    Keywords:  abstract screening; nasopharyngeal carcinoma; pretrained language models; radiotherapy complications; systematic review
    DOI:  https://doi.org/10.3390/life16091566
  7. J Med Internet Res. 2026 Sep 25. 28 e84915
       Background: Large language models (LLMs) have the potential to improve the efficiency of evidence synthesis, but their reliability in performing complex tasks such as risk-of-bias (ROB) assessment in randomized controlled trials (RCTs) remains unclear.
    Objective: This study aimed to evaluate whether LLMs can reliably assess ROB in RCTs using version 2 of the Cochrane ROB tool for randomized trials (ROB 2).
    Methods: This study was conducted between December 28, 2024, and February 28, 2025, in adherence to American Association for Public Opinion Research reporting guidelines. Twenty-nine RCTs were selected from published Cochrane systematic reviews across diverse medical fields. We developed a structured prompt engineering framework that transformed ROB 2 decision trees into logical rules for the LLM. Each RCT was independently evaluated twice by ChatGPT, with Cochrane review authors' assessments serving as the reference standard for comparison. The main outcomes were the accuracy and consistency of ROB 2 assessments at both the domain and trial levels, evaluated using accuracy, sensitivity, specificity, and F1-score. Consistency between the repeated assessments was quantified using the Cohen κ and prevalence-adjusted, bias-adjusted κ.
    Results: The LLM demonstrated a moderate aggregate domain accuracy of 73.1% (95% CI 64.7%-81.5%) in the first assessment and 75.9% (95% CI 66.3%-85.4%) in the second assessment. Domain-averaged sensitivity decreased from 61.4% (95% CI 48.1%-74.7%) to 53.4% (95% CI 41.7%-65.0%), whereas domain-averaged specificity increased from 75.8% (95% CI 65.3%-86.3%) to 81.1% (95%CI 67.1%-95%), indicating a conservative tendency in identifying a high ROB. Domain-level accuracy ranged from 62.1% to 87.9%, with the lowest accuracy observed in domain 1 and the lowest F1-scores observed in domain 2. Consistency between repeated assessments was high, with a mean agreement of 89.0% (SD 7.5%), and Cohen κ values were 0.86, 0.39, 0.56, 0.84, and 0.85 in domains 1 to 5, respectively.
    Conclusions: In this exploratory study, ChatGPT demonstrated moderate accuracy and high consistency in assessing ROB in RCTs using the ROB 2. However, its reliability diminished in complex scenarios requiring interpretation of implicit narratives or behavioral nuance. These findings suggest that LLMs may support methodological evaluations in systematic reviews by acting as automated screeners to reduce reviewer burden, but current implementation still requires expert oversight, particularly for trials involving subjective outcomes or nonstandard reporting.
    Keywords:  ChatGPT; LLMs; RCT; ROB 2; large language models; randomized controlled trial; risk of bias; systematic reviews; version 2 of the Cochrane risk-of-bias tool for randomized trials
    DOI:  https://doi.org/10.2196/84915
  8. PLoS One. 2026 ;21(9): e0358873
       BACKGROUND: Peer review processes may inadequately assess compliance with established reporting guidelines such as the Consolidated Standards of Reporting Trials (CONSORT) criteria. Large language models (LLMs) demonstrate potential for systematic manuscript evaluation; however, how consistently different models assess adherence to CONSORT guidelines in published clinical trials remains unexplored.
    METHODS: Twenty randomized controlled trials published in immunology journals between 2015 and 2016 were identified through PubMed. Three LLMs (ChatGPT-4o, Gemini 2.5 Flash, and Claude Sonnet 4.6) independently assessed compliance across 37 CONSORT 2010 subpoints. The primary endpoint was the difference between models in mean CONSORT compliance score. Secondary endpoints included inter-model agreement and the proportion of articles meeting a 90% compliance threshold. Statistical analysis employed repeated measures analysis of variance (ANOVA) with post-hoc pairwise comparisons (α = 0.05).
    RESULTS: Mean CONSORT compliance rates were: ChatGPT-4o 80.8% (95% CI 75.8-85.8%), Gemini 2.5 Flash 64.4% (95% CI 58.0-70.8%), and Claude Sonnet 4.6 54.7% (95% CI 48.0-61.3%). Using a 90% compliance threshold as a quality benchmark, ChatGPT-4o identified 25% of papers (5/20) as meeting this standard, while Gemini 2.5 Flash and Claude Sonnet 4.6 each identified none (0/20) as meeting this standard. Repeated-measures ANOVA demonstrated significant differences between models (F(2,38) = 43.01, p < 0.001, partial η² = 0.694). All pairwise comparisons were statistically significant (ChatGPT-4o versus Gemini 2.5 Flash and ChatGPT-4o versus Claude Sonnet 4.6, both p < 0.001; Gemini 2.5 Flash versus Claude Sonnet 4.6, p = 0.014).
    CONCLUSIONS: LLMs varied substantially in their assessment of CONSORT compliance in published randomized trials, with a consistent ordering: ChatGPT-4o scored compliance highest, Gemini 2.5 Flash intermediate, and Claude Sonnet 4.6 lowest. Item-level agreement between models was only moderate, with no pairwise weighted kappa reaching the 0.61 substantial-agreement threshold. This inter-model variability indicates the need for standardized evaluation protocols before LLM-assisted manuscript screening is adopted.
    DOI:  https://doi.org/10.1371/journal.pone.0358873
  9. JSLS. 2026 Jul-Sep;30(3):pii: e2026.00066. [Epub ahead of print]30(3):
       Background and Objectives: The worthiness of a medical article is not known until it is read and critically evaluated. A process using artificial intelligence (AI) was developed and assessed to do that for a specific article.
    Methods: A medical article was uploaded to an AI large language model program with instructions to answer specific bounded questions about that article only, its methodological components, assessment of findings, and appropriateness and validity of conclusions. Fifty articles were evaluated and compared using AI and human assessment.
    Results: Agreement was found 89% of the time. Disagreements were concerned with word use, context, and by experiences and human memory. No AI fabrications occurred.
    Conclusion: This method eliminated, or at least heightened awareness of the problems, biases, limitations, and confounders of an article to be read. Isolation and validity of the components that collectively produce the conclusions from what is found and said in an article is more clearly delineated by a step-by-step process, eliminates human legacy conflicts, and quickly determines the legitimacy of conclusions.
    Keywords:  AI; Artificial intelligence; LLM; large language model; methodology
    DOI:  https://doi.org/10.4293/JSLS.2026.00066
  10. Cureus. 2026 Aug;18(8): e114999
      As oncology research continues to advance, the synthesis of clinical trial evidence is becoming more demanding for healthcare professionals, yet it remains critical for improving the provision of care. While sometimes inconsistent, large language models (LLMs) have emerged as a potential solution to automate literature screening and data extraction. This technical report describes the iterative development of an AI-assisted literature synthesis tool for radiation oncology research. A screening and data extraction tool was developed in three phases using ChatGPT-4o. In Phase 1, prompts were iteratively refined to screen 840 abstracts for Phase 2/3 radiotherapy trials in small cell lung cancer (SCLC). Screening performance was compared against manual reviewer consensus using sensitivity and specificity. In Phase 2, 13 clinical trial variables were extracted from 18 eligible studies and evaluated against a manual extraction reference with a weighted scoring rubric. In Phase 3, a web-based application was developed using the optimized prompts from the first two phases and pilot-tested on eight healthcare professionals for usability and workflow relevance. In the first two phases, the initial prompts had many false positives and inconsistent extractions. Iterative prompt refinement across the second and third batches of studies improved both screening accuracy and extraction consistency. In the validation batch, the final screening model achieved 100% sensitivity and specificity, and the data extraction prompts obtained a mean score of 12.0 (SD: 0.5) out of 13. During Phase 3, pilot testers reported that the application helped them read and compare clinical trial data through concise, structured tables. Users also provided suggestions for future development, including role-dependent personalization and visualization tools. Prompt-engineered LLM models show potential for improving efficiency and accessibility of literature screening and data extraction in oncology research. Future work will integrate the feedback obtained and externally validate the tool using larger independent datasets and multidisciplinary user cohorts.
    Keywords:  artificial intelligence in healthcare; data extraction; large language models; literature screening; prompt engineering; radiation oncology; small-cell lung cancer
    DOI:  https://doi.org/10.7759/cureus.114999
  11. Med Sci (Basel). 2026 Aug 31. pii: 533. [Epub ahead of print]14(5):
       BACKGROUND/OBJECTIVES: Artificial intelligence tools have emerged as promising methodological support for systematic reviews and health technology assessment (HTA). Smart infusion pump interoperability represents a relevant case study due to its implications for medication safety, nursing workflow, and hospital quality improvement. The aim was to evaluate the performance of artificial intelligence as a methodological support tool across a systematic review, using the evidence synthesis on smart infusion pump-electronic health record interoperability as a case study.
    METHODS: A systematic review following PRISMA 2020 guidelines was conducted. Searches were performed in MEDLINE, Embase, and Cochrane Library databases. AI-assisted tools (ChatGPT GPT-4 and Open Science Reviewer) were incorporated into data extraction, reporting appraisal based on STROBE criteria, and exploratory identification of methodological limitations under strict human supervision. Concordance between AI-assisted and manual extraction was evaluated descriptively.
    RESULTS: Overall concordance between AI-assisted and manual data extraction was 82.5% (99/120) across assessed variables. Agreement was highest for structured variables, including study design (10/10; 100.0%), study identification variables (19/20; 95.0%), and participant characteristics (36/40; 90.0%). Agreement was lower for study content variables (18/30; 60.0%) and methodological appraisal (16/20; 80.0%). Among the 21 discrepancies, misclassification errors were most common (13/21; 61.9%), followed by omissions (3/21; 14.3%), incomplete data (3/21; 14.3%), and hallucinations (2/21; 9.5%). AI-assisted identification of methodological limitations showed substantial descriptive agreement with human assessments but demonstrated limited capacity for judgmental interpretations.
    CONCLUSIONS: Artificial intelligence demonstrated utility for structured review tasks such as data extraction and reporting appraisal, but showed limitations in tasks requiring interpretative and methodological judgement. Human oversight therefore remains essential throughout the review process. These findings derive from a single case study and should not be generalised beyond the evaluated context and AI tools.
    Keywords:  artificial intelligence; large language models; smart infusion pumps; systematic review
    DOI:  https://doi.org/10.3390/medsci14050533
  12. Cureus. 2026 Aug;18(8): e115164
      Background The growing adoption of artificial intelligence (AI) technologies, including large language models (LLMs), such as ChatGPT (OpenAI, San Francisco, CA) and many others, has expanded their role in statistical analysis and scientific research workflows. Despite this increasing use, conventional statistical software, such as IBM SPSS (IBM Corp., Armonk, NY), remains the accepted reference standard because of its established methodological precision, transparency, and reproducibility. Consequently, it is important to determine whether statistical results generated by LLMs achieve comparable levels of accuracy and consistency. As these systems undergo continual refinement, independent replication, comparative validation, and continuation studies are necessary to assess whether newer model generations produce findings that reliably correspond with those obtained using established statistical methodologies. Methodology Thirteen statistical procedures commonly used in clinical, medical, epidemiological, and applied health sciences research were evaluated using authentic datasets derived from previously published peer-reviewed studies and undergraduate applied health sciences coursework. Analyses included epidemiological tests, nonparametric analyses, chi-square analyses, univariable and multivariable logistic regression, Kaplan-Meier survival analysis, Cox proportional hazards regression, and measures of nominal association. For each procedure, dataset and variable names were copied directly from SPSS 31.0 data files and analyzed in GPT-5.5 using standardized prompts. AI-generated results were then compared directly with SPSS 31.0 outputs to evaluate computational agreement, methodological consistency, and inferential concordance across all statistical procedures. Results GPT-5.5 demonstrated a high degree of agreement with SPSS 31.0 across all statistical procedures evaluated, including risk and odds ratios, Wilcoxon signed-rank, Mann-Whitney U, Kruskal-Wallis, Friedman's test, one-way and two-way chi-square, binary and multivariable logistic regression, Kaplan-Meier survival analysis, Cox proportional hazards model, and phi and Cramer's V. Most analyses produced identical test statistics, effect estimates, confidence intervals, and inferential conclusions. Minor numerical differences were observed only in selected approximation statistics, including standardized Z values, Wald statistics, and Cox regression parameter estimates, and were attributable to expected implementation differences in optimization routines, convergence criteria, tie corrections, or reporting conventions. Importantly, none of these differences altered statistical significance, effect interpretation, or the overall scientific conclusions. Conclusions GPT-5.5 demonstrated excellent statistical consistency with SPSS 31.0 across a broad range of statistical analyses, producing complete or near-complete agreement in nearly all computational and inferential outcomes. Minor numerical differences were limited to implementation-specific approximations and did not affect statistical significance or substantive interpretation. These findings support the use of GPT-5.5 as a reliable adjunct for statistical verification, interpretation, and research support while reinforcing that independent methodological oversight and validation with established statistical software remain essential for scientific research.
    Keywords:  ai-assisted statistics; alignment; chatgpt; ibm spss; large language models
    DOI:  https://doi.org/10.7759/cureus.115164
  13. Obs Stud. 2026 ;12(2): 297-310
      Sensitivity analysis methods such as the Cornfield's inequality and the E-value were developed to assess the robustness of observed associations against unmeasured confounding - a major challenge in observational studies. However, the calculation and interpretation of these methods can be difficult for clinicians and interdisciplinary researchers. Recent advances in large language models (LLMs) offer accessible tools that could assist sensitivity analyses, but their reliability in this context has not been studied. We assess four widely used LLMs, ChatGPT, Claude, DeepSeek, and Gemini, on their ability to conduct sensitivity analyses using the Cornfield's inequality and the E-value. We first extract study-specific information (exposures, outcomes, measured confounders, and effect estimates) from four published observational studies in different fields. Using such information, we develop structured prompts to assess the performance of the LLMs in three aspects: (1) accuracy of E-value calculation, (2) qualitative interpretation of robustness to unmeasured confounding, and (3) suggestion of possible unmeasured confounders. To our knowledge, there has been little prior work on using LLMs for sensitivity analysis, and this study is an early investigation in this area. The results show that ChatGPT, Claude, and Gemini accurately reproduce the E-values, whereas DeepSeek shows small biases. Qualitative conclusions from all the LLMs align with the magnitude of the E-values and the reported effect sizes, and all models identify biologically and epidemiologically plausible unmeasured confounders. These findings suggest that, when guided by structured prompts, LLMs can effectively assist in evaluating unmeasured confounding, and thereby can support study design and decision-making in observational studies.
    Keywords:  Cornfield’s inequality; E-values; Large language model; Sensitivity analysis; Unmeasured confounding
    DOI:  https://doi.org/10.1353/obs.00013
  14. Am Heart J Plus. 2026 Oct;70 100891
      
    Keywords:  Discharge disposition; Health disparities; Healthcare outcomes; Hospital discharge; Hospitalized patients; Machine learning; Predictive modeling; Social determinants of health
    DOI:  https://doi.org/10.1016/j.ahjo.2026.100891
  15. J Med Internet Res. 2026 Sep 21. 28 e92090
       Background: Large language models (LLMs) exhibit extensive medical knowledge but are prone to hallucinations and show low fact-level explainability, limiting clinical adoption and regulatory compliance. Existing approaches, such as retrieval-augmented generation, partially address these issues by grounding answers in source documents; however, the aforementioned problems persist.
    Objective: We propose the application of an atomic fact-checking framework designed to enhance the reliability and explainability of LLMs in medical long-form question answering. By decomposing generated answers into discrete atomic facts and verifying each against an authoritative knowledge base of medical guidelines, this approach enables precise identification and correction of incorrect statements, alongside explicit linkage to supporting literature.
    Methods: The fact-checking algorithm operates within a retrieval-augmented generation framework: LLM-generated answers are decomposed into atomic facts (smallest and self-contained information units), each of which is assessed and corrected if FALSE. To determine an optimal strategy, the validation-question and answer (Q&A) set on prostate cancer treatment was tested under varying instructions. An extensive evaluation, including multireader assessments by human medical experts and the automated open Q&A benchmark AMEGA (Autonomous Medical Evaluation for Guideline Adherence), was conducted for the final pipeline. In addition to another radiooncologic test-Q&A set, anonymized real-world tumor board cases and an independent, established neurology-Q&A set were used. Given their transparency and accessibility advantages, we compared various open-source models in pairs of generalist models and their medical fine-tuned counterparts, with regard to performance and improvements by fact-checking.
    Results: The framework significantly reduced hallucinations and inaccuracies. Medical expert assessment and automated benchmarks demonstrated significant improvements in factual accuracy, achieving up to a 50% overall answer improvement and an 80% hallucination detection rate. Notably, the observed gain was strongest in real tumor-board questions-the most challenging dataset. Additionally, the framework achieved high explainability by tracing each atomic fact back to the most relevant chunks from the database, providing a granular, transparent explanation of the generated responses.
    Conclusions: To conclude, we present the application of an atomic fact-checking algorithm to medical Q&A. It identifies factual inaccuracies and hallucinations in LLM-generated answers, achieving the greatest gains on clinically realistic, complex questions. Correction via fact-checking improves the overall answer quality while achieving fact-wise explainability, paving the way for more credible clinical use of LLMs.
    Keywords:  LLM; RAG; atomic fact; atomic fact-checking; autoevaluation; backtracing; fact-checking; hallucination; large language model; medical Q&A; medical fine-tuned; open-source; prompt engineering; question and answer; radiation oncology; retrieval-augmented generation; rubrics
    DOI:  https://doi.org/10.2196/92090
  16. medRxiv. 2026 Sep 20. pii: 2026.09.17.26363303. [Epub ahead of print]
       BACKGROUND: Rare diseases affect an estimated 300 million people worldwide, yet the research needed to guide diagnosis and treatment is often fragmented across multiple unstructured literature sources. Natural history studies (NHS) are a key source of this evidence, but manually extracting structured information from NHS publications can be tedious and does not scale.
    METHODS: We have developed a proof-of-concept for an information extraction pipeline testing three opensource large-language models (LLMs) -- Athena-v3-AWQ, Google's Gemma3-27B, and Meta's Llama-3.1-70B-Instruct, to extract key NHS characteristics from PubMed abstracts curated from a Chan Zuckerberg Initiative disease research state model corpus (302 gold-standard and 8,338 full-corpus abstracts), and compared the models on efficiency, extraction completeness, and expert-rated accuracy.
    RESULTS: All three models processed abstracts with success rates exceeding 99%. However, Gemma achieved the best overall performance, with the highest expert-rated accuracy (68.0% of outputs rated "good" vs. 36.0% for Llama and 10.0% for Athena) and the fastest runtime on the full corpus (~16 minutes for 3,547 abstracts), despite Llama scoring higher on the automated Token F1 metric (0.874 vs. 0.723), highlighting a divergence between automated and human evaluation. Athena's lower performance was largely attributable to verbatim copying rather than synthesis of extracted content.
    CONCLUSIONS: These findings illustrate how locally deployed open-source LLMs can extract structured NHS characteristics at scale, thus supporting their use to accelerate evidence synthesis in rare disease research.
    DOI:  https://doi.org/10.64898/2026.09.17.26363303
  17. J Prosthet Dent. 2026 Sep 22. pii: S0022-3913(26)00586-X. [Epub ahead of print]
       STATEMENT OF PROBLEM: Although general-purpose large language models (LLMs) have shown some promise in supporting dental treatment planning, their performance remains suboptimal. This limitation underscores the need for specialized models tailored to restorative dentistry.
    PURPOSE: The purpose of this study was to develop and evaluate a domain-specific LLM to provide evidence-based restorative treatment planning (RTP) for endodontically treated teeth (ETT).
    MATERIAL AND METHODS: The pipeline included knowledge collection from a dataset based on 39 clinical factors, textbooks, and peer-reviewed literature to establish RTP-GPT; supervised fine-tuning with clinical treatments; integration of a retrieval-augmented generation (RAG) system; and benchmarking against general-purpose LLMs. Twenty novel scenarios were tested across 4 models. Two blind evaluators scored outputs in 5 domains: accuracy, conservativeness, recommendation completeness, justification completeness, and analytical reasoning. The Kruskal-Wallis and Dunn post hoc tests assessed differences, and intraclass correlation coefficients (ICCs) measured reproducibility across 3 testing days (α=.05).
    RESULTS: RTP-GPT achieved the best overall performance, significantly outperforming ChatGPT and RTP-GPT+RAG in all domains (Bonferroni-corrected P<.05). Gemini ranked second overall and showed stronger analytical reasoning than RTP-GPT+RAG but was less conservative than RTP-GPT (P<.001). RTP-GPT+RAG surpassed ChatGPT in conservativeness and the completeness of recommendations and justifications. Regarding reliability, RTP-GPT+RAG demonstrated the highest reproducibility with good-to-excellent ICCs, while RTP-GPT, despite superior accuracy, showed the lowest reproducibility in accuracy, conservativeness, and treatment completeness.
    CONCLUSIONS: RTP-GPT provided the strongest performance in RTP for ETT, outperforming Gemini in conservativeness and both RTP-GPT+RAG and ChatGPT across all domains. However, its limited reproducibility underscores the need for further considerations.
    DOI:  https://doi.org/10.1016/j.prosdent.2026.08.012
  18. NPJ Artif Intell. 2026 ;2(1): 69
      Recent progress in large language models (LLMs) has leveraged their in-context learning (ICL) abilities to enable quick adaptation to unseen biomedical NLP tasks. By incorporating only a few input-output examples into prompts, LLMs can rapidly perform these new tasks. While the impact of these demonstrations on LLM performance has been extensively studied, most existing approaches prioritize representativeness over diversity when selecting examples from large corpora. To address this gap, we propose Dual-Div, a diversity-enhanced data-efficient framework for demonstration selection in biomedical ICL. Dual-Div employs a two-stage retrieval and ranking process: First, it identifies a limited set of candidate examples from a corpus by optimizing both representativeness and diversity (with optional annotation for unlabeled data). Second, it ranks these candidates against test queries to select the most relevant and non-redundant demonstrations. Evaluated on three biomedical NLP tasks (named entity recognition (NER), relation extraction (RE), and text classification (TC)) using LLaMA 3.1 and Qwen 2.5 for inference, along with three retrievers (BGE-Large, BMRetriever, MedCPT), Dual-Div consistently outperforms baselines-achieving up to 5% higher macro-F1 scores-while demonstrating robustness to prompt permutations and class imbalance. Our findings establish that diversity in initial retrieval is more critical than ranking-stage optimization, and limiting demonstrations to 3-5 examples maximizes performance efficiency.
    Keywords:  Computational biology and bioinformatics; Mathematics and computing
    DOI:  https://doi.org/10.1038/s44387-026-00123-0
  19. Front Nutr. 2026 ;13 1815038
      Animal-source foods are nutritionally "plastic," with fatty acid profiles, vitamins, and minerals that can be altered by animal diets and management. Although the literature documenting these effects is extensive, it is fragmented across disciplines and experimental contexts, limiting translation into actionable guidance for food quality and human nutrition. We developed the Intelligent System for Integrating Global Human & Animal Health Technology (INSIGHT), a domain-specialized retrieval-augmented generation (RAG) system designed to synthesize evidence across animal production and human nutrition research with explicit provenance. INSIGHT employs a nine-stage RAG pipeline that integrates query expansion, hybrid retrieval, evidence reranking, and self-verification to deliver transparent, citation-linked responses. Throughout, retrieval refers to document selection, integration to the combination of retrieved evidence into a coherent evidence set, and synthesis to the LLM-based generation of grounded narrative responses. The knowledge base comprises ~4,000 peer-reviewed papers in animal science, feed composition, and human dietary research. To evaluate retrieval performance across diverse literature contexts, we developed a multi-group evaluation framework: 282 documents were randomly selected and organized into 26 semantically coherent groups of ~10 papers each. For each group, Perplexity Deep Research generated 25 question-answer pairs and identified ground-truth relevant documents. Each question was posed to INSIGHT, yielding document-level precision, recall, and F1-score metrics across 614 total queries. Generalized linear mixed models with beta regression revealed significant between-group performance variation (p < 0.0001), indicating that retrieval effectiveness depends on semantic domain characteristics. Across groups, INSIGHT achieved a mean precision of 0.77 (SD = 0.12), recall of 0.62 (SD = 0.15), and F1-score of 0.67 (SD = 0.10), all significantly exceeding a 0.5 baseline (p < 0.0001). These results demonstrate that INSIGHT provides reliable document-level retrieval across diverse topical domains, though significant between-group performance variation (p < 0.0001) indicates that retrieval effectiveness is context-dependent. The present evaluation is intentionally scoped as a controlled retrieval performance assessment; end-to-end synthesis quality evaluation using systematic review benchmarking and RAG assessment metrics represents a critical next step. Domain-specialized, evidence-grounded systems such as INSIGHT can accelerate cross-disciplinary knowledge integration, with continued development focused on improving recall in semantically complex domains, enhancing terminology normalization, and expanding corpus coverage across animal production and human nutrition research.
    Keywords:  animal nutrition; animal-source foods; decision support systems; evidence-grounded synthesis; human nutrition; large language models; literature synthesis; retrieval-augmented generation
    DOI:  https://doi.org/10.3389/fnut.2026.1815038
  20. Braz J Anesthesiol. 2026 Sep 19. pii: S0104-0014(26)00113-2. [Epub ahead of print] 844833
      
    DOI:  https://doi.org/10.1016/j.bjane.2026.844833
  21. J Am Med Inform Assoc. 2026 Sep 23. pii: ocag130. [Epub ahead of print]
       OBJECTIVE: Pregnant and breastfeeding women, along with children, represent populations persistently underrepresented in clinical research, leading to uncertainty in clinical decision-making. This study investigated the landscape of pharmacology research concerning pharmacokinetics (PK), pharmaco-epidemiology (PE), and clinical trials (CT) within maternal and pediatric patient populations.
    MATERIALS AND METHODS: A pharmacology research landscape in maternal and pediatric patient populations was analyzed using a novel co-learning framework applied to 38 million PubMed abstracts. It incorporated an evolving annotation guideline for abstract labeling, uncertain sampling, annotation error calibration, and large language model training. The co-learning framework enabled human annotators and language models to iteratively learn from each other, optimizing the classification of PK/PE/CT papers. Analysis was restricted to medications prescribed for maternal and pediatric patients, focusing on publication evidence.
    RESULTS: Co-learning significantly increased the F1-scores for classifying the CT, PE, and PK abstracts from (0.81, 0.84, 0.73) to (0.96, 0.94, 0.93), respectively. On the external validation set, the F1-score increased from (0.81, 0.60, 0.42) to (0.95, 0.88, 0.91) for CT, PE, and PK, respectively. The models identified 376 231 and 879 936 abstracts relevant to maternal and pediatric patients, respectively. In pregnant or postpartum women, 22%-66% prescribed drugs lacked either PK or PE/CT evidence. Among pediatric patients, 8%-40% prescribed drugs had no or weak PK or PE/CT evidence.
    DISCUSSION: Co-learning is a more effective approach than active learning in obtaining feedback from human annotators for optimizing the language model. With sufficiently labeled samples, a language model such as BioBERT outperformed large language models.
    CONCLUSION: In both maternal and pediatric pharmacology research, the absence of PK evidence was notably more frequent compared to PE and CT. This significant finding calls for more pharmacokinetics studies within pediatric and maternal patient populations.
    Keywords:  artificial intelligence; knowledge bases; maternal and pediatric health; natural language processing; off-label use
    DOI:  https://doi.org/10.1093/jamia/ocag130
  22. Diagnostics (Basel). 2026 Sep 18. pii: 3028. [Epub ahead of print]16(18):
      Background/Objectives: Lumbar magnetic resonance imaging (MRI) is widely overused in low back pain. We assessed the appropriateness of lumbar MRI requests against three international guidelines and two independent large language model (LLM)-based artificial intelligence (AI) assessments, related appropriateness to diagnostic yield, and derived a simple decision rule to identify imaging likely to change management. Methods: In this single-centre retrospective study, 147 consecutive adults undergoing lumbar MRI for low back pain were analysed. Each request was classified using pre-imaging clinical variables according to operationalised ACR, NICE and ACP criteria and by two blinded LLM assessments (Claude, Anthropic; ChatGPT, OpenAI). The reference outcome was a management-changing MRI (subsequent surgery or interventional procedure); a four-item decision rule was derived by exhaustive search over clinical criteria optimising the Youden index. Results: Depending on the instrument, 57.8-72.1% of requests were inappropriate (ACR 66.0%; NICE 72.1%; ACP 66.0%; AI 57.8-64.6%). Agreement between instruments was substantial to near-perfect (κ 0.69-0.94; inter-AI κ 0.77). Red flags or neurological warning features were present in only 5.4%. MRI changed management in 33.3% overall-58.1-80.5% of appropriate versus 15.1-16.8% of inappropriate requests. A rule of four criteria-red flag or neurological warning feature, neurogenic claudication, purely radicular pain, or pre-imaging consideration of surgery/intervention-predicted management-changing MRI with 84% sensitivity and 81% specificity (Youden 0.64, versus 0.50-0.59 for the guidelines) and would have reduced imaging volume by 59.2%. Conclusions: Roughly two-thirds of lumbar MRI requests were guideline-inappropriate and rarely altered management. In this exploratory derivation cohort, a simple four-item rule showed promising predictive performance, but its apparent numerical advantage over existing guideline logic was not statistically confirmed and requires independent prospective validation; two blinded LLM assessments showed substantial agreement with guideline judgements, supporting further evaluation of LLM-based pre-order screening in independently validated settings.
    Keywords:  appropriateness; artificial intelligence; clinical decision rule; low back pain; lumbar spine; magnetic resonance imaging
    DOI:  https://doi.org/10.3390/diagnostics16183028