bims-arines Biomed News
on AI in evidence synthesis
Issue of 2026–09–20
twenty papers selected by
Farhad Shokraneh, Systematic Review Consultants LTD



  1. J Clin Epidemiol. 2026 Sep 18. pii: S0895-4356(26)00393-8. [Epub ahead of print] 112517
       BACKGROUND: The Epistemonikos Sustainable Knowledge Platform (SKP) integrates a suite of machine learning tools designed to make evidence synthesis more efficient, including access to the Epistemonikos Database of Trials (ED-Trials). One feature is the use of classifiers combined with large language models (LLMs) to support abstract screening during study selection. LLM-based classifiers can be applied to populations, interventions, and study design features. Their primary aim is to reduce the proportion of records needed to be screenedby automatically excluding records with a high probability of being ineligible. We aimed to evaluate SKP's LLM-based classifiers on abstract screenings of records retrieved via: (scenario 1) a traditional systematic literature search (original ), and (scenario 2) a search conducted within ED-Trials and subsequently forwarded to SKP.
    METHODS: We used data from a completed but at the time unpublished systematic review and network meta-analysis (NMA)as reference standards. We evaluated 29 LLM-based classifier combinations in both scenarios and assessed the proportion of records excluded by automated screening as well as whether any relevant study was wrongly excluded and therefore missed. To determine the impact of missed studies, we re-ran the network meta-analyses of the original review, excluding missed studies from the model and we evaluated how missed studies impact the certainty-of-evidence ratings.
    RESULTS: LLM-based classifiers excluded by automated screening 5.8-46.6% of 1,129 records identified by the original search and 7.7-34.3% of 1,003 records based on the SKP search. More specific classifiers excluded more records than broader ones. Of the 29 combinations applied to the original search results, 10 incorrectly excluded the same two relevant studies. Based on the SKP review, four combinations missed one study, and 10 combinations missed two studies. However, exclusion of these studies resulted in minimal changes to NMA-effect estimates (risk ratio changes of 0.0-0.04 in both scenarios) and did not alter certainty-of-evidence ratings.
    DISCUSSION: LLM-based classifiers are a promising strategy for reducing the amount of records to screen while keeping the risk of missing relevant studies low. However, further evaluations using multiple-use cases across different medical topics are necessary to learn more about the generalizability of our results.
    Keywords:  artificial intelligence; classifier; evidence synthesis; large language models; literature screening; study selection
    DOI:  https://doi.org/10.1016/j.jclinepi.2026.112517
  2. Clin Public Health Guidel. 2026 Jul;3(3): e70073
    GIN AI working group
       Introduction: Artificial intelligence (AI) may support several processes of the health guideline enterprise. This article describes the development of an extension of the Guidelines International Journal (GIN)-McMaster Guideline Development Checklist (GDC) for integrating AI in the guideline enterprise. This development has been led by the GIN-AI Working Group.
    Methods: We started by prompting a large language model (LLM) for items related to the use of AI in each of the steps of the original GDC. Subsequently, the members of the working group engaged in a set of iterative discussions, resulting in item refinement and in a consensus first version of the extension. We retrospectively applied this first version to a case use of guidelines incorporating AI in their development (Allergic Rhinitis and its Impact on Asthma [ARIA] 2024-2025 guidelines), leading to further refinement and to the approval of the final extension tool.
    Results: Prompting LLMs resulted in the generation of 149 items. Of those, 117 were removed and 19 were modified by members of the working group. On the other hand, 17 new items were added during the iterative discussion process. The retrospective application of the extension led to changes in the wording of four items. The final version of the checklist extension has been approved with 49 items modifying or adding to the original GDC.
    Discussion: We have developed an extension of the GIN-McMaster GDC that encompasses a set of conduct standards that are intended to facilitate the comprehensive and transparent integration of AI in the health guideline enterprise.
    Clinical Trial Registration: Not applicable. This study is not a clinical trial.
    DOI:  https://doi.org/10.1002/gin2.70073
  3. Clin Public Health Guidel. 2026 Jul;3(3): e70079
       Background: Artificial intelligence (AI) and automation offer opportunities to enhance the efficiency, timeliness, and sustainability of living guidelines (LGs). However, how AI and automation have been integrated into existing LG development frameworks remains unclear. This scoping review represents the first step in a broader programme of work aimed at developing a framework which aims to guide responsible and coordinated AI adoption across all phases of LG development.
    Objective: To identify existing frameworks, methods, and approaches that integrate AI or automation into any stage of LG development.
    Methods: We conducted a structured search of PubMed, Embase, Web of Science, Scopus, Cochrane Database for Systematic Review and Cochrane Central Register of Controlled Trials from inception to March 12, 2026. We included peer-reviewed articles describing frameworks, models, or methods that applied AI or automation in the development of LGs. We extracted data into a standardised form, capturing study characteristics, AI methods, targeted guideline processes, and key findings and synthesised findings using a thematic narrative approach.
    Results: Of 1090 records identified, three studies met the inclusion criteria. The included studies described AI or automation applied to select components of the LG process, to support continuous evidence surveillance, to semi-automate study screening, and incremental updating of living systematic reviews. None addressed multiple stages of LG development in an integrated manner.
    Conclusions: Current evidence demonstrates fragmented and narrowly focused AI or automation applications within LG development processes. Our findings highlight opportunities for future work to develop and evaluate more comprehensive frameworks that span multiple stages of the LG lifecycle.
    Keywords:  artificial intelligence; automation; frameworks; living guidelines
    DOI:  https://doi.org/10.1002/gin2.70079
  4. Clin Public Health Guidel. 2025 Oct;2(4): e70038
       Introduction: Artificial intelligence (AI) has the potential to support processes of the guideline enterprise. To support the adequate preplanning and the transparent reporting of the use of AI in the guideline enterprise, the Guidelines International Network (GIN) proposed the development of an extension of the GIN-McMaster Guideline Development Checklist (GDC) for the use of AI in guidelines. Here we describe the protocol for the development of this extension.
    Methods: We will follow a multiphase approach to generate an extension of the GIN-McMaster GDC. We will start by prompting a large language model to suggest relevant items for the extension. We will subsequently revise the obtained outputs in an iterative process involving the different elements of the working group. This iterative process will be informed by a scoping review. We will then reach different stakeholders to test the proposed checklist extension on their easiness of use, implementation, feasibility and reasonability.
    Questions: This protocol lays the ground for the development of an extension to the GIN-McMaster GDC for the use of AI in guidelines. This extension will support the adequate use of AI in the guideline enterprise.
    Keywords:  artificial intelligence; guidelines; protocol
    DOI:  https://doi.org/10.1002/gin2.70038
  5. Brief Bioinform. 2026 Sep 01. pii: bbag502. [Epub ahead of print]27(5):
      Extracting structured information from documents or scientific papers is crucial for data sharing and retrieval. Recent advances in large language models (LLMs) have demonstrated strong capabilities in language understanding, and a number of LLM-based tools have been developed for extraction-oriented tasks. However, it's still difficult to find a universal and user-friendly tool for various practical extraction tasks. To address this challenge, we propose OmniExtract, an automatic data extraction tool with user-friendly configuration files that can adapt to various data extraction tasks. OmniExtract employs a prompt optimization method to refine task-specific prompts and achieve high extraction performance. It also supports comprehensive data extraction from both documents and tables, making it applicable to a broad range of data sources. Evaluation results show that OmniExtract obtains a high accuracy ~90% for three datasets. Furthermore, two additional data extraction applications of OmniExtract in real-world scenarios have been presented, achieving an accuracy of 92.21% and ~90% precision and recall, respectively. Specifically, OmniExtract can handle tabular files of various sizes and formats, and achieve over 99% precision and recall on table information extraction tasks. The data reliability performance shows that OmniExtract is a valuable tool for database updating. An online testing service is available at https://ngdc.cncb.ac.cn/omniextract/. The service can be deployed locally with the code in https://github.com/wyb39/OmniExtract.
    Keywords:  data mining; information extraction
    DOI:  https://doi.org/10.1093/bib/bbag502
  6. JMIR AI. 2026 Sep 16. 5 e93761
       BACKGROUND: Despite high reported accuracy on clinical and evidence appraisal tasks, AI-generated medical information may lack explicit support from source documents. This creates challenges for digital health practitioners regarding transparency, auditability, and trust when AI systems are used for evidence synthesis, guideline development, and clinical knowledge management. Large language models (LLMs) can generate fluent and seemingly correct outputs, but existing evaluations often rely on agreement with human judgments and do not directly assess whether AI-generated content is grounded in underlying evidence.
    OBJECTIVE: This study measures evidence support and hallucination in AI-generated medical information by assessing the extent to which LLM-generated risk-of-bias assessments are supported by source clinical trial reports.
    METHODS: We evaluated 3 LLMs (GPT-5, OpenAI o3-mini, and GPT-3.5) on risk-of-bias (RoB 2) assessment using all 97 randomized controlled trials for which the source Cochrane systematic review provided complete human RoB 2 annotations and full-text reports were accessible, constituting the complete reference set. No train-validation split was applied; all 97 studies were used for evaluation. Model outputs were constrained to structured RoB 2 signaling questions and domain-level judgments. For each generated claim, relevant text passages were retrieved from trial reports using the Okapi BM25 (Best Matching 25) algorithm. A verification step assigned evidence verdicts (supported, contradicted, not found, or out of scope) with verbatim quotations. We quantified evidence support rates and conservative and strict hallucination rates. Task performance was evaluated using exact and binary accuracy, sensitivity, specificity, F1-score, Youden J, and agreement with human reviewers using Cohen κ and Fleiss κ.
    RESULTS: Binary accuracy of AI-generated risk-of-bias judgments was high across domains (90%-98%), whereas exact accuracy was substantially lower (42%-71%), reflecting frequent disagreements in severity classification despite correct directional classification. GPT-5 achieved the strongest overall performance, including perfect binary accuracy for overall risk-of-bias conclusions and the highest agreement with human reviewers (quadratic κ up to 0.81). However, evidence support rates across models ranged from only 60% to 65%, with conservative hallucination rates of 34%-37%. GPT-5 showed the highest mean evidence support (64.3%) and the lowest strict hallucination rate (35.7%). Mean top-1 BM25 retrieval scores were similar across models (approximately 30-31), suggesting that differences in hallucination were not primarily attributable to differences in retrieval strength.
    CONCLUSIONS: AI-generated medical information can achieve high decision-level accuracy while still lacking documentary support in a substantial proportion of outputs. Measuring evidence support and hallucination reveals important limitations that are not captured by agreement metrics alone. Retrieval-based evidence verification provides a reproducible and transparent approach for evaluating the reliability of AI-generated medical information, with direct relevance to digital health practice, evidence-based medicine, and medical informatics.
    Keywords:  AI; digital health; evidence-based medicine; hallucinations; information quality; information retrieval; large language models; medical informatics; risk of bias; systematic review
    DOI:  https://doi.org/10.2196/93761
  7. ESMO Real World Data Digit Oncol. 2026 Sep;13 100770
       Background: Oncologists face increasing difficulty staying current with rapidly evolving clinical data, guidelines, and regulatory updates. Building and maintaining an annotated clinical trial evidence library is time- and labor-intensive. To address this challenge, we developed and validated a living oncology evidence platform (Living-OEP) for breast cancer (BC), powered by an agentic artificial intelligence (AI) system that supports daily, human-conducted, AI-augmented systematic literature review (SLR).
    Materials and methods: The agentic AI system, incorporating GPT-4.1 and o3 (OpenAI) and Claude [Anthropic, PBC, San Francisco, CA] Sonnet-4 (Anthropic), was designed to emulate expert-led, Cochrane-compliant SLR workflows. Guided by a human-developed annotation manual, the system decomposes tasks, self-debugs, and validates outputs. Training data included 29 236 clinical trial abstracts across BC, lung, and prostate cancer, each annotated with four review and 32 extraction variables. Structured data were integrated with guideline-based treatment pathways, forming a real-time, evidence-linked OEP. Accuracy was benchmarked against 1997 human annotations. Living-OEP evidence quality was compared with AI chatbots (ChatGPT [OpenAI, San Francisco, CA], Perplexity [Perplexity AI, Inc., San Francisco, CA], Consensus [Consensus, Boston, MA], and OpenEvidence [OpenEvidence Inc., Miami, FL]) using six criteria across eight BC treatment scenarios.
    Results: The agentic AI review accuracy ranged from 95.1% to 97.2%. Extraction accuracy ranged from 51.1% to 99.4%, with three variables still undergoing iterative refinement. Compared with other AI tools, the Living-OEP provided more comprehensive and accurate evidence and linked all data to original publications and Food and Drug Administration labels.
    Conclusions: A human-conducted, AI-augmented living SLR integrated with guidelines and regulatory data can provide real-time evidence support. Future studies are required to evaluate its impact on physician workflows, clinical decision-making, and implementation in oncology practice.
    Keywords:  agentic artificial intelligence; breast cancer; clinical trial annotation; evidence synthesis; systematic literature review
    DOI:  https://doi.org/10.1016/j.esmorw.2026.100770
  8. Sci Prog. 2026 Jul-Sep;109(3):109(3): 368504261489884
      The exponential growth of biomedical literature poses significant challenges for knowledge management in specialized domains such as osteoporosis. While Large Language Models (LLMs) offer advanced natural language understanding capabilities, their direct application in high-risk medical contexts is hindered by hallucination risk, static knowledge, and limited traceability. This paper proposes a domain-specific Retrieval-Augmented Generation (RAG) system tailored for osteoporosis knowledge management, integrating authoritative sources from the International Osteoporosis Foundation and China's National Health Commission. The implemented system uses FAISS indexing with all-MiniLM-L6-v2 sentence embeddings and compares three retrieval strategies--Classic RAG, a fixed-budget non-agentic Multi-Query RAG (MQ-RAG), and Agentic RAG--against an LLM-only baseline. We evaluate the four conditions using a curated test set of 210 questions across true/false, single-choice, and open-ended formats. Results from three representative LLMs--Deepseek-v3.2, ChatGPT-5, and Qwen3-max--are reported for objective accuracy and expert-rated answer quality. Inter-rater agreement for the open-ended expert evaluation was substantial (weighted Fleiss's kappa = 0.72). This work highlights the potential of agentic RAG architectures for specialized medical question-answering systems.
    Keywords:  large language model (LLM); osteoporosis; retrieval-augmented generation (RAG)
    DOI:  https://doi.org/10.1177/00368504261489884
  9. iScience. 2026 Sep 18. 29(9): 117438
      The rapid expansion of biomedical literature demands automated summarization tools that reliably condense research articles into concise, accurate summaries. We benchmarked 62 summarization methods, ranging from frequency-based and TextRank extractors to encoder-decoder models (EDMs) and large language models (LLMs), on 1,000 biomedical abstracts from 20 journals across ScienceDirect and Cell Press, using author-written highlights as reference summaries. Models were evaluated with a composite suite of lexical, semantic, and factual metrics, including ROUGE, BLEU, METEOR, embedding-based similarity, and factuality scores. General-purpose models (e.g., Mistral, GPT, and Llama) achieved the highest overall performance across lexical and semantic dimensions, outperforming reasoning-oriented (e.g., DeepSeek and Magistral) and domain-specific (e.g., BioGPT and BioMistral) models. Notably, medium-sized models outperformed large-scale models, suggesting an optimal balance between model capacity and efficiency, while classical extractive methods lagged behind neural approaches. These findings provide a systematic reference for selecting biomedical summarization tools and highlight that broad pretraining outperforms narrow domain adaptation.
    Keywords:  benchmarking; biomedical text summarization; large language models; natural language processing
    DOI:  https://doi.org/10.1016/j.isci.2026.117438
  10. Risk Anal. 2026 Oct;46(10): e70356
      How should the credibility of claims about the human health effects of exposures be evaluated when their support consists of multiple individually inconclusive lines of evidence? Health risk analysis typically draws on heterogeneous evidence streams-including observational epidemiology, systematic reviews and meta-analyses, mechanistic and toxicological evidence, and policy-facing syntheses-to support causal conclusions and risk management recommendations. However, important distinctions among observed associations, causal interpretations, mechanistic plausibility, and claims about intervention or policy effectiveness are often blurred in scientific reviews and evidence syntheses, making evaluations difficult to reproduce, audit, or systematically critique. This article introduces ASP1 (Automated Scientific Paper reviewer 1), available online at https://moirai.shinyapps.io/asp1_pdf_runner/, an AI-assisted framework for systematic, transparent evaluation of heterogeneous scientific evidence and of the conclusions drawn from it. ASP1 combines modular review components, structured claim extraction, explicit linkage between evaluative judgments and supporting text, and cross-module synthesis to expose and systematically examine the evidentiary and inferential chain connecting a document's evidence to its conclusions. We illustrate it using a mini-corpus of recent publications concerning ultra-processed foods (UPFs) and health spanning prospective cohort studies, case-control studies, umbrella reviews, mechanistic syntheses, and policy-facing evidence reviews. Results can be browsed at https://moirai.shinyapps.io/asp1_bundle_browser/. Across these diverse study designs, ASP1 consistently distinguished evidence of association from stronger causal and interventional claims and identified important limitations involving residual confounding, exposure misclassification, mechanistic uncertainty, category heterogeneity, and overextension of policy conclusions beyond direct evidentiary support. Comparison with independent expert human commentary on a recent Lancet review showed substantial overlap in major inferential concerns identified by ASP1 and by skilled human reviewers, while also revealing complementary strengths. ASP1 was generally stronger in explicit inferential decomposition, claim-specific support classification, consistency, and auditability; human experts contributed richer contextual understanding, sharper methodological intuition, and more vivid and concrete examples of evidentiary limitations. These results suggest that current AI systems guided by structured evaluative frameworks can support more transparent, reproducible, and disciplined review of heterogeneous scientific evidence at realistic scales. The greatest promise of such systems may lie less in automating scientific judgment than in helping make scientific reasoning more explicit, inspectable, auditable, and open to challenge.
    Keywords:  artificial intelligence; epidemiology; evidence synthesis; health risk assessment; large language models; ultra‐processed foods
    DOI:  https://doi.org/10.1111/risa.70356
  11. Healthcare (Basel). 2026 Sep 01. pii: 2767. [Epub ahead of print]14(17):
      Background/Objectives: Emergency department (ED) overcrowding contributes to delayed care, prolonged length of stay (LOS), resource strain, and adverse patient outcomes. This systematic review aimed to examine how artificial intelligence (AI) and machine learning (ML) have been used to address ED crowding and patient flow, with emphasis on modeling approaches, validation practices, and real-world implementation. Methods: Following PRISMA 2020 guidelines, Scopus, Embase, Ovid MEDLINE, and CENTRAL were searched for relevant studies published from 2020 onward. After deduplication, 1888 records underwent title and abstract screening using two locally deployed LLaMA models with human adjudication. Screening performance was assessed against 150 manually annotated records. Full-text eligibility assessment and structured data extraction were conducted independently by multiple reviewers, with disagreements resolved by consensus. Results: Thirty-two studies were included. Most were retrospective, single-site investigations using electronic health record, administrative, or operational data. Common outcomes included ED LOS, waiting time, occupancy, boarding, disposition, and crowding indices. Tree-based and boosting models frequently performed well, although no approach was consistently superior across tasks and settings. Most studies relied on same-site validation, while external and temporal validation were uncommon. Prospective implementation, workflow integration, model maintenance, and direct operational, clinical, economic, or equity impacts were rarely evaluated. For LLM-assisted screening, LLaMA 4 Scout achieved 84.0% accuracy, 80.0% recall, 88.9% precision, and an F1 score of 84.2%, compared with 78.0%, 67.5%, 88.5%, and 76.6%, respectively, for LLaMA 3.3 on 150 randomly sampled papers. Conclusions: AI and ML show promise for addressing ED overcrowding, but the literature remains concentrated at the model-development stage. Future research should prioritize standardized outcomes, multicenter validation, prospective implementation, and direct evaluation of operational and patient-care outcomes.
    Keywords:  artificial intelligence; emergency department overcrowding; large language models; length of stay; machine learning; patient flow; predictive modeling; systematic review
    DOI:  https://doi.org/10.3390/healthcare14172767
  12. Am Heart J Plus. 2026 Oct;70 100877
       Background: Large language models (LLMs) are increasingly used to support statistical analyses in biomedical research. However, their ability to accurately and reproducibly compute descriptive statistics directly from datasets has received limited evaluation.
    Objective: To compare the accuracy, within-modality repeatability, and between-modality consistency of ChatGPT and Claude in generating descriptive statistics from structured datasets provided through different input modalities.
    Methods: Two publicly available Stata datasets were evaluated: 'auto.dta' (74 observations, 12 variables) and 'citytemp.dta' (956 observations, 6 variables). Original variable names were replaced with generic labels. Each dataset was analyzed using three input modalities (copy-paste, Word, and Excel), with two independent repetitions per modality. Stata served as the reference standard. Accuracy was assessed for counts of missing and non-missing observations, minima, maxima, means, standard deviations, medians, quartiles, frequencies, and percentages.
    Results: ChatGPT and Claude produced identical results across all analyses. For each model, 624 categories of descriptive statistics were evaluated. Exact agreement with the Stata reference standard was observed for 576 of 624 categories (92.3%). The only deviations involved first and third quartiles (Q1-Q3) for eight variables. Post hoc analyses demonstrated that these differences were entirely attributable to the use of a different, but mathematically valid, quartile definition based on linear interpolation rather than computational errors. Within-modality repeatability and between-modality consistency were complete across the evaluated analyses.
    Conclusions: ChatGPT and Claude demonstrated excellent accuracy and consistent results across repeated analyses and input modalities for the datasets and descriptive statistics evaluated. After accounting for differences in quartile definitions, no calculation errors were identified across the 624 evaluated categories per model.
    Keywords:  AI; Artificial intelligence; ChatGPT; Claude; Descriptive statistic; LLM; Large language model; Research; Statistical analysis
    DOI:  https://doi.org/10.1016/j.ahjo.2026.100877
  13. J Am Board Fam Med. 2026 Sep;pii: 168435. [Epub ahead of print]39(1):
      
    Keywords:  Artificial Intelligence; Clinical Decision Support; Clinical Decision-Making; Evidence-Based Medicine; Family Medicine; Large Language Models; Natural Language Processing
    DOI:  https://doi.org/10.3122/jabfm.2026.260341R0
  14. Pharmacoeconomics. 2026 Sep 17.
       BACKGROUND: Conventional evidence synthesis in systematic literature reviews (SLRs) of economic evaluations typically provides a narrative description of modelling approaches used across studies and rarely examines the underlying rationale for these choices or their interpretation. This study develops and applies a thematic synthesis approach to analyse modelling decisions reported in economic evaluations.
    OBJECTIVE: The aim of this study was to introduce and apply an Approach-Rationale-Note (ARN) framework for thematic evidence synthesis in economic evaluation SLRs and demonstrate its application using economic evaluations of treatments for spinal muscular atrophy (SMA).
    METHODS: Evidence from a previously conducted SLR of economic evaluations of SMA treatments was synthesised using a deductive thematic approach. The ARN framework structured the analysis by distinguishing modelling approaches, their rationales, and contextual observations relevant to interpretation. Analytical domains and dimensions were developed iteratively to enable comparison of modelling decisions across studies, with their selection guided by the review question and the methodological issues identified in the evidence base.
    RESULTS: Substantial variation in modelling approaches was observed across key ARN domains, including model structure, health state definitions, treatment effect assumptions, extrapolation methods, and the handling of uncertainty. While several studies adopted similar modelling approaches, the rationales provided for these choices varied considerably. Health technology assessment (HTA) sources frequently highlighted uncertainty related to long-term outcomes, treatment durability, and survival extrapolation.
    CONCLUSIONS: A structured thematic synthesis of modelling decisions enables comparison not only of the approaches used in economic evaluations but also of the underlying reasoning behind those choices. The ARN framework provides a transparent and complementary approach for examining methodological decision making in economic modelling studies and may support improved interpretation of modelling assumptions in HTA and related evidence syntheses.
    DOI:  https://doi.org/10.1007/s40273-026-01653-w
  15. Stud Health Technol Inform. 2026 Sep 17. 340 234-241
      Statistical analysis of clinical data requires expertise in medical statistics. Large language models (LLMs) are increasingly used for code generation and may support both descriptive and advanced analyses, but their reliability remains uncertain. This study evaluated whether five current LLMs (GPT 5.3, Claude Sonnet 4.6, Gemini 2.5 Flash, Perplexity, and Grok) could reproduce expert-validated statistical analyses from a published clinical workflow across three statistical tasks: descriptive table generation, Kaplan-Meier survival analysis, and Cox proportional hazards modelling. All models received identical datasets and standardised prompts. Their outputs were compared with analyses performed by two experts trained in mathematical statistics and evaluated for grouping correctness, numerical accuracy, missing-value handling, and quality of generated R code. All five models produced correct descriptive statistics once dataset variables were specified explicitly. Two models failed the initial descriptive benchmark because of variable-name ambiguity, but both recovered after prompt clarification. All models reproduced the correct Kaplan-Meier p-value, although figure completeness differed. In the Cox benchmark, four models reproduced all required hazard-ratio terms, while one omitted the interaction results. These findings suggest that LLMs may support clinical research workflows, but their outputs still require careful validation before use.
    Keywords:  Artificial Intelligence; Hypertension; Large language models; Proportional Hazards Models; Pulmonary; Reproducibility of Results; Survival Analysis
    DOI:  https://doi.org/10.3233/SHTI261007
  16. Patterns (N Y). 2026 Sep 11. 7(9): 101644
      Large language models (LLMs) increasingly write code, analyze data, and orchestrate scientific workflows. This can create reproducibility challenges because LLMs blur the boundary between how an analysis is built and what it depends on at run time. Guidelines for LLM-assisted science begin with a critical choice: whether the LLM sits on the data path of the published analysis, a live step results depend on, or off the data path, producing durable artifacts such as code. Reproducible analyses require preserving data, code, and runtime; an on-path LLM becomes part of the runtime, a dependency that may change, be deprecated, or become inaccessible. We derive six recommendations: (1) keep the LLM off the data path where possible, (2) preserve LLM-generated artifacts, (3) verify results by methods suited to the LLM's role, (4) consider open-weight models, (5) record the model version, and (6) assess determinism.
    Keywords:  AI in science; data analysis; determinism; large language models; open-weight models; provenance; reporting standards; reproducibility; research software; scientific workflows
    DOI:  https://doi.org/10.1016/j.patter.2026.101644
  17. MethodsX. 2026 Dec;17 104113
      This article describes the protocol used to conduct a systematic mapping study (SMS) of software architecture. The presence of architectural and design instruction in primary studies cannot be determined through search and superficial review alone. A principled quality assessment must be conducted to identify studies that are both relevant and of high quality. LLM assistance offers practical means of performing such assessments in a timely and efficient manner. Therefore, our contribution is threefold: • Guidance for the following SMS steps: tooling setup, database search, deduplication and snowballing, primary study selection, quality assessment, and data extraction. • The rationale and methodology to perform quality assessment on primary studies regarding software architecture. • A method of using large language models (LLMs) together with human evaluators is presented. The method is supported by bespoke open-source software.
    Keywords:  Entity resolution; Large language models; Method; Quality assessment; Software architecture; Systematic mapping study
    DOI:  https://doi.org/10.1016/j.mex.2026.104113
  18. Nature. 2026 Sep 16.
      
    Keywords:  Computer science; Machine learning; Publishing
    DOI:  https://doi.org/10.1038/d41586-026-02899-2