bims-arines Biomed News
on AI in evidence synthesis
Issue of 2026–09–13
sixteen papers selected by
Farhad Shokraneh, Systematic Review Consultants LTD



  1. J Eval Clin Pract. 2026 Sep;32(6): e70596
       OBJECTIVE: To evaluate the performance and consistency of Large Language Models (LLMs) in core systematic review (SR) tasks and to introduce open-source tools for automated batch processing that provide decision rationales.
    METHODS: We assessed GPT-4o, Kimi-K2, DeepSeek-V3, and DeepSeek-R1 on five SR tasks: title/abstract screening (3550 records), full-text screening (233 texts), data extraction (112 RCTs), Risk of Bias (ROB) assessment (112 RCTs), and AMSTAR-2 assessment (20 SRs). Each model was evaluated twice to measure consistency. All outputs required supporting rationales and verbatim evidence.
    RESULTS: LLMs demonstrated proficiency across tasks, with generally high intra-model but lower inter-model consistency. In screening, models showed lower precision (0.27-0.40) but high recall (0.83-0.91) and specificity (0.83-0.91). DeepSeek-R1 and DeepSeek-V3 excelled in title/abstract and full-text screening, respectively. Data extraction accuracy was similar across models (0.78-0.82). Kimi-K2 achieved the highest ROB F1 score (0.71). AMSTAR-2 assessments were generally acceptable.
    DISCUSSION: While effective, LLMs showed variable performance across SR tasks. The mandatory output of rationales and evidence enhances transparency and allows for human verification of AI decisions.
    CONCLUSION: We provide a suite of automated tools for key SR tasks. By leveraging these tools to validate model outputs rather than starting manually, reviewers can significantly improve workflow efficiency while maintaining methodological rigour.
    Keywords:  automation; evidence‐based medicine; large language models; risk of bias; systematic reviews
    DOI:  https://doi.org/10.1111/jep.70596
  2. BMJ Digit Health Ai. 2025 ;1(1): e000017
       Objective: To evaluate and synthesise current applications of large language models (LLMs) in systematic reviews and meta-analyses (SRMAs), identify key limitations and propose an enhanced theoretical framework to improve the efficiency, scalability and reliability of evidence synthesis.
    Methods and analysis: We conducted a narrative review of recent studies applying LLMs across key SRMA stages. A total of 21 publications were analysed for model type, task application, accuracy metrics and workflow impact. Building on this evidence base, we designed a comprehensive LLM-enhanced SRMA framework that categorises LLM roles as consultants and assistants, integrates human-in-the-loop strategies and uses retrieval-augmented generation (RAG) and agent-based architectures to address critical challenges including hallucinations, bias and workflow inefficiency.
    Results: The reviewed literature demonstrated that LLMs can support various SRMA tasks with reported accuracy ranging from 61% to 99%, showing particular promise in literature screening and data extraction. Our proposed framework conceptualises modular integration of LLMs across all six SRMA stages, with LLMs serving as consultants for research question formulation and search strategy development and as assistants for task automation including abstract screening and structured data extraction. The framework incorporates RAG technology to reduce hallucinations by grounding outputs in retrieved literature and employs agent-based orchestration for complex analytical workflows. Theoretical analysis suggests potential for significant efficiency gains while maintaining methodological rigour through strategic human oversight.
    Conclusion: LLMs offer substantial theoretical potential to transform evidence synthesis by improving efficiency, scalability and consistency across SRMA workflows. The proposed LLM-enhanced framework provides a systematic, theoretically grounded approach for integrating advanced artificial intelligence capabilities into existing SRMA methodologies while preserving essential human oversight and analytical integrity. Future empirical studies are needed to validate the framework's practical effectiveness, establish implementation protocols and demonstrate real-world benefits in evidence-based medicine.
    Keywords:  BMJ Health Informatics; Global Health; Medical Informatics; Public Health
    DOI:  https://doi.org/10.1136/bmjdhai-2025-000017
  3. PLOS Digit Health. 2026 Sep;5(9): e0001666
      The scale and pace of evidence generation in digital psychiatry increasingly exceed the capacity of traditional systematic literature review (SLR) methods. Large language models (LLMs) are gaining traction in evidence synthesis, yet limited guidance exists for integrating generative AI into transparent, reproducible SLR workflows. To develop and evaluate a prompt-driven, modular decision-making framework for adaptable SLR workflows in digital psychiatry, and to compare stage-matched performance relative to a consensus-based human reference process. We conducted a mixed-methods, stage-matched comparative evaluation of three GPT-5 mode variants (Auto, Agent, and Deep Research) against a consensus-based human-led SLR workflow using a registered digital psychiatry review (PROSPERO CRD42025648122) as a case example. Nine SLR tasks were implemented within the ChatGPT interface using structured RISEN prompts: preliminary searches; research question, eligibility criteria, and search strategy development; article screening; data extraction; article summarization; critical appraisals; and descriptive results synthesis. Performance was evaluated using task-specific assessments of accuracy, sensitivity, specificity, inter-rater agreement, reporting compliance, hallucination monitoring, and workflow feasibility. All GPT-5 mode variants were feasible across SLR stages, although performance varied by task and mode. No fabricated study-level content was identified during structured hallucination auditing. Auto mode performed best for structured rule-based tasks requiring efficiency and implementation-ready outputs (e.g., eligibility criteria). Agent mode excelled in conceptual integration and interpretive tasks (e.g., research question generation, critical appraisals, descriptive synthesis). Deep Research mode most closely approximated human reasoning for higher-order synthesis tasks, performing best in full-text article screening, article summarization, and narrative synthesis. Prompt-driven LLM workflows are feasible and semi-efficacious for selected SLR tasks in digital psychiatry when deployed within a modular, human-in-the-loop decision-support framework. Findings support task-specific deployment of an adaptable modular framework, enabling healthcare researchers and clinicians to modify LLM workflows according to methodological appropriateness, confidence in performance, and institutional review practices.
    DOI:  https://doi.org/10.1371/journal.pdig.0001666
  4. Clin Transl Allergy. 2026 Sep;16(9): e70202
       OBJECTIVE: To demonstrate an evaluation of the use of artificial intelligence (AI) in supporting tasks in evidence synthesis (particularly search for primary studies) in the allergy or respirology fields.
    METHODS: We queried three AI-based platforms specialised on identifying scientific publications using several strategies (i.e., prompting strategies and search approaches) in November 2024 to identify primary studies related to health-related case studies in the allergy or respirology field. We compared how many eligible primary studies we identified using AI-based platforms versus in a systematic search in multiple electronic bibliographic databases. We also compared meta-analytical results obtained with the primary studies identified by querying AI-based platforms versus in the context of a systematic review. Finally, we developed a structured methodological framework and a reporting checklist for using these AI-based platforms.
    RESULTS: In our main case study, a systematic review of randomised controlled trials (RCTs) informing the Allergic Rhinitis and its Impact on Asthma (ARIA) guidelines, a strategy involving searching multiple AI-based platforms identified 85.7% of all full articles with DOI, but failed to identify unpublished trials registered in trial databases, resulting in an overall identification of 56.3% eligible RCTs. Meta-analytical estimates were similar when considering only the primary studies identified using AI-based strategies versus those identified by the systematic review. We observed a lower performance in the case study of systematic reviews of observational studies.
    CONCLUSIONS: We provide an example on how it is possible to evaluate AI-based platforms in the support of evidence synthesis, with a focus on the allergy and respirology fields.
    Keywords:  allergic rhinitis; artificial intelligence; evidence search; rapid evidence reviews; systematic reviews
    DOI:  https://doi.org/10.1002/clt2.70202
  5. Pediatr Radiol. 2026 Sep 10.
       BACKGROUND: Deduplication is usually treated as a routine preprocessing step in systematic reviews, yet errors introduced at this stage are not visible in Preferred Reporting Items for Systematic reviews and Meta-Analyses (PRISMA) flow diagrams and can propagate through evidence synthesis.
    OBJECTIVE: To evaluate the validity and reproducibility of commonly used deduplication tools in a paediatric neuroradiology systematic review.
    MATERIALS AND METHODS: PubMed and Embase were searched for peer-reviewed literature on paediatric subacute sclerosing panencephalitis neuroimaging, yielding 603 records. Two reviewers independently deduplicated the combined corpus using Zotero (version 8.0.4) and Rayyan (updated 21 January 2026). A manual reference standard was established by dual-reviewer adjudication of potential duplicate pairs using title, publication year, author list, and Digital Object Identifier (DOI). Tool outputs were compared against this reference standard for false-positive and false-negative decisions and for reproducibility across runs.
    RESULTS: Manual adjudication identified 120 true duplicate records, leaving 483 unique records. Zotero produced identical outputs across reviewers, indicating deterministic behaviour, but generated 16 false-positive duplicate decisions and 5 false negatives, predominantly related to DOI collision. Rayyan identified more true duplicates but produced non-identical outputs across independent runs, with 16-17 false positives and 1 false negative per run.
    CONCLUSION: Commonly used deduplication tools can introduce either deterministic bias or non-deterministic record loss before screening begins. In paediatric radiology evidence synthesis, the deduplication software version and verification strategy should be explicitly reported. Auditable secondary checks should be considered before records are removed.
    Keywords:  Deduplication; Evidence synthesis; Neuroradiology; Paediatric radiology; Reproducibility; Subacute sclerosing panencephalitis
    DOI:  https://doi.org/10.1007/s00247-026-06775-z
  6. Campbell Syst Rev. 2026 Sep;22(3): 18911803261484937
      Systematic reviews and meta-analyses aim to comprehensively identify, summarize and appraise literature. Record identification is performed within multiple databases with important overlap. Deduplication is therefore an important, time-consuming step often underreported or described only briefly. Unique identifiers like the digital object identifier (DOI®) provide a practical basis to identify and remove unambiguous duplicates, but require careful handling of missing, inconsistently formatted, or non-unique DOIs. We created a simple tool that removes duplicates based on normalized DOIs and titles, and make it available as a free-to-use web application hosted on deduplicate.it. It is designed for ease of use, while maintaining high transparency by utilizing a human-readable, simple algorithm and outputting exclusion files and a PRISMA-style flowchart alongside the deduplicated output. In a validation set of five reviews in medicine, education and social welfare, deduplicate.it was able to identify around 80% of duplicates and reduce the time dedicated to deduplication accordingly.
    Keywords:  deduplication; duplicate removal; literature screening; reference screening; systematic review
    DOI:  https://doi.org/10.1177/18911803261484937
  7. Water Res. 2026 Aug 26. pii: S0043-1354(26)01487-9. [Epub ahead of print]308(Pt A): 126813
      The integration of artificial intelligence (AI) with sustainable water management holds transformative potential for addressing global sustainability challenges. However, this progress is critically hindered by the slow, labor-intensive construction of large-scale datasets, particularly in identifying the relevant literature from thousands of candidates. Here, we present a hierarchical AI framework that ensures both high accuracy and transparency in literature screening. First, we develop a domain-tailored prompting strategy (3T+RAG) that grounds large language models (LLMs) in structured water-treatment knowledge, thereby improving classification accuracy and screening reliability. Evaluated on three manually curated domain datasets, LLMs achieved reliable literature screening with an average F1-score of 0.88, while operating 79 times faster at only 1.4% of the estimated cost of individual human annotation. To further enhance​ transparency and auditability, we introduce a multi-agent Reviewer-Reviewer-Arbiter (RRA) framework, in which two reviewer agents independently assess literature using the 3T+RAG prompting strategy, and an arbiter agent resolves disputes through reasoning-trace analysis. This architecture distinguishes reviewer disagreements from consistent decisions, allowing human review to focus on disputed articles and targeted quality control. Overall, this study establishes an efficient and reliable pathway for water-treatment literature screening, providing a robust foundation for large-scale dataset construction and accelerating sustainable water management.
    Keywords:  AI for science; Generative AI; Large language models; Literature screening; Sustainable water management; Water purification
    DOI:  https://doi.org/10.1016/j.watres.2026.126813
  8. F1000Res. 2025 ;pii: Chem Inf Sci-260. [Epub ahead of print]14
      Effective research depends on building on the knowledge found in the scientific literature. Designed to streamline literature tasks, the EPA's Abstract Sifter literature tool, now at version 8, has been continually extended and enhanced since its introduction in 2017[1]. Early enhancements to the tool have primarily focused on core tasks common to all researchers. For example, citation retrieval from PubMed has been made faster and the returned citation threshold increased to 10,000. Features that allow deeper examination of the literature have been introduced as well. A functionality called Term-mapping allows for fast, dynamic relevancy ranking of returned citations. MeSH substances, such as proteins, genes, and chemicals, can now be extracted from a retrieved corpus of citations, ranked by frequency and explored through the MeSHMine functionality. Features that facilitate user engagement with publications have also been improved: formatting and colorization ease reviewing of the abstract text and the tagging and noting citations functionality has been streamlined. Version 8 introduced multiple features that break new ground in working with chemical literature. For example, chemical entity extraction from scientific publications has been streamlined through download of PDFs and automated table extraction. Following entity extraction, the chemical names can be used as inputs to retrieve EPA's chemical identifiers, the DSSTox (Distributed Structure-Searchable Toxicity) chemical IDs (DTXSIDs). Once these identifiers have been retrieved, a wealth of chemical information is available through built-in functions accessing EPA's Computational Toxicology and Exposure application programming interface (CTX-APIs) [2]. This new functionality allows researchers to build on the EPA's efforts in chemical data assembly and curation. The Abstract Sifter version 8 is a valuable tool for researchers endeavoring to understand chemicals and their effects on the environment and biological systems.
    Keywords:  DSSTox; Literature mining; PubMed; drug discovery; knowledge mining; toxicology
    DOI:  https://doi.org/10.12688/f1000research.160617.2
  9. BMJ Digit Health Ai. 2026 ;2(1): e000034
       Objectives: This study aims to compare the reliability and accuracy of three large language models (LLMs) (Claude, Gemini and GPT) in assessing the risk of bias of nonrandomised studies using the ROBINS-I tool.
    Methods and analysis: We conducted a secondary analysis of 171 nonrandomised studies previously assessed with Risk Of Bias In Non-randomized Studies of Interventions (ROBINS-I) tool by two independent human review teams. Only studies with concordant human domain-level ratings were included. Each study was independently assessed twice by Claude, Gemini and Generative Pre-trained Transformer (GPT) using agent-based structured implementations of the ROBINS-I tool. Reliability (agreement between two runs of the same LLM) was evaluated using percent agreement and Gwet's AC1. Accuracy (agreement with human reviewers) was assessed only for studies with consistent LLM ratings, using the same metrics.
    Results: Claude demonstrated high reliability across all domains (79.5-98.0% agreement, AC1=0.729-0.975). Gemini showed moderate-to-high reliability (agreement 76.7-100%, AC1=0.680-1.0). GPT exhibited lower reliability overall, though domain-level agreement ranged from 70.9-95.6% (AC1=0.596-0.944). In terms of accuracy, Claude showed overall poor agreement with human reviewers (14.4-68.5% agreement; low AC1 values). Gemini demonstrated moderate-to-high accuracy in several domains, including deviations from intended interventions (79.6%, AC1=0.848) and measurement of outcomes (73.9%, AC1=0.702), with the highest overall agreement (40.0%, AC1=0.672). GPT showed variable accuracy, with the highest in measurement of outcomes (62.8%, AC1=0.571) and classification of interventions (57.8%, AC1=0.498), but poor performance in selection (14.3%, AC1 = -0.041) and overall agreement (23.0%, AC1=0.267).
    Conclusions: Claude was internally consistent but poorly aligned with human reviewers. Gemini achieved both high reliability and moderate-to-high accuracy, whereas GPT had lower reliability and mixed accuracy. Current off-the-shelf LLMs cannot reliably perform ROBINS-I risk of bias assessments.
    Keywords:  Artificial intelligence; Evidence-Based Medicine
    DOI:  https://doi.org/10.1136/bmjdh-2026-000034
  10. J Eval Clin Pract. 2026 Sep;32(6): e70593
       RATIONALE: The rapid expansion of medical literature has led to variability and contradictions in study findings, making it increasingly difficult to distinguish meaningful signals from noise. Much of this variability arises from methodological limitations, including confounding, selection bias, and reverse causation. Although artificial intelligence (AI)-assisted tools exist for risk-of-bias assessment, most are designed for systematic reviews and are not tailored to identifying epidemiologic biases in observational studies. Structured, scalable approaches are needed to evaluate validity in real-world evidence research.
    AIMS AND OBJECTIVES: To develop and validate EpiVise, an AI-assisted, expert-informed, rule-based framework for identifying major sources of bias in pharmacoepidemiologic studies and to assess its agreement with expert epidemiologist evaluations.
    METHODS: Recently published pharmacoepidemiologic studies from high-impact journals (post- July 2025) were independently evaluated by EpiVise and two expert epidemiologists across predefined bias domains, including measured confounding, confounding by indication, selection bias, immortal time bias, and disease latency bias. Agreement was assessed using weighted kappa statistics. In addition, synthetic study scenarios with predefined embedded biases were constructed to evaluate framework performance under controlled conditions.
    RESULTS: Among published studies (10 studies; 60 ratings), agreement between EpiVise and expert assessments was substantial (weighted κ = 0.75; 95% confidence interval [CI], 0.63-0.87). Twelve ratings (20.0%) were discordant, all limited to adjacent categories. In synthetic scenarios (10 studies; 50 ratings), agreement was also substantial, with 40 of 50 ratings concordant (80.0%) and a weighted κ of 0.72 (95% CI, 0.61-0.83).
    CONCLUSION: EpiVise demonstrated substantial agreement with expert epidemiologist assessments in both published and synthetic study evaluations. As a scalable and reproducible framework for identifying common epidemiologic biases, EpiVise may enhance evidence appraisal, peer review, and clinical or regulatory decision-making. Further validation across broader study designs and therapeutic areas is warranted.
    Keywords:  AI platform; AI validation; bias assessment; observational studies
    DOI:  https://doi.org/10.1111/jep.70593
  11. Innovation (Camb). 2026 Sep 08. 7(9): 101391
      Extracting quantitative data from the growing body of scientific literature is a challenge central to modern research across disciplines. While recent advances in large language models have significantly facilitated automation of this traditionally time-consuming task, their computational demands limit scalability and accessibility. Smaller specialized systems offer reduced computational requirements but sacrifice accuracy, domain generalization, or the extraction of contextual details. To address these gaps, we present Quinex, a domain-agnostic framework for quantitative information extraction based on comparably small language models. Quinex identifies quantities and their associated entities, properties, and qualifiers using a multi-turn question-answering approach. By addressing the data bottleneck, considering implicit properties, and optimizing extraction order and question templates, Quinex achieves state-of-the-art F1 scores of over 98% for quantities, 82% for entities, and 87% for properties across diverse scientific genres. It goes beyond identification by normalizing units, aligning them with a unit ontology, and detailing critical qualifiers such as references, spatiotemporal scopes, and determination methods. By reducing the effort required to extract structured quantitative data from texts, Quinex enables transformative applications, including automated literature reviews, quantitative search, and trend monitoring, and sets a new benchmark for scalable, accurate, and domain-agnostic quantitative information extraction.
    Keywords:  automated literature review; quantitative information extraction; scientific NLP; scientific literature mining; small language models; structured data extraction
    DOI:  https://doi.org/10.1016/j.xinn.2026.101391
  12. Asian J Psychiatr. 2026 Sep 08. pii: S1876-2018(26)00327-8. [Epub ahead of print]125 105154
      
    Keywords:  Artificial intelligence; ChatGPT; Qualitative analysis; Thematic analysis
    DOI:  https://doi.org/10.1016/j.ajp.2026.105154
  13. JAMIA Open. 2026 Oct;9(5): ooag091
       Objective: This methodological validation study evaluated a large language model (LLM), multi-step framework for thematic analysis (TA) of healthcare interviews, compared to a human-only coded reference standard.
    Materials and Methods: The unit of analysis was deidentified transcripts of interviews with geriatrics patients describing their experiences contributing contextual patient generated health data. An experienced 4-person coding team completed a primarily inductive TA in which 50 codes were synthesized into 3 themes through iterative discussion. LLM-assisted TA was performed using ChatGPT (GPT-4o and GPT-4.5 models). In a 3-step process of code generation, code refinement, and theme generation, 54 initial codes were produced, refined to 25 codes, and consolidated into 3 final themes. Quantitative evaluation of the LLM model performance comparing human-generated and LLM generated codes and themes was performed with sentence-t5-xxl embeddings and a + 0.7 cosine similarity threshold.
    Results: The cosine-similarity analysis showed a precision of 100% and a recall of 88%. Code-level similarity scores ranged from 0.70 to 0.82 whereas theme-level similarity scores ranged from 0.80 to 0.83, with manual review confirming consistent coding and thematic boundaries. Human-only coded analysis was completed in approximately 39 person-hours compared to 5 person-hours for the LLM analysis. The approximate labor cost for human-only coded analysis was $1924 or $240 per transcript and $237 inclusive of paid subscription ($29.63 per transcript) for the LLM analysis.
    Discussion: An LLM-assisted multi-step approach can closely replicate human-only TA at reduced time and cost.
    Conclusion: An LLM-assisted approach offers a practical and scalable tool for qualitative sociotechnical research when combined with human oversight.
    Keywords:  human factors; large language models; patient generated health data; qualitative analysis; sociotechnical
    DOI:  https://doi.org/10.1093/jamiaopen/ooag091
  14. Biol Methods Protoc. 2026 ;11(1): bpag032
      Clinical trial statistical programming requires 12-24 full-time-equivalent months per Phase 3 study and remains a bottleneck in pharmaceutical research. Modern artificial intelligence coding agents reason capably but lack domain-specific tools: they cannot read proprietary statistical software datasets, parse analysis specifications, or generate standards-compliant code without extensive guidance. We present ClinAgent, a skill and tool layer that augments any artificial intelligence coding agent with clinical programming capabilities through Model Context Protocol tools. Its design separates minimal data access from rich domain logic: skills package prompts, rule engines, and decision trees encoding expert knowledge, while tools provide stateless input-output for statistical software datasets, spreadsheet specifications, and log files. In this single-study proof-of-concept evaluation, we validate ClinAgent's nine skills on artifacts from a production Phase 2 cardiovascular study, with synthetic datasets spanning 13 analysis domains and 102 109 observations. All skills pass functional validation. On this small sample, deterministic components identify one error and seven warnings without false positives and match all 56 subject-level variables; corresponding confidence intervals are wide, so these point estimates should be read as upper bounds pending replication. Prompt-based specification generation, dependent on the underlying language model, reaches 72.1% derivation accuracy overall, above 96% in simple domains and below 55% in complex ones, indicating that generated specifications require expert review. Our contributions include an agent-augmentation architecture, nine validated skills, tool implementations for clinical data formats, and a validation methodology distinguishing deterministic tool correctness from language-model-dependent output.
    Keywords:  CDISC standards; Model Context Protocol; SAS (RRID: SCR_008567); artificial intelligence; automation; clinical trial methodology; pandas (RRID: SCR_018214); statistical programming
    DOI:  https://doi.org/10.1093/biomethods/bpag032
  15. Laryngorhinootologie. 2026 Sep 08.
       Objective: Large Language Models such as ChatGPT are increasingly being discussed as tools to support medical decision-making. The aim of this study was to investigate whether the quality of responses generated by ChatGPT (GPT-4) to guideline-based ENT questions improves when the underlying evidence-based S3 guideline is provided as a PDF together with the respective question.
    Methods: Thirty guideline-based questions were derived from five current S3 guidelines. ChatGPT answered each question in two runs: once without and once with provision of the respective guideline as a PDF together with the question. Three ENT specialists independently and anonymously evaluated the responses regarding content accuracy and length/conciseness (scale 1-3).
    Results: Providing the guideline as a PDF together with the respective question resulted in a significant improvement in accuracy (1.77 vs. 1.26; p = 0.003) and conciseness (1.96 vs. 1.50; p = 0.004). A significant correlation between accuracy and conciseness was observed (p < 0.001). Interrater reliability was fair to moderate.
    Conclusion: The targeted provision of evidence-based guidelines as a PDF together with the respective guideline-based ENT question significantly improves response quality in ChatGPT. These findings highlight that the integration of medically sound primary sources is crucial for the accuracy and reliability of LLMs. LLMs therefore appear to be a potentially supportive tool in clinical and scientific assistance but do not replace medical expertise. In particular, potential "hallucinations" require continued critical medical and scientific oversight.
    DOI:  https://doi.org/10.1055/a-2931-6385
  16. Account Res. 2026 Sep 09. 2725029
       OBJECTIVE: Little attention has been paid to how publication and access models influence the evidence that AI systems retrieve, process, and summarize. This article introduces the concept of AI-induced evidence skew, proposes four pathways through which it may arise, and provides a preliminary empirical illustration.
    METHODS: Nine widely used large language models (LLMs) were prompted to generate literature-supported content. The accessibility status of the references provided was compared with the percentage of open-access records identified through searches of three scholarly databases.
    RESULTS: AI-induced evidence skew is a structural distortion in AI-assisted evidence synthesis whereby outputs disproportionately represent open-access or otherwise machine-accessible literature. Four proposed pathways are differential representation of literature in training data, retrieval-access constraints, human verification practices, and system- or user-imposed access restrictions. Among 78 verifiable references generated by LLMs, a mean of 95.38% had freely available full text (87.43% were open-access), compared with a mean open-access percentage of 57.83% across scholarly database search results.
    CONCLUSIONS: Although exploratory and non-generalizable, these findings support the hypothesis that AI-assisted search and synthesis may overrepresent publicly accessible literature. The concern is not the quality of open-access literature, but that accessibility may become an unintended determinant of included evidence. Mitigation requires greater awareness, human oversight, transparent reporting of AI use and retrieval limitations, verification of AI-generated references, and improved transparency regarding AI evidence coverage. Further empirical investigation across disciplines, platforms, and AI tools is warranted.
    Keywords:  Artificial intelligence in publishing; biomedical publishing; large language models; publication models; research integrity
    DOI:  https://doi.org/10.1080/08989621.2026.2725029