bims-arines Biomed News
on AI in evidence synthesis
Issue of 2026–08–16
eighteen papers selected by
Farhad Shokraneh, Systematic Review Consultants LTD



  1. J Health Econ Outcomes Res. 2026 ;13(2): 46-54
       Background: Systematic literature reviews (SLRs) are central to evidence generation in health economics and outcomes research, but traditional SLR workflows are labor-intensive. Artificial intelligence (AI) is increasingly being explored to support evidence synthesis tasks.
    Objectives: To describe methods used to assess AI performance in screening and data extraction in SLRs, and to provide recommendations on how to ensure appropriate use of AI in literature reviews.
    Methods: We conducted an AI-assisted SLR that mirrored a traditional SLR performed by humans only and assessed the performance of the AI tools employed. AI was used to screen titles/abstracts, screen full texts, and conduct data extraction. Using the traditional SLR as the benchmark, AI performance for title/abstract screening was evaluated in terms of accuracy, recall, and precision at predefined screening milestones, followed by the accuracy assessment of full-text screening by AI. Data items extracted by AI were verified by humans and categorized as correct, incomplete, missing, incorrect, or requiring human checks. The average accuracy rate calculated on the basis of correct data items extracted was used as an indicator of AI performance in data extraction.
    Results: During title/abstract screening, accuracy and recall remained high but precision was low, increasing false positives. Full-text screening identified a minority of studies included in the traditional SLR. Mean extraction accuracy was 72.93% (range, 57.69%-88.46%). No data were hallucinated by the AI model used.
    Keywords:  artificial intelligence; automation tools; evidence synthesis; large language models; machine learning; systematic literature reviews
    DOI:  https://doi.org/10.36469/001c.165173
  2. Laryngoscope. 2026 Aug 11.
       OBJECTIVES: Systematic literature reviews (SLRs) are time-intensive and resource-consuming. While large language models (LLMs) have shown promise in encoding clinical knowledge, evidence for their performance on complex text analysis, necessary for SLRs in otolaryngology, remains limited. In this proof-of-concept study, we investigate an LLM's performance in screening relevant articles for a SLR with the novel screening of title and abstracts, reevaluation, and full-text review (STARR) protocol.
    METHODS: ENTGPT (based on GPT-4o) was compared to two human reviewers in article inclusion/exclusion decisions using the traditional and STARR screening protocols. ENTGPT was provided with inclusion and exclusion criteria, titles, abstracts, and full texts (if available) of the 850 articles retrieved in the original search. The model's decisions were compared to those made by two human reviewers.
    RESULTS: ENTGPT, using the STARR protocol, achieved 99.87% accuracy in article classification compared with human reviewers (95% CI: 0.99-1.0), including 100% specificity and 95% sensitivity. When using the traditional protocol, sensitivity declined to 35%. ENTGPT, using the traditional protocol, achieved 99.47% accuracy in article classification compared with human reviewers (95% CI: 0.99-1.0), including 100% specificity and 35% sensitivity. When using the STARR protocol, accuracy improved to 99.87% and sensitivity markedly increased to 95%.
    CONCLUSIONS: ENTGPT accurately replicated human reviewers in article selection and data extraction for an otolaryngology SLR using the STARR and traditional protocols. This performance suggests that LLMs could be employed to significantly streamline the SLR process, potentially saving substantial time and resources for researchers.
    LEVEL OF EVIDENCE: N/A.
    Keywords:  STARR protocol; large language models; screening automation; systematic literature review automation
    DOI:  https://doi.org/10.1002/lary.70801
  3. Cochrane Evid Synth Methods. 2026 Sep;4(5): e70101
       Introduction: Rapid reviews aim to deliver timely evidence for decision-makers when full systematic reviews are not possible or practical. Efficient selection of studies is challenging when questions are complex, or the evidence base is diffuse.
    Methods: In a rapid review on trial informativeness, our team used EPPI-Reviewer, a web-based systematic review platform that supports document management, screening, and machine learning prioritization, to conduct title and abstract screening. We developed a machine learning classifier model within the platform to rank records by predicted relevance based on coding structures aligned with predefined criteria.
    Results: The classifier model correctly concentrated relevant studies in the higher probability bands, which allowed most eligible records to be identified early. As screening progressed to lower probability bands, the number of newly identified records declined, indicating effective prioritization. Real-time collaboration and a clear audit trail supported consistent decision-making across reviewers. Limitations included the initial effort to train the model and potential subscription costs.
    Conclusion: Classifier assisted screening in EPPI-Reviewer improved the feasibility of conducting a rapid review on a complex topic within a limited timeframe. Although the risk of missed citations remains, this is inherent to any review method. With appropriate training and support, classifier models and platforms like EPPI-Reviewer can enhance both efficiency and transparency in rapid evidence synthesis.
    Keywords:  clinical trial; data mining; information storage and retrieval; machine learning; review
    DOI:  https://doi.org/10.1002/cesm.70101
  4. JB JS Open Access. 2026 Jul-Sep;11(3):pii: e26.00102. [Epub ahead of print]11(3):
      Systematic reviews and meta-analyses remain time-consuming and labor-intensive. We developed and validated a locally executed agentic artificial intelligence (AI) framework for deduplication, screening, and structured data extraction in systematic reviews, reported following Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA)-Transparent Reporting of AI in Comprehensive Evidence Synthesis (trAIce) guidelines. A multiagent pipeline of specialized agents for deduplication, title/abstract screening, structured data extraction, and verification was executed entirely locally to ensure data governance and reproducibility. Human-in-the-loop validation compared outputs against dual independent reviewers across 3 spine surgery systematic reviews, assessing accuracy, inter-rater agreement (Cohen's κ), time savings, and clinically critical error rates. Across 6,214 records, deduplication achieved near-perfect agreement with human reviewers (κ = 0.98), title and abstract screening yielded higher concordance than human screening (κ = 0.91) while reducing full-text review volume by 83%, and structured data extraction reached substantial agreement (κ = 0.87). The framework reduced reviewer time by 91.1% (90.8%-91.4%), a mean saving of 19.5 hours per review (p < 0.001). Clinically critical discrepancies were rare (<1%) and traceable, with no fabricated or hallucinated data introduced. A locally executed agentic AI framework has the potential to deliver accurate, efficient, and secure automation of systematic review tasks with human oversight, offering a reproducible pathway for trustworthy evidence synthesis under PRISMA-trAIce standards.
    DOI:  https://doi.org/10.2106/JBJS.OA.26.00102
  5. Evid Based Dent. 2026 Aug 08.
       OBJECTIVE: To evaluate the accuracy and reliability of three AI platforms ChatGPT, Perplexity, and Google Gemini in assessing the methodological quality of systematic reviews using the AMSTAR 2 checklist, compared with expert manual evaluation in dental research.
    METHODS: A cross-sectional comparative study was conducted to assess the performance of three AI platforms ChatGPT, Perplexity, and Google Gemini in evaluating the methodological quality of 35 systematic reviews using the AMSTAR 2 checklist. Manual assessments by a domain expert served as the reference standard. Each AI system was prompted with a standardized AMSTAR 2 query, and item-level outputs were collected for direct comparison. Key metrics included percentage agreement, error proportions, and inter-rater reliability measured by Cohen's kappa. Error proportions represent the proportion of discordant assessments out of total valid pairwise comparisons across 16 AMSTAR-2 items. Differences between LLM-generated and reference AMSTAR-2 ratings were summarized using effect estimates with corresponding 95% confidence intervals. Comparative performance across platforms was assessed based on confidence-interval overlap rather than hypothesis testing. Results are presented as effect estimates with corresponding 95% confidence intervals, without hypothesis testing or statistical dichotomization. This approach provided a robust and reproducible framework to benchmark AI-assisted quality appraisal in dental evidence synthesis.
    RESULTS: Among 35 systematic reviews assessed, Perplexity demonstrated the highest agreement with expert AMSTAR-2 ratings (error proportion: 19.0%; weighted κ_w: 0.78, 95% CI 0.71-0.85), followed by ChatGPT (error proportion: 22.9%; weighted κ_w: 0.62, 95% CI 0.54-0.70) and Google Gemini (error proportion: 43.9%; weighted κ_w: 0.41, 95% CI 0.33-0.49). Perplexity also achieved the best sensitivity (81.3%, 95% CI 76.5-85.4%) and specificity (82.7%, 95% CI 78.1-86.5%) for correctly identifying high-quality reviews. Non-overlapping 95% confidence intervals suggest meaningful differences in performances among platforms, with Perplexity showing superior agreement across all metrics. Across all platforms, agreement was generally higher for non-critical AMSTAR-2 domains involving clear and structured reporting, whereas performance was weaker for critical domains requiring interpretation of complex methodological details, risk-of-bias considerations, and evidence synthesis procedures.
    CONCLUSIONS: Perplexity demonstrated the highest accuracy and agreement with expert assessments of the methodological quality of systematic reviews, suggesting its potential as a supportive AI tool for AMSTAR-2-based appraisal in dental evidence synthesis. In contrast, systematic biases observed in ChatGPT and Google Gemini underscore the continued need for human oversight to ensure the validity of methodological assessments. Differences in agreement and error proportions were observed across all models when compared with expert AMSTAR-2 evaluations, indicating meaningful variability in methodological appraisal performance, reinforcing that AI-assisted appraisal of systematic review methodology should complement rather than replace expert human judgment in dental research.
    DOI:  https://doi.org/10.1038/s41432-026-01238-8
  6. J Med Internet Res. 2026 Aug 13. 28 e98374
       Background: Large language models (LLMs) are increasingly used in qualitative research, but their reliability compared to human analysis, especially on large, non-English datasets, is unclear. Previous studies on older models (like GPT-4) show limitations in nuance and token capacity.
    Objective: This study compared the qualitative analysis capabilities of OpenAI's GPT-5 and Google's Gemini 2.5 Pro (Gemini) with a human qualitative analysis. The study uses a large dataset of 317 Dutch newspaper articles from January 1, 2020, to December 31, 2023, investigating the sentiment toward nurses during the COVID-19 pandemic.
    Methods: The study used a 2-phase methodology. First, a thematic comparison was conducted where the human researchers, GPT-5, and Gemini independently generated inductive coding trees from the entire corpus. Second, a comparative test was performed where all 3 coders applied a predefined codebook to a 10% stratified random subsample. The human baseline was validated through double-coding by 2 independent researchers, achieving an acceptable intercoder reliability (α=0.703). The AI analysis was iterative, using model-optimized prompts and an article-by-article approach.
    Results: Both AI models successfully identified third-order themes (eg, "Health care heroes") consistent with the data. In deductive application, however, both models systematically overcoded compared to the human consensus (181 and 183 codes vs 138), resulting in low intercoder reliability against the human baseline (α=0.487) for GPT-5 and (α=0.507) for Gemini.
    Conclusions: This study suggests a potential divergence in analytical logic. The observed coding frequencies indicate that LLMs may default to semantic presence (literal frequency), whereas human coders appear to prioritize interpretive significance (contextual weight), leading to systematic overcoding. Consequently, this article argues that LLMs should not be viewed as autonomous researchers but as high-sensitivity filtering instruments requiring human calibration. This study concludes that AI can serve as a valuable assistant for qualitative researchers. Still, it benefits from a rigorous, iterative, and human-in-the-loop approach to manage methodological friction and ensure nuanced, valid analysis.
    Keywords:  AI; COVID-19; comparative study; health services research; large language models; medical informatics; natural language processing; nursing; qualitative research; sentiment analysis
    DOI:  https://doi.org/10.2196/98374
  7. J Med Internet Res. 2026 Aug 14. 28 e91215
       Background: The exponential expansion of biomedical literature has created an urgent need for efficient methods to recognize and extract population, intervention, comparison, and outcome (PICO) elements-the foundational elements of evidence-based medicine.
    Objective: This study systematically evaluated 2 complementary approaches for automating PICO recognition and extraction in medical literature: prompt engineering optimization and parameter-efficient fine-tuning (PEFT) of large language models (LLMs).
    Methods: We developed a dual-phase methodological framework: (1) systematic prompt optimization incorporating in-context learning, chain of thought (COT), and multipath reasoning strategies; and (2) PEFT of the LLM architecture using low-rank adaptation (LoRA), quantized LoRA, and freeze techniques. The PubMed-PICO and NICTA-PIBOSO benchmark datasets were used for recognition tasks, and the EBM-NLP dataset was used for extraction tasks. Performance metrics included precision, recall, and F1-score. F1-score was adopted as the major metric as it balances precision and recall.
    Results: For prompt engineering, COT achieved the overall best performance across both recognition and extraction tasks. For example, in the recognition task, COT obtained strong average F1-scores of 77.1% (SD 0.5%) for the population element and 84.5% (SD 0.4%) for the outcome element on PubMed-PICO. In the extraction task, COT achieved the highest average F1-score of 73.9% across 3 PICO elements (the population, intervention, and outcome elements) on EBM-NLP. These results suggest that, for smaller models such as those with 3B parameters, explicit step-by-step guidance in COT is more effective than more complex prompting strategies. In PEFT implementations, for example, LoRA achieved the best recognition performance (mean F1-score 91.7%, SD 0.3% for population) on PubMed-PICO, whereas quantized LoRA showed the best extraction capability (mean F1-score 79.3%, SD 0.5% for intervention) on EBM-NLP. Fine-tuned models achieved competitive performance across all datasets, with notable gains on NICTA-PIBOSO and EBM-NLP. PEFT further enhanced the model's overall performance compared with prompt engineering, with element-dependent differences across PICO categories.
    Conclusions: Our findings indicate that LLMs can effectively automate PICO recognition and extraction through 2 complementary approaches. First, prompt engineering allows the model to perform tasks directly without altering its internal settings. Second, the PEFT method further unlocks their maximum performance potential by incorporating additional fine-tuning based on prompt engineering. This work makes significant advances and provides critical insights for optimizing methodological approaches in clinical applications related to or comprising PICO extraction and recognition tasks.
    Keywords:  LLM; PICO; evidence-based medicine; large language model; natural language processing; parameter-efficient fine-tuning; population, intervention, comparison, outcome; prompt engineering
    DOI:  https://doi.org/10.2196/91215
  8. J Clin Med. 2026 Aug 02. pii: 6005. [Epub ahead of print]15(15):
      Background: The reliable integration of large language models (LLMs) into neuroimaging data extraction workflows remains unresolved. Prior benchmarking shows that exact-match accuracy underestimates LLM extraction performance, but whether inter-model consensus and variable complexity can guide automation remains unclear. We evaluated whether inter-model consensus can serve as a confidence signal for human-artificial intelligence (AI) extraction and can guide complexity-stratified workflow triage. Methods: Four frontier LLMs were queried via OpenRouter with an identical zero-shot structured prompt to extract 22 predefined variables from 91 peer-reviewed neuroimaging AI articles, yielding 2002 article-variable items per model. Variables were stratified a priori into low- (n = 7), medium- (n = 8), and high-complexity (n = 7). Performance was compared with an expert reference using exact-match and semantic-equivalence accuracy. Item-level consensus and five triage strategies characterized the efficiency-accuracy trade-off. Results: Semantic-equivalence accuracy converged to 80.5-83.4% across models despite approximately ten percentage-point exact-match differences. Unanimous 4/4 consensus occurred in 45.6% (910/1994) of items, with exact-match accuracy of 85.8%, rising to 95.3% after semantic normalization; however, 14.2% still failed to match the reference. Reliability was complexity-dependent: 96.6% for low-complexity variables, 73.2% for medium-complexity variables, and 38.1% for high-complexity variables. A hybrid strategy auto-accepting 4/4 items and routing 3/4 items to rapid verification reduced estimated review effort by approximately 59%. Conclusions: Inter-model consensus is useful, but it is incomplete and depends on variable complexity. We show that LLM-assisted extraction in neuroimaging AI is a complexity-stratified workflow design problem: low-complexity neuroimaging variables may be selectively automated, while medium-complexity variables require rapid verification, and high-complexity methodological variables should remain human-led.
    Keywords:  artificial intelligence; data extraction; evidence synthesis; large language models; neuroimaging; variable complexity
    DOI:  https://doi.org/10.3390/jcm15156005
  9. J Med Internet Res. 2026 Aug 10. 28 e89858
       Background: Although widespread antiretroviral therapy has extended the life expectancy of people living with HIV, cardiovascular disease (CVD) has emerged as a primary comorbidity. Persistent cross-specialty knowledge gaps in routine clinical practice lead to suboptimal adherence to guidelines. Integrated, evidence-based tools are urgently needed to overcome these interdisciplinary barriers. While large language models (LLMs) have demonstrated significant capabilities in medicine, no systematic evaluation has assessed their ability to facilitate multidisciplinary CVD management for people living with HIV.
    Objective: This study compared the performance of 4 mainstream AI models (DeepSeek-V3, DeepSeek-R1, ChatGPT-4o, and ChatGPT-o4-mini) against 12 human clinicians (8 infectious disease specialists and 4 cardiologists) in addressing guideline-based CVD management tasks for people living with HIV.
    Methods: Based on 4 authoritative domestic and international HIV/CVD guidelines, a structured 25-question assessment was developed via 2 rounds of Delphi consultation. Standard reference answers and an evaluation framework were finalized through expert consensus. LLM responses were generated using standardized prompts. Clinicians answered identical questions via one-on-one structured interviews, transcribed verbatim. Six multidisciplinary experts independently rated all responses across 4 dimensions-accuracy, completeness, readability, and reliability-using a 4-point ordinal scale (1=poor to 4=excellent). Cumulative link mixed models analyzed intergroup differences.
    Results: All AI models achieved significantly higher scores than clinicians across all dimensions (P<.001). The AI group's mean scores ranged from 3.44 to 3.68 (median 4, IQR 3.0-4.0; coefficient of variation=0.145-0.178). Conversely, clinicians' scores were lower (mean 1.78-2.05; median 2, IQR 1.0-3.0; coefficient of variation=0.428-0.473) with marked dispersion. DeepSeek-R1 delivered the optimal performance, significantly outperforming the other 3 models (all P<.001). Specialty-stratified analysis revealed no significant overall score difference between cardiologists and infectious disease specialists (odds ratio 0.92, 95% CI 0.84-1.01; P=.09). However, dimension-specific analysis indicated that cardiologists scored higher in accuracy (odds ratio 0.81, 95% CI 0.67-0.97; P=.03). Domain-specific divergence was evident: cardiologists outperformed infectious disease specialists in CVD risk assessment (2.26 vs 1.83), whereas infectious disease specialists led in drug adverse effect evaluation (2.23 vs 1.65).
    Conclusions: In this structured question-and-answer study, LLMs outperformed human clinicians across all metrics for HIV-associated CVD management, with DeepSeek-R1 achieving superior composite scores. These findings validate DeepSeek-R1's potential as a cross-disciplinary decision-support tool capable of integrating complex clinical knowledge, mitigating specialty gaps, and enhancing information precision. Integrating AI systems into multidisciplinary workflows, complemented by targeted clinical training, may optimize the management of complex comorbidities in people living with HIV.
    Keywords:  AI; HIV; cardiovascular disease; comorbidity management; patient education
    DOI:  https://doi.org/10.2196/89858
  10. J Evid Based Soc Work (2019). 2026 Aug 13. 1-16
       PURPOSE: Locating relevant studies is the first step of evidence-based practice, yet most searching relies on keyword matching. Artificial intelligence (AI) tools called embedding models search by meaning, but the best-known options are paid commercial services. The study asked which free embedding models best search the social work literature, whether they match the commercial standard, and whether rerankers are needed.
    MATERIALS AND METHODS: Twelve free embedding models and two commercial OpenAI models were tested on 64,956 social work records (1989-2025) using 150 curated queries. Two frontier AI judges, from families unrelated to every tool evaluated, made 50,328 blind head-to-head comparisons (nDCG@10). A judge-free known-item test (496 queries) and a blind 120-pair expert human instrument provided validation.
    RESULTS: Free tools matched or beat the commercial standard. Two free models outperformed the paid flagship; a free 300-million-parameter model beat the paid default, essentially tied for first at finding specific papers (91.9%), and with a reranker was the best configuration overall (.846). Keyword search trailed every embedding model (.604 vs. .680-.842). Rerankers rescued weak models but added nothing to the strongest. Judges agreed on 85.5% of comparisons, and committee-to-rater agreement (69-76%) matched or exceeded rater-to-rater agreement (69-71%).
    DISCUSSION: Score differences among leading models are too small to change what a searcher sees; tool choice should rest on size, speed, cost, and privacy.
    CONCLUSION: High-quality, meaning-based search of the social work literature is achievable with free tools on an ordinary computer: no subscription, no queries sent to an outside company.
    Keywords:  Literature search; evidence-based practice; information retrieval; open-source artificial intelligence; research infrastructure; semantic search
    DOI:  https://doi.org/10.1080/26408066.2026.2718944
  11. Anal Chem. 2026 Aug 11. 98(31): 22595-22608
      Large language models (LLMs) have evolved into versatile tools of the 21st century, simplifying repetitive and labor-intensive tasks in everyday life. Here, we aimed to test the feasibility of using different LLMs to assess the greenness of analytical procedures according to the "Analytical GREEnness Metric Approach" (AGREE) by extracting specific data corresponding to the 12 principles of green analytical chemistry from scientific articles. Seven open-access articles on plant natural products using different analytical techniques were evaluated with the five most popular artificial intelligence (AI) tools (ChatGPT, Copilot, Perplexity, Claude, and Gemini), which were tasked to obtain specific data, along with a justification for the selection of this data. Additionally, different versions (basic and advanced) of the same tool (Perplexity and Perplexity Pro), output types (PDF file and link to the online version), and repeatability (three times the same task) were compared. All extracted data were used to calculate AGREE scores, and the final results were compared with those obtained by experts. AI tools were able to assess greenness with a high degree of accuracy, similar to that of trained researchers. Furthermore, a verification/comparison study demonstrated the possibility of critically examining the greenness assessment of developed methods to facilitate standardizing greenness evaluation. The final application of advanced AI tools dedicated to scientific research confirmed the greenness scores, indicating high consistency between all considered evaluation methods.
    DOI:  https://doi.org/10.1021/acs.analchem.6c02529
  12. J Am Med Inform Assoc. 2026 Aug 04. pii: ocag122. [Epub ahead of print]
       OBJECTIVES: Accurate citation of relevant publications is essential for scientific integrity in biomedical research. Large language models (LLMs) excel at text generation but often hallucinate fabricated or inaccurate citations. Retrieval-augmented generation (RAG) can mitigate these errors, yet current approaches lack semantic precision in evidence retrieval. This study aims to develop a domain-specific RAG system for reliable, context-specific biomedical citation recommendations.
    MATERIALS AND METHODS: We introduce CiteSure, a sentence-level citation recommendation tool designed to deliver reliable, evidence-based, and context-specific references using LLMs. CiteSure utilizes a 2-stage retrieval-augmented generation (RAG) framework, combining a domain-specific dense retriever (BioLLM2Vec) and reranker (BioRankLLaMA), adapted from LLaMA3-8B-Instruct using biomedical-specific training data. CiteSure leverages the complementary strengths of retrieval and generative LLM models, ensuring factual precision and contextual alignment. We evaluated CiteSure on a curated Alzheimer's disease dataset, comparing it to standalone LLMs and traditional retrieval-based methods.
    RESULTS: CiteSure achieved 100% factual accuracy and the highest relevance score of 77.50%, outperforming all baselines. BioLLM2Vec retrieved relevant articles with over 80% accuracy in the top 100 candidates. BioRankLLaMA consistently outperformed baseline rerankers across MAP, MRR, and Precision@5 metrics, confirming the benefit of domain-specific adaptation and contrastive fine-tuning.
    DISCUSSION AND CONCLUSION: Our results demonstrate that CiteSure, built on a 2-stage retrieval-augmented generation framework, effectively integrates domain-specific retrieval with LLM-based generation to achieve substantial improvements over baseline approaches. Our work underscores the importance of domain-specific adaptation in biomedical citation recommendation and provides publicly available datasets, models, and code for support future research.
    Keywords:  biomedical citation recommendation; large language models; retrieval augmentation
    DOI:  https://doi.org/10.1093/jamia/ocag122
  13. Diagnostics (Basel). 2026 Aug 04. pii: 2456. [Epub ahead of print]16(15):
      Background/Objectives: Large language models (LLMs) show promise for clinical decision support, yet their accuracy in interpreting specialized medical guidelines remains uncertain. Retrieval-augmented generation (RAG) may enhance performance by grounding responses in authoritative knowledge bases. This study aimed to compare the accuracy, comprehensiveness, and safety of RAG-enhanced versus standard LLMs for answering clinical questions derived from the German S3 guideline for oral cavity carcinoma. Methods: We conducted a prospective, single-blind benchmark study evaluating six LLMs: one RAG-enhanced model (Custom GPT with guideline access), one consensus-based model (ConsensusGPT), and four standard models (DeepSeek-V3.2, Mistral Small 3.2, Qwen3-Next-80B, GPT-OSS-120B). Fifty clinical questions covering 17 guideline domains were presented to each model three times, yielding 900 evaluations. Three expert reviewers assessed responses using 5-point Likert scales for accuracy, comprehensiveness, and clarity, under a single-blind procedure, the effectiveness of which was tested by a pre-specified manipulation check. We then ran a paired within-model experiment in which each base model was queried with and without guideline access through a transparent, openly released retrieval pipeline, and scored every response with a condition-blind automated judge alongside deterministic retrieval metrics computed from the logs. Secondary outcomes included hallucination rates and guideline citation behavior. Inter-rater reliability was assessed using intraclass correlation coefficients (ICCs). Results: In a paired within-model design that held each base model fixed, adding transparent guideline retrieval improved accuracy-significantly in the three weaker open-weight models (Mistral, Qwen3, and GPT-OSS) and directionally in the already-strong DeepSeek and GPT-5 bases. Because a pre-specified blinding check found that experts could still identify retrieval-augmented answers with 98.5% accuracy, we anchored causal interpretation on measures that do not depend on the human raters, ranked by their independence: deterministic, log-derived retrieval metrics first, and then an automated, condition-blind LLM judge, whose agreement with the experts (Spearman ρ = 0.81, 95.7% within-one agreement) establishes shared calibration rather than independence from their bias. Deterministically from the retrieval logs, citation groundedness rose from 0% to 51-89% and retrieval recall@5 was 92%. On the judge, content-level hallucination fell from 42% to 4% and accuracy rose by a pooled +0.64 points (95% CI 0.47-0.80); the accuracy gain persisted after adjustment for response length (+0.48, 95% CI 0.22-0.73), which retrieval shortened rather than lengthened. The accuracy gain was large for weaker base models and small or non-significant for already-strong ones, whereas the hallucination and auditability gains were consistent across all models. The human ratings reproduced the judge's accuracy effect (+0.61, 95% CI 0.49-0.74), and GPT-5 run through the transparent pipeline showed no significant difference from the proprietary Custom GPT (judge accuracy 4.48 vs. 4.58). Conclusions: Guideline retrieval yields a reproducible, largely base-independent improvement in the safety and auditability of LLM answers to clinical guideline questions, with accuracy gains concentrated in weaker base models. Because retrieval-augmented answers are recognizable to experts, rigorous evaluation should rely on rater-independent measures, and residual hallucination continues to require human oversight.
    Keywords:  artificial intelligence; benchmark study; clinical decision support; clinical practice guidelines; evidence-based medicine; hallucination; large language models; oncology; oral cavity carcinoma; retrieval-augmented generation
    DOI:  https://doi.org/10.3390/diagnostics16152456
  14. Diagnostics (Basel). 2026 Aug 01. pii: 2435. [Epub ahead of print]16(15):
      Background/Objectives: Pharmacovigilance workflows rely heavily on unstructured text across diverse sources. Here, we systematically reviewed how large language models (LLMs) are being explored as support tools for adverse drug reaction (ADR) detection, extraction, triage, and documentation, highlighting their potential for precision medicine and big data-enabled safety monitoring. Methods: Following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses 2020 guidelines, we systematically searched PubMed, Scopus, and Web of Science for studies published between January 2022 and March 2026. Ultimately, 83 empirical studies satisfied the inclusion criteria. A narrative synthesis was conducted to address methodological heterogeneity across these studies. Results: LLM applications were concentrated in constrained information-extraction and classification tasks, including signal evaluation, clinical-note extraction, social media surveillance, and literature screening. Quantitative performance varied substantially by system design: error-correction prompting yielded an F1-score of 0.921 for ADR named entity recognition, whereas retrieval-augmented generation improved data-retrieval accuracy from 8.3% to 78.3%. Most studies were retrospective, benchmark-based, or proof-of-concept evaluations. Across 581 paired pre-consensus domain judgements, observed inter-rater agreement was 90.4% and Cohen's κ was 0.837 (95% CI 0.772-0.895). Hallucination, low specificity, prompt sensitivity, narrow datasets, and weak external validation remained common limitations. Conclusions: Current evidence supports supervised, task-specific applications of LLMs for extraction, triage, retrieval, and documentation rather than autonomous pharmacovigilance decision-making. Prospective evaluation, external validation, transparent reporting, and accountable human oversight are required before high-stakes clinical or regulatory deployment.
    Keywords:  adverse drug event; adverse drug reaction; big data; clinical notes; drug safety; large language models; literature screening; pharmacovigilance; precision medicine; signal detection
    DOI:  https://doi.org/10.3390/diagnostics16152435
  15. Public Health Pract (Oxf). 2026 Dec;12 100827
       Objectives: Health inequalities in primary care persist across many high-income countries, with populations experiencing disadvantage often receiving less care despite greater need. Traditional evidence synthesis methods struggle to keep pace with the growing volume of research to help policy makers and practitioners to take timely, evidence-informed action. Our aim was to develop a Living Evidence Map of interventions that address inequalities in primary care, using machine learning (ML) to enhance the efficiency of evidence identification, screening, and mapping.
    Study design: Evidence synthesis.
    Methods: We used EPPI-Reviewer software to train a machine learning classifier to identify relevant studies. This was complemented by citation searching, and records were manually screened in order of predicted relevance (priority screening). Included studies were coded by intervention type, disadvantaged group, and health or care outcome, which was then visualised using EPPI-Visualiser.
    Results: A total of 31,871 articles were screened, resulting in 577 primary studies, 481 systematic reviews, and 6 umbrella reviews being included in the Living Evidence Map. Most studies focused on ethnic minority groups, with common interventions including education, advice and counselling, and culturally tailored care. There was a paucity of studies targeting gender and sexual minorities, and structural interventions (e.g. funding, workforce).
    Conclusions: The Living Evidence Map offers a dynamic, policy-relevant tool for navigating the evidence base on health inequalities in primary care. It serves both researchers and decision-makers by making it easier to see what evidence exists, where it's concentrated, and where there are gaps. However, further work is needed to improve inclusion of grey literature evidence and interventions addressing intersectional disadvantage to support decision-making for end users.
    Keywords:  Health inequalities; Living evidence map; Machine learning
    DOI:  https://doi.org/10.1016/j.puhip.2026.100827
  16. BMJ Health Care Inform. 2026 Aug 11. pii: e102207. [Epub ahead of print]33(1):
       OBJECTIVES: Clinical practice guidelines are a cornerstone of evidence-based medicine, yet their implementation in routine care remains inconsistent. Large language models (LLMs), particularly with Retrieval-Augmented Generation (RAG), have shown strong performance in medical question answering, but their ability to use knowledge from German-language guidelines has not been systematically evaluated due to a lack of a dedicated benchmark. We therefore developed such a benchmark and evaluated guideline-based question answering with different LLMs and retriever configurations.
    METHODS: We developed cpgQA-DE, an expert-validated benchmark dataset of 200 multiple-choice questions derived from 10 current German clinical practice guidelines across five specialties. All questions were reviewed for correctness, relevance and complexity. The dataset includes case-based and knowledge-based questions with metadata on guideline source, specialty, relevance and difficulty. We evaluated LLM performance using a RAG-based pipeline built on a corpus of German guidelines, comparing multiple model-retriever combinations.
    RESULTS: RAG integration substantially improved accuracy across all tested models and for most retrievers. The best-performing configuration, GPT-5 combined with the multilingual-e5-large retriever, achieved an accuracy of 95%. Notably, the open-weight model gpt-oss-120b reached 90% accuracy when used with RAG.
    CONCLUSIONS: cpgQA-DE enables systematic and reproducible offline evaluation of guideline-aware question answering systems in the German healthcare context. The observed performance gains with RAG support the use of LLMs augmented with quality-assured external knowledge. Such systems may help bridge the evidence-practice gap by providing guideline-based recommendations at the point of care, while strong performance of open-weight models suggests potential for on-premises deployment in privacy-sensitive clinical environments.
    Keywords:  Decision Support Systems, Clinical; Evidence-Based Medicine; Health Information Systems; Large Language Models; Medical Informatics Applications
    DOI:  https://doi.org/10.1136/bmjhci-2026-102207
  17. PLoS One. 2026 ;21(8): e0355603
      The rapid growth of user-generated textual content on the internet has intensified the need for accurate and scalable text classification methods. However, supervised learning approaches remain heavily constrained by the high cost and effort required for manual data annotation, particularly in large and heterogeneous datasets. To address this challenge, this paper proposes a novel hybrid active learning framework for efficient classification of unlabeled text data. The proposed approach integrates multiple classical machine learning classifiers-Support Vector Machines, Logistic Regression, Naive Bayes, and Random Forest-within a hybrid ensemble architecture, combined with a pool-based active learning strategy to iteratively select the most informative unlabeled instances for annotation. Textual data are transformed into numerical representations using several feature extraction techniques, including Bag-of-Words, TF-IDF, Word2Vec, and BERT-based embeddings, allowing for a comprehensive evaluation of representation effectiveness. Extensive experiments are conducted on four diverse benchmark datasets from healthcare, finance, spam detection, and e-commerce domains. The results consistently demonstrate that the proposed hybrid active learning model outperforms traditional ensemble classifiers across all datasets and evaluation metrics. In particular, TF-IDF-based hybrid ensembles achieve the highest gains in accuracy, precision, recall, and F1 score, while requiring substantially fewer labeled instances. Furthermore, the proposed framework exhibits strong robustness in imbalanced classification scenarios, significantly improving minority class detection. Overall, the findings confirm that combining hybrid ensemble learning with active learning offers an effective, lightweight, and cost-efficient alternative to purely transformer-based approaches, making it well-suited for real-world text classification tasks where labeled data are scarce or expensive.
    DOI:  https://doi.org/10.1371/journal.pone.0355603
  18. Nature. 2026 Aug;656(8127): 532
      
    Keywords:  Ethics; Machine learning; Scientific community
    DOI:  https://doi.org/10.1038/d41586-026-02490-9