J Med Internet Res. 2026 Aug 03. 28
e90046
Background: Inference-time retrieval augmentation is increasingly used to improve the traceability and verifiability of large language model (LLM) applications in health care. Evaluation practices for text-based retrieval-augmented generation (RAG) and graph-structured RAG (GraphRAG) systems remain heterogeneous, which limits comparison across studies and complicates judgments about clinical readiness.
Objective: This review mapped evaluation methods for inference-time retrieval-augmented and graph-structured retrieval-augmented LLM systems in health care and characterized how evaluation constructs are defined, operationalized, and reported across system layers and evaluation-setting categories.
Methods: We conducted a scoping review in accordance with PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews), with search reporting informed by PRISMA-S (PRISMA literature search extension). Searches were conducted through May 14, 2026, in PubMed (MEDLINE), Web of Science Core Collection, IEEE Xplore, ACM Digital Library, arXiv, and medRxiv, with backward and forward citation tracking of included studies. Eligible records described health care-relevant LLM systems using inference-time RAG and reported at least 1 evaluation component. Data were charted on study characteristics, system design, retrieval-layer evaluation, evidence linkage, safety-related and GraphRAG-specific evaluation, and selected reporting and governance characteristics. We also constructed an evidence-and-gap map cross-classifying evaluation-setting categories with key evaluation domains.
Results: A total of 157 studies met the inclusion criteria. Clinical question answering was the most frequently represented application (89/157, 56.7%), followed by clinical decision support (70/157, 44.6%). Most evaluations were conducted in offline-only settings (140/157, 89.2%), whereas 17/157 (10.8%) studies reported workflow-facing, prospective, or deployment-level evaluation. Independent retrieval-layer evaluation was reported in 47/157 (29.9%) studies. Grounding and faithfulness evaluation was reported in 41/157 (26.1%) studies, and fine-grained evidence verification was reported in 22/157 (14%) studies. Human evaluation was reported in 94/157 (59.9%) studies, but interrater reliability was reported in 26/94 (27.7%) studies. LLM-as-judge evaluation was reported in 41/157 (26.1%) studies, with bias-control measures reported in 15/41 (36.6%) studies. Formal safety-related evaluation was reported in 45/157 (28.7%) studies. Among 27 (17.2%) GraphRAG studies, intermediate-artifact evaluation was reported in 11/27 (40.7%) studies, and graph construction evaluation was reported in 6/27 (22.2%) studies. The evidence-and-gap map showed limited coverage of fine-grained verification, contradiction handling, safety evaluation, LLM-as-judge safeguards, GraphRAG construction evaluation, and GraphRAG intermediate-artifact evaluation in workflow-facing, prospective, or deployment-level settings.
Conclusions: Evaluation of health care RAG and GraphRAG systems has expanded rapidly, yet reporting and operational definitions remain inconsistent across evaluation layers. Current evidence remains concentrated in offline evaluation, with limited workflow-facing, prospective, or deployment-level assessment of retrieval quality, fine-grained evidence linkage, safety, LLM-as-judge safeguards, GraphRAG construction quality, and GraphRAG intermediate artifacts. This review maps these gaps across evaluation-setting categories and translates them into synthesis-informed evaluation considerations. These findings suggest that future evaluation may need to move beyond end-to-end benchmark performance toward more transparent, layer-specific, safety-oriented, and clinically contextualized assessment before workflow-facing implementation.
Keywords: GraphRAG; artificial intelligence; clinical decision support systems; evaluation studies as topic; hallucination; information storage and retrieval; large language models; natural language processing; retrieval-augmented generation; scoping review