J Med Internet Res. 2026 Jul 31. 28
e98519
Background: Continuous glucose monitoring (CGM) is central to diabetes care, but explaining CGM patterns consistently and empathetically remains time-intensive in clinical practice. Large language model (LLM)-based systems may support patient-facing interpretation of CGM data, but evidence remains limited for retrieval-grounded tools evaluated against clinician-authored responses in counseling scenarios. The system was intended for CGM interpretation and communication support rather than autonomous therapeutic decision-making.
Objective: This study aimed to evaluate whether a retrieval-grounded LLM-based conversational agent (CA) could support patient understanding of CGM data and preparation for diabetes consultations by generating responses to questions arising during CGM-informed diabetes counseling, with quality comparable to clinician-authored responses.
Methods: We developed a scaffolded LLM-based CA for CGM interpretation and diabetes counseling support. The system was designed to provide plain-language explanations of CGM patterns and responses to diabetes management questions while avoiding directive or individualized medical advice, such as recommending medication initiation, dose adjustment, or regimen changes. Around 12 CGM-informed cases, each comprising a deidentified CGM trace, a synthetic patient vignette, and accompanying CGM visual materials, were constructed from using available clinical datasets. Between October 2025 and February 2026, 6 senior UK diabetes clinicians each reviewed 2 assigned cases and answered 24 questions (12 per case). In a source-masked multirater evaluation, each CA-generated and clinician-authored response was independently rated by 3 clinicians on 6 quality dimensions, including clinical accuracy, guideline adherence, actionability, personalization, communication clarity, and empathy. Safety flags and perceived source labels were also recorded. The primary analysis used linear mixed effects models with random intercepts for case and rater.
Results: A total of 288 unique responses (144 CA and 144 clinician responses) were evaluated, generating 864 ratings. CA-generated responses received higher quality scores than clinician-authored responses under controlled vignette-based conditions, with mean scores of 4.37 (SD 0.57) versus 3.58 (SD 0.90) and an estimated mean difference of 0.782 points on a 5-point scale (95% CI 0.692-0.872; P<.001). This pattern was observed across all 6 categories of patient questions. The largest estimated differences were for empathy (mean difference 1.062, 95% CI 0.948-1.177) and actionability (0.992, 95% CI 0.877-1.106). Safety flag distributions were similar between CA and clinician responses, with major concerns rare in both groups (n=3, 0.7% each). Although CA responses were longer, additional analyses adjusting for word count did not indicate that response length explained the overall quality difference.
Conclusions: Scaffolded LLM-based systems may have value as adjunct tools for CGM review, patient education, and preconsultation preparation by supporting standardized explanatory tasks. However, these findings should be interpreted in light of the vignette-based design, restricted datasets, and a small clinician panel, and they do not establish suitability for autonomous therapeutic decision-making, medication adjustment, or unsupervised real-world use. Prospective validation in clinical workflows is needed before implementation.
Keywords: clinical evaluation; continuous glucose monitoring; conversational agent; diabetes care; large language model; patient-facing AI; retrieval augmented generation