bims-aukdir Biomed News
on Automated knowledge discovery in diabetes research
Issue of 2026–08–09
eleven papers selected by
Mott Given



  1. Sci Rep. 2026 Aug 05. pii: 24161. [Epub ahead of print]16(1):
      The global prevalence of diabetic retinopathy (DR) is increasing in parallel with the rising burden of diabetes, posing a substantial public health challenge. Current automated DR screening systems often lack lesion-level interpretable, robustness against class imbalance, and reliable uncertainty estimation, limiting their applicability in real-world clinical settings. The study suggests a clinical-guided deep learning framework of diabetic retinopathy (CG-DRNet) that considers lesion-conscious attention, adversarial data augmentation, and Bayesian uncertainty measurement to achieve reliable and explainable DR severity detection with objective to allow early-stage detection, particularly mild nonproliferative diabetic retinopathy (NPDR), and reciprocate clinical diagnostic procedures. The proposed framework uses a multi-task deep learning framework and lesion-aware attention network to explicitly predict microaneurysms, hemorrhages, exudates, and neovascularization. The system consists of a conditional generative adversarial network (CWGAN-GP) that is employed to overcome extreme imbalance in classes by generating clinically realistic images of minority classes of the fundus. Monte Carlo dropout is used to model Bayesian uncertainty to estimate predictive confidence, and an uncertainty-informed semi-supervised strategy of learning is used to enhance data efficiency. Evaluation of the framework is carried out on publicly available fundus image datasets APTOS 2019, Messidor-2 and Clinical metadata with the use of PyTorch 2.0.1. The proposed CG-DRNet reached 93.8% accuracy on APTOS 2019 and 91.2% on Messidor-2 with only 2.6% difference between generalization and the original model, macro F1-score of 0.891, quadratic weighted kappa of 0.912, referable DR detection AUC of 0.963 and expected calibration error of 0.034. The framework was shown to have 84.7% sensitivity when detecting Grade 2 + with 67 ms inference time and therefore has been shown to be viable in clinical use.
    Keywords:  APTOS 2019; Diabetic retinopathy; Fundus imaging; Monte Carlo dropout; Uncertainty quantification
    DOI:  https://doi.org/10.1038/s41598-026-58690-w
  2. Front Med Technol. 2026 ;8 1826143
      Diabetic Retinopathy or DR is one of the leading causes of blindness in the working age population across the world, making its early detection and accurate classification as one of the major challenges in healthcare and medical imaging. Over the years various approaches have been developed like conventional deep learning models to classify DR or grade DR according to its severity. However, it has been observed that some of the major challenges encountered by various approaches is the lack of a lightweight feature focused learning mechanism. In this study, the authors propose a Variational Autoencoder (VAE) with disentanglement factor (Beta) and Logistic Regression based framework to classify DR in both ways, binary and multilevel classification. The proposed architecture aims to extract the fine features and textures of the retinal images compress it into a latent vector via the encoder and then expand the image map into a reconstructed image which aids the Logistic Regression classifier to classify the images and distinguish the severity of the disease. The proposed model was evaluated on two well-known datasets, APTOS 2019 dataset and DDR dataset. In case of binary classification, the model achieved an accuracy of 98.64% on the APTOS 2019 dataset and a 97.83% accuracy when evaluated on the DDR dataset. On multilevel classification, the model recorded an accuracy of 97.80% and 97.23% on APTOS 2019 dataset and DDR dataset respectively. These findings highlight the potential of the proposed method as an accurate and effective tool for automated DR screening and severity grading.
    Keywords:  diabetic retinopathy; disentanglement factor; features focused mechanism; logistic regression classifier; severity grading; variational autoencoder
    DOI:  https://doi.org/10.3389/fmedt.2026.1826143
  3. Front Digit Health. 2026 ;8 1816806
       Background: Machine learning models used for diabetes risk prediction may encode age-related biases that reduce diagnostic accuracy for specific demographic groups. Adversarial debiasing with a gradient reversal layer (GRL) offers a theoretically principled approach to learning representations that are invariant to a protected attribute; however, its practical effectiveness under realistic conditions of subgroup imbalance in healthcare datasets has not been fully characterised.
    Research question: Does adversarial debiasing with a GRL improve age-equitable diabetes prediction, and how do its fairness effects vary across different data partitions?
    Methods: Adversarial debiasing was evaluated for age-bias mitigation in diabetes prediction using the publicly available Pima Indians Diabetes Database (n = 768). All eight dataset predictors were used; three age groups (<30, 30-50, and >50 years) were derived from age for fairness evaluation. An adversarial neural model with a gradient reversal layer was compared against a logistic regression baseline. Features were standardised using a scaler fitted on training data only. The train-test split was stratified by diabetes outcome. Overall performance metrics (accuracy, recall, ROC-AUC) and the recall parity gap across age groups were computed on a primary labelled test partition (n = 154); robustness was assessed across five independent random seeds (0-4).
    Results: On the primary test partition, the adversarial model improved recall for the smallest age group [>50 years: 0.5556 → 0.7778, +22.22 percentage points (pp)] while maintaining comparable overall discrimination (ROC-AUC: 0.7852 → 0.7896, +0.45 pp). However, the recall parity gap increased from 0.0996 to 0.2153 (+11.57 pp), reflecting a concurrent decline in recall for the <30-year group (-6.25 pp). Across five random seeds, the mean recall parity gap showed a modest mean reduction (0.3282 → 0.3033, -2.49 pp), but with high variability (SD > 0.27) exceeding the mean difference. The adversarial model reduced the fairness gap in three of five seeds, increased it in one, and produced no change in one.
    Conclusion: Adversarial debiasing can improve predictive recall for underrepresented demographic subgroups but does not guarantee consistent fairness improvements across data partitions, particularly when subgroup sample sizes are small. Multi-seed evaluation is essential for reliable fairness assessment; single train-test splits are insufficient.
    Keywords:  adversarial debiasing; age bias; algorithmic fairness; diabetes prediction; gradient reversal layer; healthcare machine learning; partition dependency; recall parity
    DOI:  https://doi.org/10.3389/fdgth.2026.1816806
  4. Front Med (Lausanne). 2026 ;13 1871693
       Background: Patients with coexisting type 2 diabetes mellitus (T2DM) and hypertension (HTN) face a synergistically elevated risk of major adverse cardiovascular events (MACE). Evidence for prediction models developed specifically in established T2DM-HTN comorbidity population remains limited.
    Objective: To methodologically explore and preliminarily evaluate an interpretable machine learning framework for 1-year MACE prediction in hospitalized patients with coexisting T2DM and HTN using routine clinical data.
    Methods: This retrospective study included 1,054 hospitalized patients with T2DM and HTN, of whom 249 (23.6%) experienced MACE during 1-year follow-up. The dataset was randomly divided into training (60%), validation (20%), and independent test (20%) cohorts using stratified sampling. LASSO regression was applied for feature selection from 69 clinical variables. Four algorithms, including logistic regression, random forest, support vector machine, and XGBoost, were developed and compared. Model performance was assessed using discrimination, calibration, and clinical utility metrics. SHapley Additive exPlanations (SHAP) were used to interpret the final model.
    Results: LASSO identified six stable predictors: HbA1c, age, hypertension duration, cystatin C (CysC), T2DM duration, and carotid intima-media thickness (CIMT). Sex was additionally incorporated based on clinical relevance. Multivariable logistic regression showed that HbA1c, age, hypertension duration, T2DM duration, CysC, and CIMT were associated with 1-year MACE risk, whereas sex was not statistically significant. Logistic regression showed the best relative balance between discrimination, calibration, and simplicity on the validation set, although learning curves indicated limited incremental improvement with increasing training sample size. After isotonic regression recalibration, the final logistic regression model achieved an ROC-AUC of 0.828, a PR-AUC of 0.656, and a Brier score of 0.116 on the independent test set. Decision curve analysis indicated potential clinical net benefit. SHAP linked model predictions to glycemic burden, aging, cumulative disease exposure, renal-related risk, and subclinical atherosclerosis.
    Conclusion: An interpretable logistic regression model based on seven routine clinical variables showed relatively good internal performance for predicting 1-year composite MACE risk in hospitalized patients with coexisting T2DM and HTN. CysC provided additional prognostic information beyond its conventional role as a renal filtration marker, although this association should be interpreted as prognostic rather than causal. External validation is required before the model can be considered for clinical decision support.
    Keywords:  SHAP interpretability; cystatin c; hypertension; machine learning; major adverse cardiovascular events; type 2 diabetes
    DOI:  https://doi.org/10.3389/fmed.2026.1871693
  5. Diabetes Res Clin Pract. 2026 Aug 05. pii: S0168-8227(26)00405-5. [Epub ahead of print] 113485
      This scoping review synthesized clinical validation evidence of artificial intelligence (AI) algorithms for diabetes-related foot ulcer (DRFU) detection and assessment and identified factors influencing AI performance in real-world settings. A systematic search was conducted in PubMed, MEDLINE, CINAHL, Scopus, and Google Scholar, following the Arksey and O'Malley framework and reported using PRISMA-ScR. Eligible studies involved adult patients with diabetes and foot ulcers, utilized learning-based AI models, and reported clinical validation outcomes. Eleven studies published between 2020 and 2026 were included from eight countries. Diagnostic performance varied across studies, with sensitivity of 91-100%, specificity of 20-96.8%, and intraclass correlation coefficients of 0.825-0.998 for wound measurement reliability. AI systems reduced manual area overestimation by 13.4-25.2%. Influencing factors were mapped across three NASSS framework domains: technological factors including image quality and algorithmic misclassification, adopter-level barriers including digital literacy limitations, and organisational system-level constraints including infrastructure instability and data privacy concerns. These findings suggest early-stage evidence of promising AI diagnostic performance. However, the evidence base remains limited by small sample sizes, heterogeneous designs, and the inability of current systems to assess deeper wound features. Prospective multi-centre studies with standardised protocols are needed before routine clinical adoption can be recommended.
    Keywords:  Deep Learning; Human-AI comparison; Influencing factors; Machine learning; Metrics performance; Real-world implementation
    DOI:  https://doi.org/10.1016/j.diabres.2026.113485
  6. Front Endocrinol (Lausanne). 2026 ;17 1933440
      
    Keywords:  clinical decision support; continuous glucose monitoring; diabetes mellitus; digital health; implementation science; machine learning; telemedicine
    DOI:  https://doi.org/10.3389/fendo.2026.1933440
  7. Front Endocrinol (Lausanne). 2026 ;17 1805548
       Introduction: Type 2 diabetes mellitus (T2DM) is among the most rapidly increasing metabolic disorders worldwide. Membrane proteins, integral components of biological membranes, are pivotal in insulin signal transduction and significantly contribute to the pathogenesis of T2DM. However, the systematic investigation of membrane proteins linked to T2DM remains insufficient.
    Methods: The T2DM dataset and membrane protein-related genes were sourced from the GEO database and the Uniprot website, respectively. Bioinformatics methodologies, including Gene Ontology (GO) analysis, Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analysis, and protein-protein interaction (PPI) network analysis, were employed to assess the differentially expressed membrane protein-related genes between the normal control group and the T2DM group. Subsequently, machine learning algorithms, including Gaussian Mixture Model (GMM), Random Forest (RF), and Support Vector Machine (SVM), were utilized to identify hub genes. Following this, a clinical diagnostic model was constructed, and the receiver operating characteristic (ROC) curve was plotted. Candidate genes were further examined in palmitic acid (PA)-induced NES2Y and HepG2 cell models and in high-fat diet/streptozotocin-induced T2DM mice. Finally, Functional effects were assessed by gene knockdown, reverse transcription quantitative PCR (RT-qPCR), western blotting (WB), immunofluorescence (IF), co-immunoprecipitation, histological examination, and lipid staining.
    Results: A total of 1,856 DEGs were identified between T2DM and control samples, including 599 upregulated genes and 1,257 downregulated genes, among which 42 were membrane proteinrelated T2DM DEGs. Machine learning algorithms were then applied to pinpoint two target genes: ICAM1 and EZR, the latter of which encodes Ezrin, a membrane-cytoskeleton linker protein. The ROC curve analysis showed that the diagnostic model exhibited strong predictive capability, with an AUC value of 0.95. PA treatment increased lipid accumulation and EZR/ICAM1 mRNA and Ezrin/ICAM1 protein expression in the cell models. Ezrin co-immunoprecipitated with the insulin receptor. In PA-treated HepG2 cells, ICAM1 or EZR knockdown increased p85α and AKT phosphorylation and GLUT4 protein expression. Ezrin and ICAM1 were also elevated in the livers of T2DM mice.
    Discussion: ICAM1 and EZR may serve as potential diagnostic biomarkers and candidate therapeutic targets associated with T2DM. Furthermore, ICAM1 and EZR may be associated with insulin resistance by reducing glucose transport efficiency, potentially through modulation of the PI3K-AKT signaling pathway.
    Keywords:  biomarker genes; diagnostic model; differentially expressed genes; membrane protein-related genes; type 2 diabetes mellitus
    DOI:  https://doi.org/10.3389/fendo.2026.1805548
  8. Int J Popul Data Sci. 2026 ;11(5): 3567
       Objective: To develop and evaluate a Transformer-based temporal representation learning framework for predicting incident diabetes using longitudinal laboratory data, and to compare its performance with state-of-the-art pretrained models and traditional machine-learning baselines.
    Approach: We deterministically linked a cardiac registry cohort (2015-2019) in Alberta, Canada, and retrieved three years of laboratory tests preceding each patient's diabetes diagnosis. Each laboratory event was modelled as a triplet consisting of test type, test value, and time gap with the previous test, to preserve temporal structure. Our Transformer architecture learned contextual embeddings that capture both intra-test dynamics and inter-test interactions, indicating the onset of diabetes. Performance was benchmarked against two pretrained Transformer models (Moment, PatchTST) and an XGBoost baseline trained on aggregated test summaries (mean, standard deviation, minimum, maximum).
    Results: The final cohort included 30,462 patients and 19 laboratory test types relevant to diabetes. The proposed method achieved the highest predictive accuracy (AUC = 0.92), outperforming Moment (0.80), PatchTST (0.74), and XGBoost (0.88). It also demonstrated superior sensitivity (0.70) and positive predictive value (0.777) while maintaining high specificity (0.938) and negative predictive value (0.911). Conclusion: Modelling raw laboratory trajectories with a dedicated temporal Transformer substantially improves diabetes prediction compared with both pretrained sequence models and aggregation-based baselines.
    Implications: This framework provides a generalizable pathway for leveraging routine laboratory data in early disease detection. It highlights the clinical value of sequence-level modelling and supports the integration of temporal representation learning into population-level surveillance and decision-support systems.
    DOI:  https://doi.org/10.23889/ijpds.v11i5.3567
  9. Ophthalmic Epidemiol. 2026 Aug 05. 1-8
       PURPOSE: Diabetic retinopathy remains a leading cause of blindness in the United States. Autonomous artificial intelligence (AI) systems for screening have received US Food and Drug Administration authorization, and a dedicated Medicare billing code for autonomous point-of-care retinal screening was introduced in 2021. This study aimed to characterize observed utilization, billing National Provider Identifier (NPI) uptake, provider-type distribution, and geographic spread of autonomous AI screening in Medicare from 2021 to 2023.
    METHODS: This study performed a retrospective descriptive analysis of the 2021-2023 Medicare Physician & Other Practitioners-by Provider and Service public-use files. Current Procedural Terminology (CPT) 92229 was the primary focus, with CPT 92227 and 92228 included as contextual comparators.
    RESULTS: Autonomous AI screening increased nearly tenfold, from 143 services in 2021 to 1,427 in 2023. Unique billing NPIs increased from 5 to 39. Optometry accounted for 57.7% of services, followed by family practice (13.9%), endocrinology (12.5%), and internal medicine (11.9%). The number of states with observed use increased from 3 to 16, and urban providers accounted for 96.6% of observed services. In 2023, observed use corresponded to approximately 15.3 services per 100,000 Medicare beneficiaries identified as having diabetes using fee-for-service claims. Estimated Medicare payments captured in the file increased from $3,538 to $42,889.
    CONCLUSION: Autonomous AI screening showed rapid early growth in Medicare Part B after introduction of code 92229, with expansion across billing NPIs, provider types, and states. Absolute use nonetheless remained small, indicating early billing uptake rather than meaningful population penetration.
    Keywords:  Artificial intelligence; Medicare claims; autonomous screening; diabetic retinopathy screening; health services research
    DOI:  https://doi.org/10.1080/09286586.2026.2714988
  10. J Med Internet Res. 2026 Jul 31. 28 e98519
       Background: Continuous glucose monitoring (CGM) is central to diabetes care, but explaining CGM patterns consistently and empathetically remains time-intensive in clinical practice. Large language model (LLM)-based systems may support patient-facing interpretation of CGM data, but evidence remains limited for retrieval-grounded tools evaluated against clinician-authored responses in counseling scenarios. The system was intended for CGM interpretation and communication support rather than autonomous therapeutic decision-making.
    Objective: This study aimed to evaluate whether a retrieval-grounded LLM-based conversational agent (CA) could support patient understanding of CGM data and preparation for diabetes consultations by generating responses to questions arising during CGM-informed diabetes counseling, with quality comparable to clinician-authored responses.
    Methods: We developed a scaffolded LLM-based CA for CGM interpretation and diabetes counseling support. The system was designed to provide plain-language explanations of CGM patterns and responses to diabetes management questions while avoiding directive or individualized medical advice, such as recommending medication initiation, dose adjustment, or regimen changes. Around 12 CGM-informed cases, each comprising a deidentified CGM trace, a synthetic patient vignette, and accompanying CGM visual materials, were constructed from using available clinical datasets. Between October 2025 and February 2026, 6 senior UK diabetes clinicians each reviewed 2 assigned cases and answered 24 questions (12 per case). In a source-masked multirater evaluation, each CA-generated and clinician-authored response was independently rated by 3 clinicians on 6 quality dimensions, including clinical accuracy, guideline adherence, actionability, personalization, communication clarity, and empathy. Safety flags and perceived source labels were also recorded. The primary analysis used linear mixed effects models with random intercepts for case and rater.
    Results: A total of 288 unique responses (144 CA and 144 clinician responses) were evaluated, generating 864 ratings. CA-generated responses received higher quality scores than clinician-authored responses under controlled vignette-based conditions, with mean scores of 4.37 (SD 0.57) versus 3.58 (SD 0.90) and an estimated mean difference of 0.782 points on a 5-point scale (95% CI 0.692-0.872; P<.001). This pattern was observed across all 6 categories of patient questions. The largest estimated differences were for empathy (mean difference 1.062, 95% CI 0.948-1.177) and actionability (0.992, 95% CI 0.877-1.106). Safety flag distributions were similar between CA and clinician responses, with major concerns rare in both groups (n=3, 0.7% each). Although CA responses were longer, additional analyses adjusting for word count did not indicate that response length explained the overall quality difference.
    Conclusions: Scaffolded LLM-based systems may have value as adjunct tools for CGM review, patient education, and preconsultation preparation by supporting standardized explanatory tasks. However, these findings should be interpreted in light of the vignette-based design, restricted datasets, and a small clinician panel, and they do not establish suitability for autonomous therapeutic decision-making, medication adjustment, or unsupervised real-world use. Prospective validation in clinical workflows is needed before implementation.
    Keywords:  clinical evaluation; continuous glucose monitoring; conversational agent; diabetes care; large language model; patient-facing AI; retrieval augmented generation
    DOI:  https://doi.org/10.2196/98519
  11. Front Endocrinol (Lausanne). 2026 ;17 1895366
       Background: Diabetes-related foot disease requires timely recognition of neuropathic risk, ulceration, infection, ischemia, offloading needs, and recurrence risk. Publicly accessible large language models (LLMs) may provide patient-facing information, but reproducible prompt construction for benchmarking such outputs remains insufficiently characterized.
    Objective: This study aimed to develop and apply a domain- and source-balanced prompt framework for benchmarking patient-facing diabetic foot information generated by publicly accessible LLMs under default single-turn public-interface conditions.
    Methods: A 24-item benchmark prompt set was generated using a domain- and source-balanced framework incorporating public-query sources and guideline-derived decision-critical content. Six clinical domains were crossed with four source categories: Google Trends, Baidu Zhidao, a PubMed-indexed Chinese diabetic foot guideline, and PubMed-indexed international diabetic foot guidelines. Each prompt was submitted once to GPT-5.5 Thinking, DeepSeek-V4, Gemini 3.1 Pro, Grok 4.3, and Qwen3.6-Max-Preview, yielding 120 responses. Response quality was assessed using DISCERN, EQIP, and GQS; visible transparency-related features were evaluated using JAMA benchmark criteria; readability was assessed using six formulas; and an exploratory potential clinical-risk flag (PCF) screened for overt short-term harm signals. Formal claim-level factual-accuracy review, guideline-concordance adjudication, and hallucination-frequency analysis were not performed.
    Results: Significant metric-specific differences were observed across models. Grok 4.3 recorded the highest observed mean DISCERN, EQIP, GQS, and JAMA-based visible transparency-related scores, whereas DeepSeek-V4 showed the lowest observed mean scores for several readability-grade metrics and the highest mean FRES. Visible transparency-related scores remained low across models. No response was rated as PCF 1 or PCF 2. No response met all predefined readability targets.
    Conclusions: The proposed prompt framework provides a structured basis for public-interface LLM benchmarking in diabetic foot education. Default responses showed metric-specific variation, limited visible transparency, and inadequate readability, and should not be relied upon independently for high-risk diabetic foot decision-making without clinician oversight.
    Keywords:  benchmarking; diabetic foot; domain- and source-balanced prompt framework; large language models; patient-facing information
    DOI:  https://doi.org/10.3389/fendo.2026.1895366