J Gen Intern Med. 2026 Oct 02.
BACKGROUND: The use of artificial intelligence (AI) models as reviewers of scientific content raises concerns about potential biases related to author identity and about the reproducibility of their evaluations. We assessed whether AI-based reviewers exhibit gender or geographic bias and evaluated the reproducibility of their scoring of scientific abstracts.
METHODS: We randomly selected 10 general internal medicine journals indexed in the Journal Citation Reports (impact factor ≥ 1.5). For each journal, five original research abstracts were retrieved in April 2026 from the Web of Science. Each abstract (n = 50) was assigned to four fictional author identities (African female, African male, American female, American male). Abstracts were independently evaluated twice by two AI models (ChatGPT and Claude), scoring quality/novelty/acceptance on a 0-10 scale, yielding 800 evaluations. Bias was assessed using multivariable ordinal logistic regression, and reproducibility using percent agreement and Fleiss' kappa.
RESULTS: For ChatGPT, quality and novelty scores were identical across identities (median [IQR] = 7 [1] and 5 [2], respectively), with minimal variation in acceptance (median [IQR] = 6-7 [1-2]). For Claude, all scores were identical across identities (quality = 7 [1], novelty = 5 [2], acceptance = 6 [2]). Multivariable analyses showed no overall association between author identity and scores, with the exception of acceptance for Claude at the global level, although no individual comparisons were statistically significant. Reproducibility was high for both models, with percent agreement > 0.98 and substantial/excellent agreement (κ = 0.74-0.89).
CONCLUSIONS: In this controlled setting, we detected no consistent evidence of gender or geographic bias, and reproducibility was high, although the restricted score range may have contributed to the observed agreement. Further studies are needed to assess the validity of LLM-based evaluation and compare it with human peer review.
Keywords: AI; ChatGPT; Claude; abstract; artificial intelligence; bias; country; gender; geographic; large language model; peer review; reliability; reproducibility