Evaluating large language models as grant reviewers: a comparative study of prompt engineering strategies
Abstract
Grant application review is resource-intensive and subject to inter-rater variability. Large language models (LLMs) may augment this process, but their reliability in grant evaluation remains unexplored. This exploratory pilot study compared LLM-generated grant reviews to human expert reviews across three prompt engineering strategies: zero-shot (no examples), few-shot (multiple training examples), and few-shot with strict scoring instructions. Three commercial LLMs (OpenAI GPT-5 Nano, Google Gemini 2.5 Flash, xAI Grok 4 Fast Reasoning) each reviewed three de-identified faculty seed grant applications ($5,000 awards) under each condition, generating 63 LLM reviews compared against 12 human expert reviews. Among the three conditions tested, few-shot prompting produced scores closest to human reviewers (mean difference: −2.57 points on a 100-point scale, p = 0.671, Glass's Δ = −0.19). Zero-shot prompting showed moderate optimism bias (+4.55 points, p = 0.347) and strict instructions produced substantial under-scoring (−10.68 points, p = 0.097). LLMs demonstrated 49%–59% lower score dispersion than humans (49% in Experiment 1 and 59% in Experiments 2–3), which likely reflects range restriction. Chance-corrected agreement between LLM and human funding recommendations was at or below chance in every condition (Cohen's κ : −0.11–0.05), indicating LLM funding recommendations provided no meaningful signal beyond chance. These preliminary findings suggest prompt engineering affects LLM scoring behavior in this sample, and that LLMs cannot yet reliably replicate human grant review scores or funding recommendations but may provide adjunctive data for consideration.