Jun 2026
An Empirical Study of Security Calibration in Large Language Models for Code
The first large-scale empirical study of security calibration in LLM-generated code is presented, showing that although architectural gating improves calibration on controlled benchmarks, calibration deteriorates in realistic repository-level settings, increasing the risk of high-confidence vulnerable outputs.
Mohammed Latif Siddiq, Md Nafiu Rahman, Joanna C. S. Santos
· arXiv.org · 0 citations