Benchmarking Prompt Engineering Against Fine-Tuning for Multi-Label Vulnerability Detection in Solidity Smart Contracts: An Empirical Study
Abstract
Large Language Models (LLMs) are increasingly being deployed for smart contract security, yet a fundamental question remains unresolved for practitioners: when confronted with the realistic, multi-label setting where a single contract may harbor several concurrent vulnerabilities, which deployment strategy is more effective — prompt engineering or fine-tuning? This paper provides a rigorous empirical benchmark to answer that question. We evaluate five prompt engineering strategies (Zero-shot, Chain-of-Thought, Structured Chain-of-Thought, Step-back, and Few-shot) across five state-of-the-art open-source LLMs (7B – 24B parameters), and compare them against the same models fine-tuned with QLoRA on a standardized multi-label dataset of real-world Ethereum contracts. Our benchmark yields three specific, quantified findings. First, the performance gap is large and consistent: the best prompt engineering configuration achieves a macro F1-score of 0.311, while fine-tuned models reach 0.79, a gap that holds across all five model architectures. Second, scale alone did not compensate for the absence of task-specific training in this setting: a 24B-parameter model under the best prompting strategy was outperformed by a 7B fine-tuned model, suggesting that the benefit of scale was not sufficient to close the paradigm gap on this structured multi-label task. Third, advanced reasoning strategies (Chain-of-Thought, Structured CoT, Step-back) provided negligible improvement over the zero-shot baseline for this task, suggesting that the evaluated prompt-only approaches struggled with structured multi-label output requirements rather than with reasoning quality per se. These results hold for the five open-source model architectures, five prompting strategies, and Slither-derived benchmark labels used in this study. These findings provide actionable guidance for researchers and practitioners designing LLM-based security tools for Solidity smart contracts.