Toward Cost-Efficient Automated Requirements Traceability with Large Language Models
Abstract
Automated requirements traceability is a critical activity in software engineering, supporting impact analysis, verification and validation, and regulatory compliance. As modern software-intensive systems grow in scale and complexity, maintaining trace links manually becomes increasingly impractical. Recent work has shown that Large Language Models (LLMs) can perform requirements traceability effectively; however, applying powerful LLMs exhaustively to all artifact pairs incurs substantial monetary cost and often depends on closed, proprietary models, collectively limiting scalability, transparency, and practical adoption. This paper investigates costefficient alternatives to exhaustive LLM-based traceability. We propose LightTraceLLM, a framework of two ensemble-based approaches: LightTraceLLM_MV, which aggregates predictions from multiple lightweight LLMs using majority voting, and LightTraceLLM_HE, which selectively invokes a high-end LLM only for ambiguous cases. We also compare against LiSSA, a prior cost-efficient embedding-based approach, and TraceLLM, a high-end LLM baseline, to provide a comprehensive picture of the cost-performance tradeoff space. An extensive empirical evaluation across multiple benchmark datasets shows that the proposed approaches achieve traceability performance comparable to the high-end baseline while reducing monetary inference cost by approximately $\mathbf{6 7 - 9 9 \%}$, depending on the method and dataset.