A Fairness–Utility Evaluation Framework for Assessing Large Language Models (LLMs)
Abstract
Large language models are increasingly used in contexts where their outputs can affect people directly, including hiring, admissions, and lending. This growing role makes it important to consider not only how well these models perform, but also whether their behavior is fair. Although many fairness metrics, bias benchmarks, and mitigation techniques have been proposed, there is still limited guidance on how fairness and predictive performance can be evaluated together through a consistent procedure. This study presents a five-stage Fairness–Utility Evaluation Framework that brings together dataset selection and preparation, model assessment, fairness evaluation, utility evaluation, and trade-off analysis and reporting. Rather than introducing a new metric, the framework organizes these elements into a structured and reproducible evaluation process. Its application is illustrated using LLaMA-2–7B, CrowS-Pairs, and WinoBias to show how the five stages can be followed in practice; the example is intended to demonstrate the procedure rather than provide new experimental results. Documentation is incorporated throughout the process so that evaluation choices, measures, and outcomes can be traced and reviewed. The resulting framework provides researchers and professionals with a structured way to consider fairness and performance together when evaluating an LLM for potential deployment.