Hot off the Press: Pareto-Optimal Fronts for Benchmarking Symbolic Regression Algorithms
Abstract
Symbolic regression (SR) is commonly benchmarked using relative Pareto analysis, where an algorithm is evaluated according to whether it dominates the other methods included in the comparison. While useful, this perspective does not reveal the true attainable limits of short, interpretable expressions, and conclusions may change depending on which competing methods are selected. In this work, we advocate benchmarking with absolute Pareto-optimal (APO) fronts instead. We construct APO fronts for 34 datasets from SRBench by exhaustively searching over short symbolic expressions and by fitting numerical constants with eight commonly used local optimization methods. The resulting fronts provide dataset-specific baselines for the best accuracy-complexity trade-offs attainable within a fixed primitive set. Comparing against the SR methods reported in SRBench shows that many current algorithms remain far from these fronts, especially in the regime of short expressions. We also find that the final fronts are relatively stable across numerical optimizers, suggesting that structural search is a more important bottleneck than constant fitting. These APO fronts provide a reusable benchmarking asset for the evolutionary computation community and a more stable reference point for measuring progress in SR.