LLCE-SQL: Data-Efficient Text-to-SQL Training via Logic-Linking Driven Collaborative Evolution
Abstract
Training high-performance Text-to-SQL models typically requires massive synthetic datasets, yet existing augmentation methods suffer from semantic drift, structural homogenization, and limited complexity coverage, leading to diminishing returns at scale. This paper proposes LLCE-SQL, a Logic-Linking driven Collaborative Evolution framework that generates high information-density training samples with robust logical consistency and controllable complexity. Central to the framework is the Logic-Linking mechanism, which establishes an atomic-level isomorphic mapping between SQL syntactic structures and natural language semantic expressions, enforcing strict consistency constraints to ensure precise cross-modal alignment. For sample evolution, we propose the SQL-MLU (SQL Meaning Logical Units) hierarchical evolution strategy with three stages: atomic attribute perturbation, intra-slice logical scaling, and macro-topological reconstruction. We also explore an optional Logic-Linking graph based Evidence-Driven CoT (ED-CoT) module that delivers explicit schema and logic traces. In addition, LLCE-SQL incorporates a multi-dimensional screening mechanism using logical viability, structured semantic entropy, and global state awareness to drive evolution toward enhanced logical depth and expressive diversity. With only 8 K training samples, LLCE-SQL achieves 86.1% execution accuracy on Spider Dev and attains a high data-efficiency score, supporting the effectiveness of logic-linked sample evolution.