The ProgramTab framework is proposed, which guides LLMs employing in-context learning to perform tabular data preprocessing with Python code, as well as the momentous contents extraction with row and column extraction and SQL generation, demonstrating that the ProgramTab framework effectively deals with table-based reasoning tasks and outperforms all LLM-based baselines.
Abstract
Table-based reasoning with large language models (LLMs), which requires reasoning based on natural language questions and structured tabular data, has gained widespread attention. However, a series of issues still constrain the application of this task. The previous approaches suffered from significant performance degradation when faced with large tables due to the difficulty of long text modeling and the limitation of input length for LLMs. The text-to-SQL approach is used to efficiently extract key information from tables and generate smaller sub-tables. However, tabular data, especially web tables, often lack the necessary structure and consistency, making them unsuitable for performing mathematical logic operations using SQL queries. We propose the ProgramTab framework, which guides LLMs employing in-context learning to perform tabular data preprocessing with Python code, as well as the momentous contents extraction with row and column extraction and SQL generation. The experiment results on table reasoning datasets demonstrate that the ProgramTab framework effectively deals with table-based reasoning tasks and outperforms all LLM-based baselines.
LIMIT(Less Is More for Instruction Tuning in Text-to-SQL), a data-centric framework that demonstrates strong database reasoning can emerge from an extremely compact training set when examples are strategically selected, is proposed, suggesting that careful data curation, rather than scale, is the key to efficient Text-...
Hao-Yuan Ma, Heng-Wei Liu, Linjuan Wu et al.· 0 citations
SPOC-SQL is proposed, which decomposes Text-to-SQL into four sequential subtasks following standard SQL execution logic and designs stage-specific optimization strategies for the model to learn key decisions, with the objective of enhancing structured decision-making during query construction.
Yingnan Chen, Chun Ding, Tian-Shi Xu et al.· 0 citations
This survey conducts a comprehensive review of the literature on table mining with Large Language Models, exploring the challenges unique to this domain, such as heterogeneous table structures, contextual ambiguity, and domain-specific knowledge requirements.
Mingyue Cheng, Qingyang Mao, Qi Liu et al.· ACM Computing Surveys· 12 citations
Experiments on WikiTQ and SLQA show that localization is particularly effective for lookup and local reasoning questions, while adaptive selection between localized and full-table reasoning achieves the best overall performance, highlighting that long-table QA requires deciding not only how to localize, but also when t...
Experimental evaluation on 300 realistic pattern mining tasks demonstrates consistent improvements in algorithm configuration accuracy, parameter compliance, and dataset specification correctness across zero-shot, one-shot, and few-shot settings, highlighting the effectiveness of inference-time domain grounding for ena...
Madhavi Palla, Uday Kiran Rage, Arjun Chakravarthi Pogaku· International Journal of Dat...· 0 citations
Table Question Answering (TableQA) requires reasoning over natural language questions and structured tables, and remains challenging due to noisy evidence and complex multi-step reasoning. Recent Large Language Model (LLM)-based approaches typically adopt decomposition–reasoning–validation pipelines that combine Chain-...
Zhen Yang, Zi-Wei Du, Ming-Han Zhang et al.· ACM Transactions on Informat...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.