Enterprise RAG deployments face a critical reliability gap: while LLMs satisfy 80% of individual constraints, only 26.8% of responses meet all requirements simultaneously, revealing a 57-point orchestration gap. Existing benchmarks assume clean retrieval with simple queries, failing to capture production conditions whe...
Chain-of-Thought (CoT) improves the reasoning ability of Large Language Models (LLMs) but incurs substantial computation and context costs. Existing methods either lose intermediate information through hard pruning or lack a principled criterion for continuous compression. We present A*-Thought-V2, a geometric dynamics...
Xiao-An Xu, Si-Yuan Liu, Shuo Wang et al.· 0 citations
TrustDABench is introduced, a benchmark that operationalizes two diagnostic questions of LLM reliability and robustness and suggests that stronger evidence-boundary recognition and representation-invariant reasoning are still needed for reliable structured-data analysis.
Boshen Shi, Yize Liu, Chen Zhao et al.· 0 citations
This work advocates for Joint Online-Offline Fine-Tuning as a superior paradigm that breaks the convention of restricting offline data to SFT and online data to RFT, and provides the first comprehensive survey focusing specifically on the synchronization of data provenance.
Tai-Hang Zhen, Guang Yang, Chen-Zhang Li et al.· Proceedings of the Thirty-Fi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.