GDP-RAG is proposed, a plan-based framework that targets only the information delta based on three simple design choices, and achieves the highest accuracy among all compared systems while maintaining a cost-of-pass of 0.51.
Wei-Chieh Chou, Xuan-Bo Chen, Jian-Zhe Lin et al.· arXiv.org· 3 citations
Large Audio Language Models (LALMs) can follow diverse instructions to synthesize speech in specified styles. However, complex instructions that require simultaneous control over pitch dynamics, speaking rate, and emotional tone often exceed what a single-pass generation can faithfully realize. While recent reasoning m...
Chee-En Yu, Yi-Cheng Lin, Sung-Feng Huang et al.· 0 citations
Agentic retrieval-augmented generation (RAG) enables language models to adapt retrieval based on previously retrieved evidence, but it remains unclear whether such adaptive orchestration consistently outperforms well-designed static pipelines. We conduct a controlled comparison of agentic and static RAG for Taiwanese h...
Kai-Hsin Chen, Wei-Yuan Chen, Xuan-Bo Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.