Humans often solve spatial problems by mentally simulating visual transformations. In contrast, conventional vision-language models (VLMs) reason primarily through language. We investigate whether VLMs can solve spatial problems by reasoning with both text and generated visual states. To this end, we introduce WM-VLM,...
Yuheng Zha, Yi-Lei Wang, Qiyue Gao et al.· 0 citations
The proposed NumCache, which compresses SEC filings into KV caches initialized from numerically dense regions and trained directly on financial QAs, is evaluated, which highlights cache-based retrieval with number-preserving representations as an effective approach for long-context financial QA.
Eftychia Makri, Peiwen Li, Yidong Jiang et al.· Proceedings of the 32nd ACM...· 0 citations
Large Language Models (LLMs) are increasingly deployed in financial applications, particularly for interpreting U.S. Securities and Exchange Commission (SEC) filings. However, financial QA over these filings is challenging, as they are extremely long, numerically dense, and often require cross-document reasoning. Exist...
Eftychia Makri, Peiwen Li, Yidong Jiang et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.