SciFigBench is introduced, a diagnostic VLM benchmark for scientific figure understanding that jointly evaluates perception, reasoning, and behavioral reliability under uncertainty and proposes the Admittance-Resistance-Inductance (A-R-I) framework to evaluate whether models acknowledge insufficient evidence, resist mi...
Paul Osemudiame Oamen, Owusu-Banahene Osei, Ananya Mukherjee et al.· 0 citations
The results show that annotation reporting in NLP has improved over time but remains uneven, and they establish a scalable framework and bare-minimum reporting recommendations for making human annotation more reliable, reproducible, and interpretable.
M. Kunilovskaya, Gagan Bhatia, Lisa Sophie Albertelli et al.· arXiv.org· 0 citations
DaEdiTikZ is introduced, the first large-scale dataset of revision-derived scientific figure edits, constructed by mining 391K plausible TikZ edit pairs from arXiv, GitHub, and TeX SE and inferring 781K directed edit instructions with a VLM conditioned on rendered figures and TikZ code.
C. Greisinger, Zhixue Zhao, Steffen Eger· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.