Scientific literature synthesis agents increasingly rely on proprietary online services, limiting reproducibility, privacy, and offline deployment. To address this challenge, we introduce SciLENS Scientific Localized Evidence Navigation and Synthesis), a fully local autonomous agent framework operating on a dual-tier i...
Le-Qi Zheng, Jin-Bo Su, Yu-Ying Li et al.· 1 citation
CapGeo-Bench is proposed, a benchmark of 4,641 high-quality figure-caption pairs equipped with a fine-grained keypoint-based evaluation metric that provides high-quality captions consistently and substantially boosts performance of MLLMs, empirically validating the visual perception bottleneck in geometric reasoning.
Yu-Ying Li, Si-Yi Qian, Hao Liang et al.· 5 citations
SLATE (Slide-based Learning Assessment for Teaching Effectiveness), the first benchmark that evaluates AI-generated language teaching slides through instructional effectiveness and learner knowledge acquisition, reveals a dissociation between artifact quality and instructional effectiveness.
Jing-Zhuo Wu, Jia-Jun Zhang, Liu Yi et al.· 0 citations
Large language models (LLMs) have shown strong potential for translating natural-language (NL) requirements into PL/SQL programs, attracting increasing attention from the database community. However, existing NL-to-PL/SQL efforts primarily focus on directly generating PL/SQL from complete NL requirements. In practice,...
Hang Zhang, Chao-Kun Wang, Yun Pan et al.· 0 citations
Gradient-Aligned Reward (GAR), which operates in the policy's own gradient space: truncated backpropagation through the output projection layer extracts a compact gradient vector for each rollout, and cosine similarity with an expert-anchor gradient yields a dense, reasoning-aware reward with less than 9% wall-clock ov...
Le-Qi Zheng, Jin-Bo Su, Fang Niu et al.· 2 citations
OrderProbe is introduced, a deterministic benchmark for structural reconstruction using fixed four-character expressions in Chinese, Japanese, and Korean, which have a unique canonical order and thus support exact-match scoring.
Juris Policy Optimization (JPO), a post-training framework for structured legal reasoning in Chinese criminal judgment prediction, is proposed and experiments show that JPO consistently improves both judgment prediction and reasoning quality over supervised fine-tuning and reinforcement learning baselines.
Zhao-Lu Kang, Yan-Tao Liu, Tailong Luo et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.