Reinforcement learning over a frozen reasoner has become a common recipe for teaching a policy which external tools to invoke. We show that this recipe becomes structurally mismatched in specialist scientific settings where the complete tool-subset space is enumerable. There, a small set of recurring computational capa...
Hao-Yue Liu, Xiao-Yu Ma, Ye-Heng Chen et al.· 0 citations
This paper introduces SEPO (Structural, Evidence-grounded Prompt Optimization), a multi-trajectory prompt optimiser centred on edit-effect lineage feedback that makes prompt optimisation addressable, attributable, and actionable.
Xiao-Yu Ma, Hao-Yue Liu, Yiwen Li et al.· 0 citations
HN-CLIP is introduced, which uses the text encoder's own text-text geometry to construct per-negative adaptive similarity margins, and improves all six tested fine-tuning frameworks on the in-domain benchmarks and reaches the strongest full-data baseline with only 20% of the training data.
Hao-Yue Liu, Ye-Heng Chen, Zhi-Chao Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.