Aviation knowledge question answering cannot directly assess the operational effectiveness and safety compliance of large language models throughout aircraft emergency procedures. We introduce AeroCopilotBench and its executable cockpit environment, ACOE, which define state-transition rules, task goals, and trajectory-...
Yu-Chen Yuan, Zheng-Huang Wu, Yuan-Gan Li et al.· 0 citations
Autoregressive (AR) models suffer from local greediness, while diffusion language models (DLMs) often lack the strict causal structure required for reasoning. To combine the advantages and overcome the drawbacks of the dual, we propose Causal Latent Revision (CaLR), a framework that reformulates reasoning as constraine...
Wei Cai, Jian Zhao, Yu-Chen Yuan et al.· 0 citations
QEncodeBench tasks large language models with encoding classical constraint problems as phase oracles and scores the generated circuits with an adversarially self-validated verifier that decides full solution-set equivalence up to a global phase, with ancillas restored and resource budgets enforced.
Xu-Jun Che, Han-Han Wu, Yu-Chen Yuan et al.· 1 citation
A safety-gated evaluation framework in which a trajectory succeeds only when all task goals are achieved without violating any hard safety constraint, while safe goal progress and trajectory safety are measured separately is established.
Yu-Chen Yuan, Zheng-Huang Wu, Yuan-Gan Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.