Student-Centric Answer Sampling (SCAS) is proposed, a framework that selects from verified teacher-generated answers according to their estimated student-centric learning cost and is derived by a token-wise gradient decomposition and used to guide answer selection during training.
Zhengyu Hu, Zheyuan Xiao, Linxin Song et al.· 0 citations
ProMoS is introduced, the first unsupervised generalist GAD framework, which detects anomalies by modeling the abundant normality in unlabeled data, and proposes prototype-guided soft-label distillation to align teacher and student in a shared prototype space, enhancing cross-graph generalizability.
Yiming Xu, Zihan Chen, Z. Peng et al.· arXiv.org· 0 citations
A fixed encoder is studied in a fixed encoder and trajectory reachability metrics (TRM), a small temporal pairwise cost trained from logged trajectories and used to rank predicted endpoints of a candidate action sequence against a goal is introduced.
Lian Li, Shengzhi Wang, Li-Bin Qiu et al.· 2 citations
This research demonstrates that the model, ArchesClimate -- SSP, does not simply imitate scenarios seen during training, but is actually capable of modeling the response of a climate state to diverse forcings, an important step towards reliable and rapid climate model scenario generation.
Graham Clyne, Julia Kaltenborn, Peer Nowack et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
ProFIL (**Pro**be-**Filtered Reinforcement Learning) is introduced to reduce theater, increase chain-of-thought faithfulness, and shrink chain length in a single, drop-in extension to Group Relative Policy Optimization (GRPO).
Swapnil Parekh, Naman Goyal· arXiv.org· 2 citations
This work proposes ADMM-Q, a novel weight quantization algorithm that considers the layer-wise quantization problem, based on a combinatorial variant of the Alternating Direction Method of Multipliers (ADMM).
Ryan Lucas, Mehdi Makni, Xiang Meng et al.· arXiv.org· 1 citation
A dynamic per-layer scalar derived by adapting the LARS/LAMB trust-ratio principle to the orthogonalized setting, where the standard denominator candidates---the raw momentum norm or the polar-factor norm---either live in the wrong unit space or carry no update-scale information.
These results expose a rate-granularity trade-off: PairAlign does not uniformly outperform denser tokenizers on every local metric, but provides a lower-rate symbolic interface preserving ordered and relational structure.
AOPD replaces ineffective negative reinforcement with localized divergence minimization in non-positive advantage regions while preserving positive reinforcement learning and maintains higher policy entropy during training and better capability retention during sequential tool-use adaptation.
Nan Jia, Haojin Yang, Xing-Chen Ma et al.· arXiv.org· 18 citations· ⚡5
The results suggest transformers rotate semantic content into spectrally quiet regions during contextualized processing, where, in some architectures, interventions may reduce grammatical disruption relative to high-variance steering.
Pratyush Acharya, Nuraj Rimal, H. Dhakal· 0 citations
AirFM-DDA is proposed, an Air-interface Foundation Model in the Delay-Doppler-Angle (DDA) domain, which reparameterizes CSI into the DDA domain to resolve multipath components along physically meaningful axes and employs window-based attention with frame-structure-aware positional encoding.
Kejia Bian, Meixia Tao, Jianhua Mo et al.· arXiv.org· 6 citations· ⚡1
The platform supports an end-to-end workflow encompassing EIS preprocessing with selectable impedance representations, agent setup and training, ECM generation for new measurements, and visualization-based evaluation and analysis of agent decision-making.
A. Jaberi, Yonatan Kurniawan, Robert Black et al.· 0 citations
A weeklong summer workshop brought higher education faculty to campus to explore how AI and machine learning materials can be adapted for their classrooms.
Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.