The result shows that even a fixed-dimensional latent space suffices to achieve vanishing approximation error as the network budget increases, and it is proved that the optimal worst-case uniform approximation error over the unit ball ofolder functions on $[0,1]^d$ has the sharp order.
This paper formalise and study a succinct version of the compatibility problem, encoding conditional distributions as arithmetic circuits, and shows that, for succinct circuit representations of conditionals, the compatibility problem is intractable.
This work evaluates three dense and mixture-of-experts models on BBQ and BBQ-V under seven conditions spanning batching, quantization, benchmark reduction, and their combinations, and compares accuracy, bias severity and prevalence, reasoning quality, subgroup behavior, subset-membership stability, runtime, and measured GPU energy against a full-benchmark BF16 baseline.
Ahmed El kady, Aravind Narayanan, Rehana Noorani et al.· 0 citations
The findings suggest that the teacher models used to generate preference data can interact with alignment training objectives in unexpected ways, generalizing to undesirable and potentially harmful behaviors like sycophantic agreement.
C. Blank, Z. Ying, Christopher Potts et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This work introduces Soft Latent Thinking, a method that replaces the LM head during reasoning with a lightweight projector, enabling autoregressive rollout in embedding space where reasoning steps remain continuous rather than tokenized.
N. Koriagin, Yaroslav Aksenov, George Bredis et al.· 0 citations
Though the construction provably evaluates Boolean expressions -- a universal symbolic computation -- of arbitrary length perfectly, in other experiments it is demonstrated that the transformer variant can learn and generalize perfectly on other common length generalization benchmarks, including modular arithmetic and ListOps.
Takuya Ito, Ruchir Puri, Murray Campbell et al.· 0 citations
It is found that learning concentrates on low log-probability tokens, and using a single fixed negative advantage matches the performance of teacher-provided ones, suggesting that OPD works largely by suppressing low log-probability tokens, which requires no teacher.
This tutorial provides a comprehensive introduction to rotational equivariance, starting from the physical and geometric intuition behind coordinate independence and building up the necessary machinery from geometric deep learning, group theory, and representation theory.
Normalized Low-Rank Adaptation (NoRA) is introduced, a simple yet effective method that normalizes the down-projection matrices during training, improving standard LoRA without requiring repeated normalization throughout training.
Jiale Kang, Zi-Yin Yue, Zheng Zhan et al.· 0 citations
TSPFN is introduced, a foundation model that redesigns TabPFN's architecture for time series data and yields a unified, generalizable framework capable of learning the specificities of medical time series.
J´er´emie Stym-Popper, Clément Rambour, Federica Granese et al.· 0 citations
LiFT, a language-informed cross-modal framework built on Flow Matching for trend-guided 3D molecular generation across both de novo design and scaffold hopping, and suggests that language-derived chemical priors provide effective trend-level guidance for 3D molecular generation.
Tian-Yu Gao, Zhi-Kai Su, Jia-Shu Li et al.· 0 citations
This work develops a self-contained methodology for learning parameter-efficient rotational transformations based on Riemannian optimization and empirically validate the proposed rotation-based steering scheme, demonstrating its superiority in intervention efficiency.
Kirill Bunin, Dmitry Bylinkin, Vladimir Aletov et al.· 0 citations
A weeklong summer workshop brought higher education faculty to campus to explore how AI and machine learning materials can be adapted for their classrooms.
Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.