This work removes a model's camera input and replaces it with memories from prior drives at the same location, suggesting that a high NAVSIM score does not require a planner to react to the current traffic scene and should be treated with caution.
Christian Löwens, Thorben Funke, Alexandru Condurache· 0 citations
This work presents a framework for learning continuous latent representations of admissible partial differential equations by embedding a scientific inductive bias directly into the training distribution, and shows that embedding a scientific inductive bias in the training distribution enables the learning of compact and geometrically meaningful hypothesis manifolds.
Experiments show that knowledge-aligned SFT can reduce factual hallucinations on WildHalu and Biography while largely preserving general capabilities and confirm that SFT targets beyond the base model's knowledge drive hallucination behavior.
AR Becker, Jakob Kemmler, David Thulke et al.· 0 citations
CoJEPA shows that combining objectives with complementary inductive biases can substitute for scale, encouraging future work to invest in smarter training objectives over ever-larger models.
Gabriel Meseguer-Brocal, Yuexuan Kong, Romain Hennequin· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Deploying a new control policy for voltage control in active distribution grids requires evidence that physical limits will be satisfied before the policy is tested on the physical grid. This assessment is difficult for two reasons. First, simulations cannot capture every disturbance, modeling error, and device interaction present in the real grid. Second, historical measurements reflect operation under existing control policies, whereas a new policy may drive the grid into different operating conditions. To address these challenges, we propose Distributionally Robust Conformal Safety Screening (DR-CSS), a policy-agnostic framework for pre-deployment, scenario-by-scenario screening of a new control policy using historical data and a nominal simulator. For each new scenario, the simulator predicts a future voltage trajectory for the whole grid; DR-CSS then constructs a conformal safety interval around this prediction using historical simulation-to-reality errors. The interval is further enlarged to account for closed-loop changes induced by the deployment of the new policy and its interactions with the remaining controllers. To the best of our knowledge, DR-CSS is the first framework in power systems to combine historical data from an existing control policy with an imperfect simulator for pre-deployment safety screening of a new policy. Experiments on the IEEE 33-bus and IEEE 141-bus systems evaluate the deployment of learning-based voltage control policies and show that DR-CSS identifies all unsafe test scenarios. To reduce unnecessary warnings on safe scenarios, we adapt the safety intervals to different operating conditions and gradually introduce new policies with recalibration after each stage. These extensions increase the informational value of the safety screening and support safer deployment decisions in active distribution grids.
Sarra Bouchkati, P. Ellinas, Adriana Geisler et al.· 0 citations
It is observed that the correlation between speakers'L1 distance and ASR error rates yields a systematic effect on English Speech, with its strength varying across datasets and models.
Tingyu Cheng, L. Clemmensen, Sneha Das· 0 citations
TopoCompress is introduced, a training-free and model-agnostic framework that compresses long contexts by selecting coherent semantic spans by selecting coherent semantic spans and achieves performance comparable to the strongest baseline while using a 4x smaller compression budget.
End-to-end weather forecasting systems produce skillful global gridded and station forecasts directly from raw Earth observations, replacing the numerical weather prediction pipeline, including data assimilation, at a fraction of its cost. These systems are deterministic and issue no uncertainty. Here we render the Aardvark Weather model probabilistic by attaching one stochastic mechanism to each component: learned, input-dependent noise at the observation encoder, capturing aleatoric uncertainty inherited from the observing system, and Monte Carlo dropout in the processor, capturing epistemic uncertainty in the learned dynamics. The resulting nested ensemble attributes forecast spread to the two sources through a law-of-total-variance decomposition, cross-checked by withholding observation streams. Probabilistic finetuning significantly improves the mean forecast, by 4.2% on average across variables and lead times. The ensemble is calibrated against ERA5 through the medium range (spread-skill ratio 0.98), keeps station RMSE within 2.4% of the deterministic model while beating it in CRPS at every lead time, and trails the operational ECMWF ensemble. The encoder branch behaves as observation-driven uncertainty. Component-attributed uncertainty makes end-to-end forecasts more transparent, a step toward observation-driven digital twins of the atmosphere.
Rodrigo Almeida, Noelia Otero, Jost Arndt et al.· 0 citations
This work introduces the first end-to-end neuromorphic spike-encoding and evaluation of the TIMIT dataset and quantifies the pipeline's efficiency with hardware-agnostic metrics based on the quantitative spiking activity.
Valentin Meunier, Amélie Gruel, Pierre Lewden et al.· 0 citations
SingProbe is introduced, a lightweight intrinsic runtime guard that directly reuses hidden states produced during LLM inference and operates alongside autoregressive decoding and extends this paradigm to medical generation through SingProbe-Med, which selectively activates risk-directed decoding interventions only when clinically relevant risks emerge.
Results show that route-supervised frontier selection can improve budgeted search without altering biochemical generation, although performance remains dependent on frontier construction and reaction ranking.
Philippe Meyer, Guillaume Gricourt, T. Duigou et al.· 0 citations
MIOH provides a controlled framework for analyzing multi-image object hallucination and serves as a critical evaluation tool for developing more reliable multimodal AI systems.
Joonki Min, Chaeyun Kim, H. Choi et al.· 1 citation
A weeklong summer workshop brought higher education faculty to campus to explore how AI and machine learning materials can be adapted for their classrooms.
Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.