This work introduces Minimal Intervention Reinforcement Learning (MInTRL), which expands the exploration frontier through sparse, local interventions in otherwise on-policy rollouts, and establishes minimal intervention as an effective paradigm for enhancing on-policy RL.
Ming-Yu Chen, Ye-Fan Tao, Gerald Friedland et al.· 0 citations
The Human-LLM Reflection Framework is introduced, a controlled two-pass protocol comparing human and LLM revision under identical conditions across self-, peer-, and cross-agent settings, using an information-theoretic analysis based on per-iteration cross-entropy reduction.
The degradation rate across neural models, both sentence embeddings and decoder-only LLMs, is studied, and how consistent it is depends on the scale of the noise: under word-level noise, models with very different architectures decline along nearly the same curve, while under character-level noise they separate.