This work introduces a method that maintains near-baseline accuracy, induces usage-based modularity by sparsely routing inputs to neuron groups, and encourages specialization of these modules, such that their activations are correlated with input classes.
CastClaw is presented, a human-in-the-loop autonomous forecasting system built through forecasting-oriented harness engineering that connects data, specialized models, analytical tools, user input, and a versioned execution record in one runtime.
Xiaoyu Tao, Mingyue Cheng, Ze Guo et al.· 0 citations
Results provide initial evidence that KV states can serve as transferable computational representations rather than strictly model-local caches, and motivate context mobility as a systems abstraction for reducing redundant prefill across heterogeneous LLM and multi-agent inference workflows.
Yi Li, Dongming Jiang, Yi Zhao et al.· 0 citations
Gradient descent (GD) is explicit Euler for gradient flow, but a state-accurate continuous-time surrogate need not remain accurate after differentiation. At every fixed nonresonant step size, ordinary automatic differentiation exactly differentiates the executed hard-ReLU GD program. We prove that, over a fixed finite horizon, the GD states converge and these exact discrete derivatives approach an event-free regional propagator, whereas the derivative of the limiting flow also contains speed-normalized activation-event transfers. A prepoint Stieltjes representation separates the absolutely continuous regional Hessian from atomic interface curvature; one nonzero gradient jump produces an exactly rank-one endpoint discrepancy, and global convexity prevents complete multi-event cancellation whenever an event is strict. Nevertheless, a standard family of globally 1-strongly convex residual-ReLU squared-loss risks realizes arbitrarily large reciprocal sensitivity ratios on open initialization sets, with a uniform transversality margin. The same discrete-versus-flow decomposition extends to parameters and reverse-mode adjoints; resolved smoothing in the scalar or autonomous-normal regime and consistent event localization recover the flow sensitivity. The results concern deterministic full-batch, finite-horizon dynamics with a stable finite itinerary of separated same-direction transverse events; they are consistency theorems, not prevalence claims for large-scale training.
Xiao-Yang Li, Run-Ni Zhou· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This model interleaves reasoning, tool calls, and returns in one left-to-right generation, trained by a supervised warm-up and then outcome-level reinforcement learning against a programmatic reward read directly off the gold call chain, which leaves no learned critic and no judge in the training loop.
Armin Dariani, Sifan Wu, Bang Liu et al.· 0 citations
An externally anchored leading-order effective field linking empirical dynamics, an interpretable mechanism and neural computation is identified linking empirical dynamics, an interpretable mechanism and neural computation in closed-loop human-AI systems.
Min-Lin Wu, Xu Fang, Yi-Cheng Zhang et al.· 0 citations
A Mean-Calibrated Kernel UCB (MCK-UCB) algorithm is proposed that turns each incomplete sales record into a reliable guide for both inventory and price decisions, using data from past rounds with similar market conditions.
Ze-An Han, Jing Liang, Ruihan Lin et al.· 0 citations
This work calls on the stream-learning community to make bounded resource usage a first-class design objective alongside drift adaptation, and proposes concrete steps toward this goal, including an API through which stream learners can explicitly expose and respect resource budgets.
Sebastian Buschjäger, N. Gunasekara, H. Gomes· 0 citations
This work introduces Selection-Aware Semantic Stress Testing (\SASST{}), which learns a task reweighting from pre-execution features on discovery tasks and evaluates the same paired comparison on separate confirmation tasks.
Yang Xu, Chenang Li, Jiefu Zhang et al.· 0 citations
S3C-LLM is introduced, a skill-guided and code-grounded agentic LLM for spectrum-to-structure elucidation that consistently outperforms current general LLMs and spectrum-specific models across spectra, while using less than 1/10th of SpectraLLM's training corpus.
Xuan-Le Zhao, Xinyu Cai, Xiang Cheng et al.· 0 citations
This work proposes code surrogate gradient as the first order signal in deployable code space to acceleate optimization, and performs guided search to preserve deployment faithfulness in fine-tuning low-bit models across different quantization datatypes.
This work validates the engineering feasibility of running industrial-scale trillion-parameter LLM-driven biomedical computing tasks on consumer hardware, establishing a new low-barrier paradigm for AI-powered early stage drug discovery.
A weeklong summer workshop brought higher education faculty to campus to explore how AI and machine learning materials can be adapted for their classrooms.
Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.