Skip to content

Author

Md Najmus Swaqeeb

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

CODS: Iterative Bellman-Residual Data Selection for Reusable Offline Reinforcement Learning

CODS is introduced, a critic-guided selector that alternates between fitting an algorithm-matched critic and acquiring high-residual transitions before freezing a reusable subset, a reusable selection procedure, not a formal coreset guarantee.

Ibne Farabi Shihab, Sanjeda Akter, A. Ahsan et al. · 0 citations
Preprint Aug 2026

Opportunity Is Not Realizability: Selection-Valid Diagnostics for Multi-LLM Routing

Oracle routing measures how much a pool of language models could gain from per-query selection, but the diagnostic has two flaws: testing against a best fixed model selected on the same examples invalidates paired inference, and a full-information oracle sees outcomes no deployable router observes. We separate three estimands (outcome-oracle opportunity, the Bayes-optimal gain from a declared pre-answer signal, and the held-out gain of a learned router) and prove selection-valid confidence intervals that survive choosing the best fixed model or the best member of a router family, a signal-information sandwich, and a $(1-1/e)$ greedy guarantee for building compact pools from submodular complementary coverage. On eight checkpoints from six families over four benchmarks, selection-valid intervals certify a population oracle gap of $9.7$--$30.7$ points on every task, yet the strongest deployable prompt router recovers only $7.5$--$14.4\%$ of it, and the simultaneous interval for the best of eleven tested policies has lower limit zero throughout. The realizable share of oracle opportunity is small and certifiable: strong routers beat the best fixed model, and most of the gap remains.

Ibne Farabi Shihab, A. Ahsan, Md Najmus Swaqeeb · 0 citations