Skip to content

The Fellowship of the Query: Learning Retrieval Actions

Sep 2026 · 0 citations · 24 references
Computer Science

TL;DR

Overall, trajectory supervision improves action prediction and evidence-recording behaviour in this evaluated pipeline.

Abstract

Retrieval-augmented question answering requires control decisions about when to decompose a question, search, reformulate, extract evidence, synthesize facts, verify progress, and stop. We study whether trajectory fine-tuning can improve small language models (SLMs) as next-action controllers. We additionally evaluate a low-resource setting in which a single SLM serves as both the controller and the final-answer generator. From accepted teacher search traces, we build a seven-way action-prediction task, where the model predicts the next structured teacher action from the current trajectory state, and evaluate LoRA-supervised fine-tuning across SLMs and xSLMs as controllers. On 1,646 held-out action examples, Granite 4.1 3B trained on 13,194 actions reaches macro-F1 0.6536, compared with 0.1736 for zero-shot prompting of the same model and 0.5399 for a TF-IDF logistic-regression baseline. In an end-to-end controller/generator swap evaluation over 149 held-out trajectories, using the fine-tuned model for both roles improves Exact Match from 0.7530 to 0.7946 and token F1 from 0.7783 to 0.8295 compared with using the base model as both controller and generator. The cross-role conditions show that the fine-tuned controller increases evidence-fact recording when the generator is fixed, while controller-only final-answer gains are not statistically clear. Overall, trajectory supervision improves action prediction and evidence-recording behaviour in this evaluated pipeline. Code is available at https://github.com/padas-lab-de/agent-action-controller

View source

Similar papers

Preprint Aug 2026

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL

Implicit multimodal in-context learning compresses demonstrations into internal interventions, ranging from static task vectors to query-conditioned transformations and attention routing. Despite their common goal, these methods differ substantially in how the intervention depends on the query and where it modifies the...

Jiaqian Li · 0 citations
#artificial intelligence Preprint Sep 2026

Rethinking On-Policy Distillation of Large Language Models II: One Training Example

On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. Existing work has mainly studied its algorithmic behavior, leaving the role of training data unclear. We examine this role at the data-minimal limit by training on a single query. One-shot OPD keeps impro...

Zi-Xuan Fu, Bing-Xiang He, Yu-Xin Zuo et al. · 12 citations · ⚡1
Preprint Aug 2026

Training a Knowledge Base: Supervised Structure Learning for Agent-Curated Document Stores

This work evaluates the curated store with a test set, on two contamination-free benchmarks: KBGym, a fictional-universe generator the authors release, and PhantomWiki, a fictional-universe generator they release, and PhantomWiki, a fictional-universe generator they release.

Yu Pan, Hongfeng Yu · 0 citations
#machine learning Preprint Sep 2026

You Only Edit Once: Incentivizing In-Context Capability of LLMs via Local Demonstration Refinement

In-context learning (ICL) is crucial for boosting the inference performance of large language models (LLMs). However, the effectiveness of ICL in LLMs is greatly influenced by the choice of demonstration sets. Exhaustive searches over these sets are combinatorial, and existing selectors often rely on relevance or likel...

Jia-Rong Wen, Qi Wang, Yun Qu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Search-Aware Reinforcement Learning for Multi-Component Query Understanding in Roblox Game Search

Query understanding (QU) plays a critical role in production search systems, translating raw user queries into search execution plans that drive downstream retrieval and ranking. While large language models (LLMs) have enabled QU to be framed as a structured multi-task generation problem (e.g., intent classification, q...

Nayoung Choi, Sheng-Jian Chen, Xiao-Kai Wei et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.