Skip to content
#edge computing Preprint

FARM: Fundamental Agentic Reward Model For Multi-task Wireless Network Optimization

Oct 2026 · 0 citations · 25 references
Engineering Computer Science

TL;DR

FARM (Fundamental Agentic Reward Model for Multi-task Wireless Network Optimization), a reward-space transfer framework that shifts cross-task knowledge reuse from policy space to trajectory-level decision evaluation, is proposed.

Abstract

Future wireless networks require learning agents to adapt across heterogeneous channel conditions, traffic patterns, quality-of-service (QoS) requirements, objectives, and operational constraints. Reusing decision knowledge across such tasks is challenging because conventional multi-task and transfer reinforcement learning methods primarily share or transfer policies, coupling transferable knowledge with task-dependent action mappings. This paper proposes FARM (Fundamental Agentic Reward Model for Multi-task Wireless Network Optimization), a reward-space transfer framework that shifts cross-task knowledge reuse from policy space to trajectory-level decision evaluation. FARM introduces an Agentic Reward Model (ARM) that learns a task-conditioned reward prior from heterogeneous source-task trajectories and provides auxiliary guidance for task-specific policy optimization. In Stage I, ARM jointly models task conditions, temporal trajectory dependencies, and objective-dependent reward structures while each source task retains its own controller. In Stage II, the learned reward prior is frozen and reused to guide the adaptation of a target-specific controller for previously unseen tasks, without transferring source-task policies. Experiments on heterogeneous multi-access edge computing (MEC) tasks show that FARM achieves a mean late-stage gain of 29.8% over Single-task SAC on unseen Rate-Latency targets, compared with 16.1% for CRA Transfer, and reaches a 46.2% gain on the moderate-OOD FAR-M case. Further analysis shows that both Mamba and Transformer trajectory encoders support Reward-Space Transfer, while Mamba provides improved robustness as longer history dependencies are introduced.

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54

Related blog posts

Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

Microsoft Research Blog Sep 29, 2026

Introducing Quine: An AI research system designed for the complexity of biology

Biology doesn't operate in silos, and neither should the AI representation of it. Quine is an early-stage research effort to create a multimodal world model of biology. By connecting insights across biological scales and modalities, Quine helps scientists computationally search a space far larger than intuition allows and prioritize hypotheses before they reach the lab. Experimental results provide important feedback, helping researchers sharpen future research directions. The post Introducing Q…

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.