Decision Transformers (DTs) have emerged as a powerful paradigm for sequential decision-making in offline learning, yet they face two intrinsic deficiencies: the challenge of specifying optimal Return-to-Go (RTG) targets and a fundamental lack of exploration beyond the offline dataset. While prior studies have attempted to address these issues separately, their solutions introduce new problems. In this paper, we introduce LOGIC (Learning Optimal Goals with an Integrated Critic), a holistic framework that resolves these dual challenges in a unified manner. LOGIC transforms the exploration paradigm: instead of exploring at the action level, it employs an integrated critic to guide a generator in discovering optimal, high-value goals (RTGs). This strategic shift enables robust exploration beyond the dataset's frontier while avoiding the fragility of complex multi-loss objectives. Simultaneously, a Backward Consistency Refinement (BCR) module, applied during both training and inference, ensures that all generated plans remain mathematically and semantically valid. Extensive experiments on diverse auto-bidding scenes demonstrate that LOGIC significantly outperforms state-of-the-art baselines, achieving superior cumulative returns and model robustness, thereby providing a principled and unified solution to goal-oriented offline learning.
Shihao Shu, Rujie Zhong, Hao Wang et al.· Proceedings of the 32nd ACM...· 0 citations
LlamaExtract Agentic Plus ranks first on all three metrics, with accuracy comparable to coding agents at a fraction of the cost, and is the first to score value accuracy, record completeness at scale, grounding, and measured cost together.
Boyang Zhang, Adrian Lyjak, Elizabeth Stewart et al.· 1 citation