Skip to content
Preprint

Investigating Knowledge Transfer Across Interactive Dialogue Games

Aug 2026 · 0 citations · 30 references
Computer Science

TL;DR

This paper investigates how knowledge transfers across different dialogue games by finetuning LLM models on games from the clembench suite and finds that some games benefit more from transfer than finetuning, and that the visuospatial family transfers best.

Abstract

Dialogue games represent a challenging setting where complex cognitive skills are required to accomplish tasks while coordinating with other players. Considering that language represents an interface for both understanding the game rules and executing actions, it is reasonable to assume that training on a specific language game will enhance specific capabilities that might be relevant for other tasks as well. Motivated by this rationale, in this paper, we investigate how knowledge transfers across different dialogue games. We study transferability by finetuning LLM models on games from the clembench suite (Chalamalasetti et al., 2023) and performing two analyses: i) we derive a task-transferability graph using a binary integer optimization program from Zamir et al. (2018), using task performance as the main metric; and ii) we compute task vectors (Ilharco et al., 2022) for each game to study similarities across finetuned models and their task transferability. In our first analysis, we find that some games benefit more from transfer than finetuning, and that the visuospatial family (e.g., exploration games) transfers best. With our task vector analysis instead, we find that similarity-based approaches capture game-role relationships but almost no transferability patterns, suggesting that more complex metrics are required.

View source

Similar papers

#natural language process... Preprint Aug 2026

First Make It Playable, Then Make It Good: Staged Interaction Learning for Small Dialogue-Game Agents

It is suggested that imitating full trajectories helps with playability, while turn-level and teacher-guided training usually improve decision-making and increase the overall score, and small models are performant simply by using careful curation strategies rather than aggressive changes.

Syed Mahbubul Huq, P. Madhyastha · 0 citations
#artificial intelligence Preprint Sep 2026

Modular Discovery of General Game-Playing Algorithms with Large Language Models

General Game Playing across arbitrary games from rules alone remains challenging due to differing algorithmic requirements across game classes and strict decision-time constraints. Rather than hand-designing search heuristics for specific domains, can we leverage Large Language Models (LLMs) to discover general game-pl...

Zun Li, John Schultz, Marc Lanctot et al. · 0 citations
Book Open access Aug 2026

MMID: Multi-turn Multimodal Interactive Dialogue Benchmark

The proposed Multi-turn Multimodal Interactive Dialogue (MMID) Benchmark enables comprehensive evaluation of the Perception, Memorization, and Reasoning abilities of MLLMs, and reveals MLLMs perform well with text-based input but degrade with images, requiring improved leverage fine-grained visual cues.

Seulgi Kim, Juoh Sun, Sumin Kim et al. · 0 citations
#artificial intelligence Preprint Sep 2026

A Qualitative Model for Reasoning about Path and Support

Spatial reasoning abilities correlate strongly with performance in STEM fields. Games offer a compelling medium for training these critical skills in developing children who have a natural proclivity for play. However, to facilitate human-like tutoring and player guidance, these games require an AI agent capable of mak...

Abhishek Jaiswal, Zoe Falomir · 0 citations
Open access Aug 2026

Gaming with AI: A Hybrid Reinforcement Learning, Large Language Model, and Procedural Content Generation Framework for Enhancing Player Engagement and User Experience

Most existing game AI research examines individual mechanisms—Dynamic Difficulty Adjustment (DDA), large language model (LLM)-driven NPC dialogue, and Procedural Content Generation (PCG)—in isolation, leaving open the question of how these subsystems interact when evaluated together against a shared user-experience (UX...

Abhinav, Amandeep, Dharmender Kumar, Suraj, Keshav · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.