Skip to content

Towards Autonomous and Auditable Medical Imaging Model Development

Jul 2026 · arXiv.org · Vol abs/2607.10522 · 0 citations · 81 references
Computer Science

TL;DR

AMID is introduced, an autonomous multi-agent framework for medical imaging model development that outperformed evaluated general-purpose MLE systems and approached or matched strong human-designed challenge solutions across heterogeneous tasks.

Abstract

Large language model (LLM) agents are beginning to automate machine learning engineering (MLE) by coupling planning, code execution, debugging, and empirical feedback. Translating this capability to medical imaging remains difficult because each task imposes modality-specific experimentation and strict requirements for validation protocols and prediction artifacts. Here we introduce AMID, an autonomous multi-agent framework for medical imaging model development. AMID first proposes Data-Conditioned Method Planning, which refines coarse task-level search spaces into executable, parallelizable method lanes grounded in task-specific data analysis and runnable medical-imaging resources. It then develops Verification-Guided Two-Stage Optimization, moving from broad early exploration of diverse method lanes to selective exploitation of promising candidates while enforcing strict verification of validation protocols, metric computation, and prediction artifacts throughout the optimization. Across 20 medical imaging challenge tasks spanning diverse modalities and prediction types, AMID outperformed evaluated general-purpose MLE systems and, on several tasks, approached or matched strong human-designed challenge solutions. These results suggest that AMID can turn task-specific medical imaging model development from bespoke manual engineering into an agentic workflow for producing high-performing and auditable model artifacts across heterogeneous tasks.

View source

Similar papers

#small language model Preprint Aug 2026

Towards Fully Automated Medical Imaging Code Generation via Validation-based Context Engineering

This work proposes AutoMedImg, a multi-agent framework for fully automated medical image processing code generation that achieves zero human intervention, with Dice scores of up to 0.90 for segmentation tasks and 99% accuracy for classification.

Zi-Xiao Zhao, Jing Sun, Zhe Hou et al. · 0 citations
Review Open access Sep 2026

Transforming large language models into medical specialists via knowledge injection.

While general-purpose large language models (LLMs) demonstrate remarkable capabilities, their clinical application demands rigorous adaptation to ensure safety and accuracy. This review presents a comprehensive framework for transforming LLMs into trustworthy medical specialists. We detail three core knowledge-injectio...

Kiduk Kim, Jeong Min Song, Dong Yeong Kim et al. · 0 citations
Review Aug 2026

Can Coding Agents Build Robust Baselines? A Skill-Based Approach for Automating the Medical Imaging Model-Development Pipeline

Developing competitive deep learning baselines for medical imaging remains a highly iterative process requiring literature review, implementation, experimentation, and expert refinement. Existing automation approaches typically optimize isolated components, such as architecture search or hyperparameter tuning, rather t...

Eugenia Moris, José Ignacio Orlando · 0 citations
#artificial intelligence Preprint Sep 2026

BioDyad: Synchronize Biomedical Discovery and Machine Learning Engineering

Agentic biomedical machine learning (ML) draws on complementary advances in biomedical evidence acquisition and executable program search. Existing systems connect aspects of these capabilities, but coordinating them throughout program search remains challenging. New evidence must guide candidate construction, executio...

Xing-Bo Du, Fadli Aulawi Al Ghiffari, Le Song et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MLToolBench: Learning Tool-Augmented Agents for Machine Learning Development

Machine learning engineering (MLE) agents have made substantial progress, but learning through ML experimentation remains costly in time and computation. Synthetic environments reduce these costs while introducing variations in data and experimental settings that require task-specific diagnosis. Access to diagnostic to...

Xin Yu, Li-Zhu Zhang, Jiamu Bai et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.