Skip to content

Numbat: Building and Verifying a Self-Contained Machine-Learning Stack

Sep 2026 · 2 citations · 40 references
Computer Science

TL;DR

This work reports on the construction and verification of numbat, a machine-learning stack written in one general-purpose language (Zig) with no third-party runtime dependencies, and treats a widely used reference implementation as an executable specification and verify against it at five levels.

Abstract

Machine-learning systems are built almost exclusively on a few large Python-orchestrated frameworks, and they inherit those stacks'engineering costs: environments of hundreds of version-coupled packages, separate export toolchains for deployment, and the split between the language research is written in and the language products ship in. We report on the construction and verification of numbat, a machine-learning stack written in one general-purpose language (Zig) with no third-party runtime dependencies. The stack spans tensor computation, automatic differentiation, neural-network modules, mixed precision, multi-GPU training, data loading and monitoring; an SDK exposes it behind a stable, additively versioned C ABI of over 1,400 entry points, with bindings for six languages; and its clinical domain planes encode regulatory requirements as executable acceptance gates rather than documentation. Verifying such a stack is the harder half of building it: a defective training run rarely fails, it converges quietly to a slightly worse model. We treat a widely used reference implementation as an executable specification and verify against it at five levels, from operator gradient checks to an automated trajectory gate against a same-machine reference run - the arrangement our companion study formalizes as a trajectory-level differential oracle. The protocol surfaced ten silent recipe divergences, which we catalog with mechanisms and symptoms. As the acceptance test, we train a 25.9M-parameter detector of the YOLOv8m class from random initialization on COCO 2017 for the full 500-epoch schedule: the exported weights score 0.4956 mAP50-95 under the official protocol, scored by the reference stack's own validator (published endpoint 0.502), with single-GPU step time at parity on identical hardware. Weights, per-epoch metrics and the full run manifest are released.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Flama: a Python framework for development and deployment of production-ready APIs, machine learning, and LLM services

Flama is presented, an open-source Python framework for developing and deploying production-ready web APIs, machine learning services, and large-language-model (LLM) applications that unifies REST API development, predictive model serving, and generative AI inference in one architecture.

José A. Perdiguero López, Miguel A. Durán-Olivencia · 0 citations
#machine learning Preprint Sep 2026

OPEN-1B: A Fully Auditable Training Run

Open-1B, a model trained under this regime, is released, together with its full pretraining dataset, every intermediate checkpoint, the training codebase, and the audit harness needed to reproduce and verify any step of its training.

John Donaghy, B. Wilcox, Oğuzhan Ersoy et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Design Docs Are All You Need: An AI-native Machine-Learning Performance Tool

SMART is described, a rigorous symbolic performance-modeling library for ML systems whose main branch contains almost no code, and Regenerated implementations reproduce hand-audited reference models to round-off precision, suggesting that design docs can be the durable artifact for ML-systems co-design tools.

Samuel Kushnir, Kimia Noorbakhsh, Kavya Sreedhar et al. · 0 citations
Preprint Aug 2026

Evolving Executable Pipeline Programs for AutoML with Language Models

This work presents LACE, an AutoML framework that instead searches over complete executable pipeline programs: an evolutionary loop maintains a population of scikit-learn-compatible Python classes, and a large language model acts as the variation operator.

Sofoklis Kitharidis, C. Veenman, J. V. van Rijn et al. · 0 citations
#artificial intelligence Preprint Sep 2026

RLX: A Unified Multi-Backend Tensor Compiler and Distributed Runtime in Rust

Production machine learning (ML) stacks often split graph compilation and kernel execution across different layers and languages, making backend behavior, deployment guarantees, and performance fallbacks hard to reason about end-to-end. RLX addresses this gap with a single Rust codebase that combines compiler and runti...

Eugene Hauptmann, Nataliya Kosmyna · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.