The North Small Translate model is presented, an open-weight, LLM-based machine translation (MT) model with instruction-following capabilities built on the same foundation as Cohere's Command A Plus, a mixture-of-experts architecture with 25 billion active parameters out of 218 billion total parameters.
Abstract
We present North Small Translate, an open-weight, LLM-based machine translation (MT) model with instruction-following capabilities built on the same foundation as Cohere's Command A Plus, a mixture-of-experts architecture with 25 billion active parameters out of 218 billion total parameters. North Small Translate is trained using difficulty sampling to obtain challenging documents and a five-step training protocol combining supervised fine-tuning, direct preference optimization, and online reinforcement learning. We prioritized throughput through a non-reasoning base model and supplemented with optional agentic capabilities to unlock translation quality gains. North Small Translate is trained to perform MT-related tasks, including post-editing and quality estimation, as well as related tasks such as general instruction following. The model achieves top MT performance across 50 languages in the class of models under 1T parameters, with no need to run expensive reasoning at inference time.
Large reasoning models (LRMs) have shown exceptional performance in complex tasks such as mathematics and coding. In the field of machine translation (MT), reinforcement learning (RL) has been utilized to enhance the quality of translations. However, traditional RL approaches rely heavily on the base model’s inherent...
Zengkui Sun, Jia-Li Zeng, Jiaan Wang et al.· Transactions of the Associat...· 0 citations
This work studies reference-free post-training for multilingual machine translation with open large language models and finds that on-policy distillation reaches, but does not surpass, the quality frontier achieved by RL with checkpoint interpolation.
Chris Han, Pengzhi Gao, Pei Fu et al.· 0 citations
This work proposes augmenting existing benchmarks to increase translation difficulty by combining adversarial optimization with a differentiable translation difficulty estimator, and uses gradients from a combined difficulty and fluency objective to iteratively replace tokens in Adversarial Translation Optimization (AT...
William Kalikman, Šimon Sukup, Michal Tesnar et al.· European Association for Mac...· 2 citations
In LLM-based speech translation, transcription-based chain-of-thought (CoT) suffers from a mismatch between reference transcripts used in supervised fine-tuning (SFT) and model-generated transcripts at inference. To address this, we propose joint recognition and translation fine-tuning via group relative policy optimiz...
The retrieval-augmented many-shot translation pipeline from the AmericasNLP 2026 system is adapted to translate between English and eleven North-Eastern Indian languages in both directions to solve the WMT26 Low-Resource Indic Language Translation shared task.
Aashish Dhawan, Christopher Driggers-Ellis, Dzmitry Kasinets et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.