Skip to content

JET: Justification Evaluation in Transformer

Sep 2026 · 0 citations · 15 references
Computer Science

TL;DR

The accuracy-throughput comparison covers model, hardware, and reasoning choices, with Jev as an external reference, and support local decision inference from existing models supports local decision inference from existing models.

Abstract

JET uses pretrained language and vision-language models to select among a finite set of answers without additional training. It evaluates candidate likelihoods directly and shares computation across candidates. Experiments on desktop CPUs and consumer GPUs assess decision accuracy and execution cost. Qwen3.6-35B-A3B achieves 87.48% accuracy on the full MMLU test set and 3.69 requests per second on a separately timed MMLU subset. The accuracy-throughput comparison covers model, hardware, and reasoning choices, with Jev as an external reference. Controlled execution experiments show 2.18-2.23-fold speedups from prefix reuse and cache management, and a 30.8% reduction in process time from input preparation optimizations, with unchanged outputs. Optional reasoning has a task-dependent accuracy-throughput trade-off. These results support local decision inference from existing models.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

vla.simd: Efficient CPU Inference for Language-Conditioned Manipulation

This work presents vla.simd, a CPU inference engine that combines shared SIMD micro-kernels, reusable computation, and target-specific optimization, and introduces IMPACT, an ACT-based policy with cached text representations and language-modulated visual features.

Khanh Duy Nguyen, Hoang M. Truong, A. T. Le · 0 citations
#artificial intelligence Preprint Sep 2026

SCX Router: Streaming Zero-Shot Model Selection with a Decoder-KV Classifier and a Real-World Task Ontology

A lightweight GLiClass-based router is introduced, a lightweight GLiClass-based router that assigns a suitability score to each inference-time model label without autoregressive generation, and a released 0.6B-parameter checkpoint combines a Qwen3 decoder with a shallow bidirectional scorer.

Ihor Stepanov, Aleksandr Smechov, Mykhailo Shtopko et al. · 0 citations
#machine learning Preprint Oct 2026

How Much Can Language Models Gain from Test-Time Computation?

How much can test-time computation improve a language model, and at what cost? Test-time scaling is widely proposed as a substitute for larger models, but existing comparisons mostly evaluate one domain at a time and rarely charge selection to the budget. We introduce SELF-POT, a benchmark and evaluation framework that...

Bang Yang, Jing-Yuan Li, Jia-Jun Fan et al. · 0 citations
#machine learning Preprint Aug 2026

Task-Aware Spectral Pruning: A Mixture-of-Masks Framework for Efficient LLM Inference

Task-Aware Spectral Pruning (TASP), a post-training framework that calibrates module-level spectral descriptors against measured task-specific ablation effects, closes grouped-query-attention and SwiGLU dependencies during sparse-mask construction, and routes each user turn to one compiled mask that remains fixed throu...

Ibne Farabi Shihab, Fariya Afrin, Sanjeda Akter et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.