AI Briefing
Sign in

Open-Weight 'Jeff' Models Match Jev Benchmarks with ~30ms Latency on Local Hardware

·2026.09.29 05:17

Key point

The 2B parameter Jeff model achieves an 83.1% score on a five-benchmark panel, slightly exceeding the published 83.0% score of the larger Jev model.

Details

A new set of open-weight models named Jeff, fine-tuned from Qwen3.5 and Gemma, delivers ultra-fast zero-shot classification with latencies as low as 28 ms on an M4 Max. The models are designed for "System 1" decision-making, providing calibrated probabilities for multiple-choice scenarios in a single forward pass without text generation.

Benchmark Performance

The Jeff-Qwen3.5-2B model achieved an 83.1% average score across a panel of five benchmarks (BBH, Financial PhraseBank, JudgeBench, RAGTruth, WinoGrande), marginally surpassing the published 83.0% score of the larger Jev model. The smaller Jeff-Qwen3.5-0.8B scored 79.1%.

  • Calibration: The 2B model showed a calibration error of 0.028, compared to Jev's ≈0.06.
  • Classification vs. Reasoning: While Jeff excels at classification tasks (e.g., 96% on Financial PhraseBank), it lags behind Jev in multi-step reasoning (e.g., 64–68% on BBH vs. Jev's 94%).

Training and Infrastructure

All training and data generation were performed on local hardware, avoiding cloud dependencies:

  • Training: Conducted on a single RTX PRO 6000 (96GB VRAM). The 0.8B model trained in 2 hours, and the 2B model in 3.5 hours.
  • Synthetic Data: Approximately 31k synthetic training questions were generated and verified by Qwen3.8-Flash-Next running on two DGX Sparks.
  • Dataset: The final training set included public datasets converted into decision formats, alongside the synthetic data.

Practical Implications

The author emphasizes that these small models function as classifiers rather than planners. In game-based zero-shot tests (Doom, Frogger, Pac-Man), the Jeff 0.8B model outperformed the 2B variant and matched Jev's published Doom score (6.55 kills) while operating at ~29 ms per decision, compared to Jev's ~212 ms API latency. The models are released under the Apache 2.0 license with a Jev-compatible API.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.