PostHog Releases Jeeves: Reasoning-Enhanced Decision Model Based on Qwen3.5-9B
Key point
Jeeves achieves 0.935 on JevBench overall, outperforming the baseline Jev model's 0.866 by integrating reasoning chains and a block-4 diffusion drafter.
Details
PostHog has released Jeeves, an open-source decision-making model that improves upon the Jev architecture by incorporating reasoning capabilities. Built on Qwen3.5-9B with LoRA fine-tuning and a specialized pointer head, Jeeves is designed to handle yes/no, multiple-choice, and rating queries in a single request while maintaining compatibility with the Jev API.
Performance Benchmarks
Jeeves demonstrates significant improvements in decision accuracy and calibration compared to the baseline Jev model and Kev-9B:
- JevBench Overall (231 items): Jeeves scores 0.935 vs. Jev's 0.866.
- JevBench Hard (111 items): Jeeves scores 0.865 vs. Jev's 0.730.
- Test Overall (Out-of-domain): Jeeves scores 0.889 vs. Jev's 0.857 and Kev-9B's 0.822.
- Calibration (ECE): Jeeves achieves a lower error of 0.037 compared to Jev's 0.049.
While Jeeves excels in rule-based and contrastive policy tasks (scoring 1.000 in held-out rule structures), it shows a trade-off in pure knowledge retrieval, scoring lower than Jev on MMLU (0.793 vs. 0.900) and MMLU-Pro (0.739 vs. 0.840).
Architecture and Inference Efficiency
The model utilizes a block-4 diffusion drafter to accelerate inference, achieving 176 tokens/sec for single questions (1.6x speedup over plain greedy decoding) and approximately 960 tokens/sec when batched. The architecture employs a pointer head that calculates scores via scaled dot products between the decision token's hidden state and option key projections.
- Latency: Without reasoning, median latency is ~0.3s. With full reasoning, median latency rises to 3.3s, with a p90 tail of 17.1s.
- Reasoning Impact: Including reasoning chains improves test split accuracy from 0.804 to 0.840.
- Optimization: Users can balance accuracy and speed using
max_thinkandnothink_thresholdparameters; for example, limiting thinking tokens to 768 reduces median latency to 2.0s with a slight accuracy drop to 0.806.
Training and Limitations
Jeeves was trained using SFT (2 epochs, 19,126 questions) followed by CISPO reinforcement learning (early stopped at step 402 to prevent over-sharpening). Key limitations include higher tail latency during complex reasoning and reduced interpretability of thinking chains due to the absence of language consistency rewards. The project requires Python 3.12 and a CUDA GPU (Hopper architecture recommended for FP8 kernels).
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.