AI Briefing
KOSign in

Jared Palmer Releases Kev 1.0: Open-Source Decision Models from 0.8B to 27B Parameters

·2026.10.01 09:00

Key point

The release includes Kev-27B and an updated Kev-9B, offering Apache-2.0 weights compatible with TypeSafe's Jev API for local and cloud deployment.

Details

Jared Palmer has released Kev 1.0, a family of four open-source decision models ranging from 0.8B to 27B parameters. The release introduces Kev-27B and an updated Kev-9B, all available with Apache-2.0 weights on Hugging Face. These models are designed to provide a probability for each answer option, compatible with TypeSafe's Jev API, allowing applications to switch endpoints and model names easily.

Model Performance and Architecture

Kev-27B leads in overall evaluations, with Kev-9B and Kev-4B performing within a percentage point on consumer complaints and matching on developer-tool decisions. The models use a pointer head to score supplied answer options directly, avoiding token generation for answers. This architecture allows the document to be processed once, with each question using cached computation, significantly reducing latency for multiple queries on the same document.

Key performance metrics include:

  • Kev-27B: Best overall results, requires an H200 or B200 for long documents.
  • Kev-9B: Workflow accuracy close to Kev-27B on smaller GPUs like H100.
  • Kev-4B: Suitable for local use and fine-tuning, runs on L40S or 32 GB Macs.
  • Kev-0.8B: High-throughput, narrowly defined decisions, runs on L4 or Apple Silicon Macs.

Training and Evaluation

Kev-27B was fine-tuned on Qwen3.8-27B, using a broader corpus of approximately 146,000 examples and 337,000 questions. The training process included examples for tone, grounding, prompt injection, and personal-data detection. Kev-9B received an additional fine-tuning pass to close gaps in policy reasoning and developer-tool decisions, improving accuracy from 58% to 83% on policy tasks.

Evaluation results are based on the author's harness, not independent leaderboards. Kev-27B demonstrates strong calibration, with its confidence scores closely matching observed accuracy. The models are tested on tasks excluded from fine-tuning to assess generalization, with Kev-27B scoring 52.3 on a chance-corrected index compared to 41.0 for Kev-9B.

Deployment and Fine-Tuning

The release includes official skills for fine-tuning and deployment, compatible with coding agents like Claude Code and Codex. Kev-4B can be fine-tuned on custom data with a small cost of about $1 in GPU time. The models support local execution on Apple Silicon via MLX, with Kev-4B answering five questions about a short text in 721 ms on a 32 GB M5 Mac. Users can deploy Kev to Modal behind an authenticated HTTPS endpoint using the provided skills.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.