AI Briefing
KOSign in

Aplomb 1 Tops Typed Decisions Benchmark with Best KL and Brier Scores

·2026.10.10 00:58

Key point

The open-weight Aplomb 1 model achieves a KL divergence of 0.123 and a Brier score of 0.065 on the Typed Decisions leaderboard.

Details

Aplomb 1 has secured the top position on Typed Decisions, an official Hugging Face benchmark for probabilistic decisions, outperforming all other models including a 27B parameter competitor.

Benchmark Performance

The model achieved the lowest KL divergence from gold at 0.123, which is 38% lower than the next best model, and the lowest Brier score at 0.065. It also recorded the highest accuracy (73.7%) among all models under 6B parameters. The benchmark evaluates models on 400 cases and 2,000 decisions, requiring probability distributions for five typed questions per unstructured text input.

Production Capabilities

Unlike general LLMs that generate tokens without explicit confidence measures, Aplomb 1 returns probability distributions for every answer, tool, enum, and boolean argument. This allows software to act autonomously when the model is confident and escalate to larger models when uncertainty is high.

Technical Specifications and Availability

Aplomb 1 supports text, JSON, images, video, and audio inputs with a context window of up to 1M tokens. On the provider's API, it processes short questions in approximately 15 ms and reads 1M-token documents in about 3 seconds. Pricing is set at $0.02 per 1M input tokens with free output. The model weights are open and available on Hugging Face.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.