AI Briefing
KO
Pick

Inkling-Small Released

·2026.08.05 15:30

Key point

Thinking Machines Lab released an open-weight model one-quarter the size of Inkling.

1 / 2

Details

Thinking Machines Lab released Inkling-Small, an open-weight multimodal MoE model with 276B total parameters and 12B active per token, under the Apache 2.0 license.

While its sibling model Inkling has 975B total and 41B active parameters, Inkling-Small is designed to be approximately one-quarter the size. It activates 6 out of 256 routing experts plus 2 shared experts per token, supporting a maximum 1M token context and text, image, and audio inputs.

The preview checkpoint adds on-policy distillation using Inkling as the teacher and 2 weeks of agentic coding reinforcement learning. As a result, it achieved 64.7% on Terminal Bench 2.1 and 80.2% on SWE-bench Verified, matching or surpassing the sibling model in some coding and reasoning evaluations.

The BF16 checkpoint requires large memory, but the NVFP4 checkpoint requires a minimum of 180GB, allowing it to run on a single NVIDIA B300, significantly improving accessibility. However, a Model Acceptable Use Policy applies separately from Apache 2.0.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.