Local Multimodal Week
Key point
A roundup of new local and open-source multimodal models including Kimi K2.6 and Qwen3.6.
Details
This roundup gathers last week's highlights in local and open-source multimodal AI.
Moonshot Kimi K2.6 features a 1T/32B MoE architecture, 256K context, native INT4, and a 400M MoonViT vision encoder. Among its four variants, Agent Swarm supports 300 sub-agents and 4,000 coordinated steps, and the post states it scored HLE-Full with tools 54.0, surpassing GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro.
Alibaba Qwen3.6-35B-A3B is a natively multimodal model with 3B active / 35B sparse MoE. It defaults to 262K context, extendable to 1.01M via YaRN, and was released under Apache 2.0. Reported performance figures include SWE-Bench Verified 73.4, Terminal-Bench 2.0 51.5, AIME 2026 92.7, and VideoMMMU 83.7. The newly added Thinking Preservation feature retains reasoning traces across conversations.
Tencent HY-World 2.0 was introduced as an open-source 3D world model that produces editable mesh, 3DGS, and point cloud outputs that can be directly imported into Unity, Unreal, Blender, and Isaac Sim. Its earlier WorldMirror 2.0 component operates with roughly 1.2B params, BF16, and 12-24 GB VRAM.
Motif-Video 2B is an open-source video model based on a 2B DiT, handling both T2V and I2V from a single checkpoint at 720p / 121 frames. It claims the best open-source score with VBench Total 83.76%, outperforming Wan2.1-14B with 7x fewer parameters. However, a caveat notes that in blind testing, Wan2.1-14B performed better on temporal stability and human detail.
AniGen, presented at SIGGRAPH 2026, generates a fully rigged 3D model from a single image, jointly producing shape, skeleton, and skinning via S³ Fields to improve rigging consistency.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.