AI Briefing
KOSign in

Blockway Releases Agens Volundr 32B Preview with Hybrid Architecture Using Only 18 KV Cache Layers

·2026.10.05 21:58

Key point

The Apache-2.0 licensed model uses a hybrid architecture with 54 KDA layers and 17 BCSA layers to minimize KV cache usage while supporting a 262K context window.

Details

Blockway, a Hong Kong-based team, has released Agens Volundr 32B Preview, the first model built on their proprietary hybrid architecture. The model is designed to optimize memory usage for long-context tasks by ensuring that only 18 of 72 layers maintain a KV cache, addressing the bottleneck where KV cache size, rather than model weights, limits deployment on local hardware.

Hybrid Architecture Design

The model employs a 72-layer dense architecture where every layer processes every token, but with distinct attention mechanisms:

  • 54 KDA (Kimi Delta Attention) layers: Utilize linear attention with a fixed-size recurrent state, requiring no KV cache.
  • 17 BCSA layers: Implement compressed-sparse attention with an exact window over the last 4,096 tokens, pooling older context 4:1 into blocks and using a learned indexer to access the top 512 blocks.
  • 1 full-attention layer: Located at layer 72.
  • Engram: A hashed n-gram memory stored in host RAM, attached at two specific layers.
  • mHC: Uses 4 residual streams instead of the standard single stream.

Performance and Benchmarks

The model supports a 262K context window. In single-user tests using a custom sglang build, BF16 inference on two 48 GB GPUs achieved decode speeds of 25.1 tok/s at 1K context, remaining near 24.0 tok/s (24.1 at 8K-64K, 23.9 at 128K). INT4 quantization on a single 48 GB GPU reached 31.0 tok/s at 1K context.

In comparative benchmarks against Qwen3.8-27B, Agens Volundr 32B Preview showed improvements in coding and math tasks:

  • LiveCodeBench v6: +4.2 points
  • HumanEval: +4.3 points
  • AIME 2025: +2.9 points
  • MATH-500: +1.6 points

However, it trails in agentic tasks, scoring 74.2 on tau2-bench (vs. 79-80 for Qwen) and 44 on SWE-bench Verified (50-task subset) (vs. 58-64). Blockway notes that closing this gap is the primary focus for the full v1 release.

Availability and Limitations

The model is released under the Apache-2.0 license. It currently requires a specific sglang build provided by Blockway, as stock sglang and vLLM cannot load it yet. GGUF and llama.cpp support are planned but not currently available. The team emphasizes that this is a preview model still undergoing training, with the full version expected to continue pre-training to approximately 10B tokens and include training on long agentic sessions.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.