AI Briefing
KOSign in

Qwen3.8 27B | 1 x R9700: 262K context, half a million tokens of reusable cache, ~180 tok/s. And yes, let's talk about the "3-bit" :)

·2026.10.04 22:12

Key point

New mixed-precision 3-bit weights for Qwen3.8 27B enable a 569,878-token reusable cache on a single AMD R9700, accelerating agent workflows by up to 12x for first-token latency with a 2.9-point MMLU-Pro drop vs MXFP4.

Details

A new release of Qwen3.8 27B for the AMD Radeon AI PRO R9700 (32 GB VRAM) introduces mixed-precision 3-bit weights, delivering 262,144 tokens of context per request and 569,878 tokens of reusable prefix cache. This configuration accelerates coding agent workflows, reducing time-to-first-token by up to 12x in multi-agent scenarios and speeding up whole-session completion by 6.1x compared to previous non-cached modes.

Quantization Details

The release is not a blanket 3-bit quantization. Only the large projection matrices (MLP, attention, and Gated DeltaNet recurrent projections), comprising 24.3B of the 27B parameters, are quantized to 3-bit.

  • Method: Matrices are rotated to spread outliers and GPTQ-calibrated on ~293K tokens.
  • Activations: Token generation uses FP8 activations; prompt reading uses 4-bit activations on rotated matrices (W3A4).
  • Other Components: Embeddings, output head, and norms remain in MXFP4 precision.
  • Storage: The 3-bit weights are a 9.55 GB add-on to the existing MXFP4 download.

Performance and Accuracy

Compared to the previous MXFP4 release, the 3-bit mode offers a trade-off between knowledge recall and speed/cache capacity:

  • Speed: Decode speed is within noise of the previous release (single stream ~156 tok/s, aggregate ~428 tok/s for 8 requests).
  • Cache: KV cache capacity increases to 569,878 tokens (vs. 174,634 in MXFP4).
  • Accuracy:
    • GSM8K: 95.45 (vs. 95.68 for MXFP4)
    • HumanEval: 92.68 (vs. 95.12 for MXFP4)
    • MMLU-Pro: 59.93 (vs. 62.57 for MXFP4)
    • The 2.9-point drop in MMLU-Pro is the only metric outside statistical noise, reflecting the cost of lower precision on knowledge recall tasks.

Agent Workflow Impact

The new --mode long-kv4 leverages prefix caching to drastically reduce latency for multi-turn agent sessions:

  • 20-turn conversation: Whole session time drops from 1,237 s to 204 s (6.1x faster).
  • Average wait for first token: Drops from 63.6 s to 7.8 s (8x faster).
  • Multi-agent repo sharing: 3 agents sharing a 100K-token repo see a 5.9x speedup in whole-session time.
  • Cached retrieval: A new question about a previously read 258K-token document takes 2.7 s instead of 134 s.

Availability and Requirements

  • Implementation: Runs on stock vLLM with a plugin (paiton-vllm-plugin).
  • System RAM: Requires 16 GB of system RAM to pin the embedding table. If RAM is below 13.5 GiB, the table stays on GPU, reducing cache to 451,879 tokens.
  • 512K Mode: An opt-in --mode long-512k extends context to 524,288 tokens per request but requires the full system RAM and has slower cold reads (~6 minutes for 500K tokens).
  • Limitations: Vision support is not yet available in the coding or 512K modes.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.