AI Briefing
KO

Cerebras Serves Qwen 3.8 27B Model at 1,500 Tokens per Second

·2026.09.04 06:23

Key point

Cerebras has started serving the Qwen 3.8 27B model on its public endpoint at a speed of approximately 1,500 tokens per second.

Details

Cerebras is serving the Qwen 3.8 27B model via its public API endpoint, with an inference speed of approximately 1,500 tokens per second. This model has 27 billion parameters, is available on both free and pay-as-you-go tiers, and supports context lengths of 64k and 128k respectively.

Model Pruning and Quantization Policy

Cerebras specified that all models served on the public endpoint are unpruned versions in their original state. The REAP (Weight-Based Expert Activation Pruning) model, developed for research purposes, is available on Hugging Face but is not served on the production API.

Weight quantization is partially applied for memory efficiency, but computations are performed with high precision by dequantizing. In particular, attention, KV cache, and activations remain unquantized to preserve model performance.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.