Cerebras Serves Qwen 3.8 27B Model at 1,500 Tokens per Second
Key point
Cerebras has started serving the Qwen 3.8 27B model on its public endpoint at a speed of approximately 1,500 tokens per second.
Details
Cerebras is serving the Qwen 3.8 27B model via its public API endpoint, with an inference speed of approximately 1,500 tokens per second. This model has 27 billion parameters, is available on both free and pay-as-you-go tiers, and supports context lengths of 64k and 128k respectively.
Model Pruning and Quantization Policy
Cerebras specified that all models served on the public endpoint are unpruned versions in their original state. The REAP (Weight-Based Expert Activation Pruning) model, developed for research purposes, is available on Hugging Face but is not served on the production API.
Weight quantization is partially applied for memory efficiency, but computations are performed with high precision by dequantizing. In particular, attention, KV cache, and activations remain unquantized to preserve model performance.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.