Sber Releases GigaChat-3.5-Reasoning
Key point
Sber released the MoE-based GigaChat-3.5-Reasoning, achieving DeepSeek V4 Flash-level performance while reducing inference token usage by 37%.
Details
Sber has released a new reasoning-specialized model, GigaChat-3.5-Reasoning. This model adopts a 432B-A28B MoE (Mixture of Experts) architecture and applies Gated DeltaNet technology to improve long-context processing efficiency.
Training Method and Performance
The model was developed by first training domain-specific expert models for code, math, and general tasks using the CISPO algorithm, then distilling them into a single model via On-policy distillation. Internal evaluations show that the model achieves performance similar to DeepSeek V4 Flash Preview while reducing token usage by 37% during inference.
Deployment and Accessibility
Model weights are available on Hugging Face under the MIT License, allowing anyone to freely use and modify them. Additionally, the model can be tested directly via the reasoning tab on the official website giga.chat.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.