Audio-Native LLM Based on GigaChat 3.1 Released
·2026.07.26 18:59
Key point
GigaChat Audio 10B, an audio-native LLM built on GigaChat 3.1 Lightning, has been released.
Details
GigaChat Audio 10B, an audio-native LLM built on the GigaChat 3.1 Lightning text model, has been released on Hugging Face.
This model uses a Conformer speech encoder and a modality adapter to feed audio embeddings directly into a Mixture-of-Experts (MoE) decoder, adding speech understanding capabilities while preserving the quality of the original text model.
Key Features:
- Audio Question Answering and Classification: Directly understands audio data to answer questions or classify it
- Temporal Grounding: Locates specific events within long audio or performs summarization with timestamps
- Supports tool-use and text-only tasks
- TimeGround-1M Dataset: Trained using a dedicated dataset pairing long-form audio with time-aligned annotations
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.