Agnes-3.0-Flash 33B Released, Featuring Hybrid Attention
Key point
The Agnes-3.0-Flash 33B model, released on Hugging Face, supports a hybrid attention architecture and a 262k context window.
Details
The Agnes-3.0-Flash 33B model has been released on Hugging Face. This model supports a context window of 262,144 tokens and features text, image, and video understanding as well as tool-calling capabilities.
Hybrid Attention Architecture
The core of the model is its hybrid attention structure. Of the total 72 layers, 54 apply the 'gated delta rule,' which uses recursive states independent of sequence length, while only the remaining 18 layers use standard global attention. This minimizes the KV cache, which grows with context length, thereby improving memory efficiency.
Key Technical Specifications
- Layer Configuration: 72 layers (54 delta rule + 18 global attention, 3:1 ratio)
- Hidden Size: 5120
- Attention Structure: Global attention uses 24 query heads / 4 KV heads (GQA); delta rule layers use 16 key heads / 48 value heads
- Position Encoding: 3-axis rotary (mrope) applied for text, height, and width
- Vision Tower: 27 layers, patch size 16, 2x2 spatial merging
Meanwhile, the benchmark site Artificial Analysis lists this model as a 'Proprietary model,' and discrepancies have been raised regarding data consistency, as the specifications, benchmark results, and context information differ from those on Hugging Face.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.