AI Briefing
KO

Codestral Mamba

·2024.07.16 09:00

Key point

Mistral AI unveiled Codestral Mamba, based on the Mamba architecture capable of linear inference and unlimited sequence handling.

Details

As part of its new architecture research, Mistral AI has unveiled Codestral Mamba. This model was designed with the help of Albert Gu and Tri Dao, and is freely available for use, modification, and distribution.

Unlike existing Transformer models, Mamba models offer the advantage of linear time inference and can theoretically model sequences of infinite length. This efficiency is highly advantageous for code productivity tasks, and the model shows performance on par with the latest SOTA Transformer-based models.

Key features and specifications are as follows:

  • Supports in-context retrieval performance for up to 256k tokens
  • An Instructed model with approximately 7.28 billion (7.28B) parameters
  • Provided under the Apache 2.0 license for free use

Users can deploy the model via the mistral-inference SDK, TensorRT-LLM, and HuggingFace, and can also try it directly on Mistral's la Plateforme.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.