AI Briefing
KO

FuriosaAI Wraps Up 2024 with Llama 3.1 Milestone

·2024.12.19 09:00

Key point

FuriosaAI's AI inference chip **RNGD** demonstrated outstanding throughput and power efficiency on **Llama 3.1** models.

Details

FuriosaAI's AI semiconductor solution RNGD (Renegade) closed out 2024 by recording strong performance metrics in real-world enterprise deployments using Llama 3.1 8B and 70B models.

RNGD processes 3,200-3,300 tokens per second (TPS) when running Llama 3.1-8B, while stably delivering 40-60 TPS even in single-user scenarios. It also demonstrates excellent power efficiency, consuming just 180W per card.

For the Llama 3.1-70B model, effective execution is possible with just 2 RNGD cards, and optimization is underway targeting 8,000 TPS on a server equipped with 8 cards.

On the software side, SDK v2024.3.0 supports tensor parallelism, torch.compile, and HuggingFace Optimum integration. Additionally, SDK v2024.1.0 provides developers with an optimized environment, including high-performance inference technologies such as PagedAttention, Block KV Cache, and Continuous Batching.

Furthermore, FuriosaAI strengthened its leadership to expand into the global market by bringing on Alex Liu, formerly of NETINT Technologies, as Senior Vice President (SVP) of Product and Business.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.