AI Briefing
KO

Tensor Contraction Processor: AI Chip Architecture

·2024.06.28 09:00

Key point

FuriosaAI has unveiled the TCP architecture and RNGD chip to overcome the power efficiency and programmability limits of GPUs.

Details

AI hardware must simultaneously achieve high programmability and power efficiency to meet explosive computational demand. However, existing GPUs have limitations: they are difficult to optimize as models change, and they consume massive power exceeding 1,000W per chip.

To solve this, FuriosaAI developed the Tensor Contraction Processor (TCP) architecture. TCP is designed around core AI operations, efficiently managing data and memory, and powers high-performance generative AI models like Llama 3 through its next-generation chip, RNGD (Renegade).

TCP is co-designed with the software stack to provide high programmability. In particular, through a universal compiler that processes the entire model as a single fused operation, even models with new architectures can be automatically optimized and deployed.

Data movement consumes up to 10,000 times more energy than the computation itself. TCP maximizes energy efficiency by reusing data stored in on-chip memory to minimize data movement. RNGD is scheduled for release later this year.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.