AI Briefing
KO

4-Chiplet AI SoC With 16Gbps UCIe-Advanced Die-to-Die Interface and Full-Chip Scalable Mesh for Large-Scale AI Inference

·2026.01.02 12:38

Key point

A 4-chiplet AI SoC implements a full-chip mesh with 16Gbps UCIe-Advanced.

1 / 2

Details

Rebellions unveiled an AI SoC composed of 4 chiplets at ISSCC 2026. The core is a 16Gbps die-to-die interface based on UCIe-Advanced, combined with a full-chip scalable mesh that operates as one unit across chip boundaries.

This architecture targets large-scale AI inferencing, particularly workloads like LLM serving where compute and memory bottlenecks alternate. Rather than simply bundling chiplets together, it is designed to treat the entire package as a single scalable system, maintaining latency and coherence while securing scalability.

The features emphasized in the paper/whitepaper are as follows.

  • Unified mixed-precision compute: Processes FP8/FP16/FP32 together in a single core to increase compute density
  • Predictive DMA + on-chip mesh: Combines software-coordinated DMA with mesh to reduce memory access bottlenecks
  • Hierarchical synchronization: Coordinates distributed execution via a central sync manager and peer-to-peer communication

In numbers, 2.8x higher compute density, 2.7TB/s effective bandwidth, 2 PFLOPS (FP8), and a 3.3x faster mesh fabric are presented. Chip-to-chip connectivity is described as 1TB/s per channel with an inter-chiplet latency of 11ns, and each chiplet uses 3 UCIe channels to maintain horizontal mesh continuity.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.