AI Briefing
KOSign in

Rebellions details Rebel100s interconnect hierarchy for AI inference

·2026.10.05 00:00

Key point

The Rebel100s accelerator integrates switch functionality into I/O chiplets to support up to 256 sockets without external switches.

1 / 3

Details

Rebellions has detailed the architecture of its Rebel100s accelerator, emphasizing that interconnect hierarchies are as critical as compute and memory for AI inference. While AI training is largely dominated by Nvidia and AMD, inference remains a wide-open market sensitive to latency and token generation costs. The Rebel100s aims to address this by creating a clean-slate hierarchy of on-die, in-package, and rackscale interconnects.

Rebel100 Architecture

The initial Rebel100 accelerator, previously known as the Rebel Quad, consists of four NPU chiplets linked by UCIe-Advanced interconnects. The chiplets are manufactured using Samsung’s SF4X process (4nm). Each Rebel-Single chiplet contains 16 neural cores with tensor and vector math units, 4 MB of L1 SRAM, and 32 L2 memory slices connected by a mesh network with 16 TB/sec bandwidth.

The Rebel100 package includes one integrated silicon capacitor (ISC) and one twelve-high stack of HBM3E memory per chiplet, resulting in an aggregate capacity of 144 GB and 4.8 TB/sec of bandwidth. The custom UCIe implementation includes a streaming protocol layer and multipathing to handle PHY failures. The four-chiplet complex delivers 1 petaflops at FP16 precision and 2 petaflops at FP8 precision.

Rebel100s and Integrated Switching

The follow-on product, the Rebel100s, adds a pair of Rebel IO chiplets to the complex. These I/O dies provide native Ethernet routing and encapsulate AXI memory protocols, allowing the system to scale up to 256 sockets in the first generation. This design absorbs switch functionality directly into the socket, eliminating the need for external switches for scale-up domains.

Key specifications for the Rebel100s include:

  • Bandwidth: The pair of Rebel IO chiplets delivers an aggregate scale-up bandwidth of 1,600 GB/sec, close to the 1,800 GB/sec offered by Nvidia’s Blackwell B200/B300 GPUs.
  • Connectivity: Each Rebel IO chiplet features two PCIe 6.0/CXL 3.0 controllers and four Ethernet ports supporting 800 Gb/sec scale-up connectivity.
  • Latency: An NPU-to-NPU hop across distinct sockets takes approximately 1 microsecond, with a target of 500 nanoseconds.
  • Synchronization: A custom hardware synchronizer eliminates handshake delays by sending sync signals as post-envelopes, improving inference performance.

Currently, Rebellions supports a single rack with 64 Rebel100s sockets using copper Ethernet cables. Expanding beyond this limit to four racks will require optical cables. The system supports flexible topologies, allowing memory server nodes to be configured without external switching hardware for smaller clusters.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.