AI Briefing
KOSign in

openTPU: Open-Source AI Accelerator Developed by AI Runs LLMs on FPGA

·2026.10.07 01:23

Key point

The openTPU project presents an open-source AI accelerator developed by AI agents, capable of running models like LFM2.5-230M and Qwen3.5 on a Xilinx Kintex-7 FPGA with bit-exact simulation results.

1 / 3

Details

AI-Designed Hardware for AI Inference

The openTPU project presents an open-source AI accelerator explicitly developed by AI agents, exploring the limits of AI in hardware design. The system aims to answer whether AI can build the chip that runs its own inference. The repository contains the full stack, including SystemVerilog RTL, an instruction set architecture, a bit-exact simulator, a kernel compiler, and host software.

Performance on Inspur YPCB-00338 Card

The accelerator runs on an Inspur YPCB-00338 card featuring a Xilinx Kintex-7 xc7k480t FPGA and two DDR3 channels. It successfully executes ten modern models with real weights, producing outputs that match the simulator bit-for-bit. Key performance metrics for selected models (measured on the card) include:

  • LFM2.5-230M (int8): 59.0 tok/s decode
  • LFM2.5-230M (4-bit): 85.8 tok/s decode
  • Qwen3-0.6B (int8): 21.6 tok/s decode
  • Qwen3.5-0.8B (int8): 17.6 tok/s decode
  • Gemma 4 E2B (4-bit): 10.57 tok/s decode
  • LFM2-2.6B (int8): 6.05 tok/s decode
  • SmolLM3-3B (int8): 5.00 tok/s decode
  • Phi-4-mini (3.8B) (int8): 3.99 tok/s decode
  • Qwen3.5-2B (int8): 8.02 tok/s decode
  • Qwen3.5-4B (4-bit): 5.88 tok/s decode

The design prioritizes simplicity, with no cache or hidden scheduling. Decode performance is bound by DRAM efficiency, reading 82% to 85% of the DDR3-1066 peak bandwidth. 4-bit quantization increases decode speed by 40% to 45% compared to int8. The system also supports Mixture-of-Experts models like LFM2.5-8B-A1B, achieving 98.5% of peak bandwidth during decoding. The project is licensed under Apache License 2.0.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.