Skymizer Unveils 700B LLM Inference Card
Key point
The HTX301-based PCIe card supports 700B LLM inference at about 240W.
Details
Ahead of COMPUTEX 2026, Skymizer unveiled the HTX301 inference chip and the HyperThought architecture. The company said that by putting 6 HTX301 chips and 384GB memory into a single PCIe card, it can run 700B-parameter model inference locally at about 240W.
- HyperThought separates prefill and decode in LLM inference.
- The GPU handles compute-dense prefill, while HTX301 handles memory-bandwidth-heavy decode.
- The software coordinates prefill/decode pools with a KV-cache manager, a phase-aware scheduler, and a dynamic placement engine.
The platform scales from 1 chip to 6 chips and 32GB to 384GB, targeting models from 4B to 700B parameters. Skymizer explained that this reduces dependence on GPU clusters and high-power cooling infrastructure, enabling data sovereignty and low latency on-premises.
Detailed platform roadmap will be further disclosed at the COMPUTEX 2026 on-site presentation.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.