Prime Inference Launches Fast, Reliable Serving for Frontier Open Models
Key point
The platform processes approximately 1 trillion tokens daily internally and offers serverless and reserved capacity options.
Details
Prime has launched Prime Inference, a new serving platform designed to complete its continuous learning loop by deploying frontier open models to real users. The service supports both serverless endpoints for variable demand and reserved capacity for persistent workloads, running on NVIDIA Blackwell infrastructure with Vera Rubin coming soon. Internally, the system has been processing approximately 1 trillion tokens per day for large-scale RL rollouts and synthetic data generation prior to its public release.
Architecture and Performance
The platform separates the public API from the model fleet to allow seamless capacity shifts and failovers. It utilizes NVIDIA Dynamo for orchestration and vLLM for execution, implementing prefill/decode disaggregation to optimize agent workloads. This architecture reduces p90 inter-token latency by approximately 40% compared to shared GPU setups. The system also employs Mooncake for host DRAM caching to preserve conversation history and reduce re-computation.
GLM-5.3 Deployment
The first public model available is GLM-5.3, live on OpenRouter since September 22. It is currently one of the fastest GLM-5.3 endpoints on the platform, achieving 100% uptime since launch and a near-zero tool-call error rate. The deployment targets 100 tokens per second per user, with optimizations achieving 101 tok/s/user and 100 output tok/s per GPU at a prefill-to-decode ratio of 1:4.
Technical Optimizations
- NVFP4 KV Compression: Reduces MLA cache row size from 576 to 352 bytes, increasing total cache capacity by approximately 50% without accuracy loss.
- BLHNC Layout: A block-major KV layout reduces transfer descriptors by 10x and mean transfer time by 47% compared to the previous layer-major layout.
- Reliable Tool Calls: Integration of xgrammar in vLLM enforces tool-call grammar during decoding, eliminating silent failures and ensuring schema conformance for agent actions.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.