AI Briefing
KOSign in

Dotwave Engine Serves 56 Concurrent Full-Duplex Voice Sessions on One H100, Cutting Costs by 98%

·2026.10.02 22:55

Key point

The new inference engine achieves 56 concurrent sessions on a single NVIDIA H100 SXM 80 GB by eliminating CPU-GPU overhead and batching continuous inference workloads.

Details

Real-time AI models, such as full-duplex voice agents, require continuous inference where the model listens and speaks simultaneously, maintaining state and meeting strict 80 ms deadlines for every audio frame. Traditional request-based inference servers like vLLM, SGLang, TensorRT-LLM, and Triton struggle with this workload because they rely on batching that cannot wait for a full session to end, leading to 75% GPU idle time in NVIDIA's reference stack for Nemotron VoiceChat 11B. In that reference setup, a single session occupies an entire H100, and adding a second session causes 8–15% of audio beats to arrive late.

The Solution: Persistent GPU Programs

Dotwave's inference engine addresses these inefficiencies by keeping the model resident on the GPU as a single program, eliminating the CPU-GPU coordination overhead for every beat. Key optimizations include:

  • Batched Session Advancement: All live sessions due on the same tick run as one batch, reading model weights once for the group rather than per session.
  • Pre-compiled Scheduling: The compiler fixes the schedule and memory layout ahead of time, removing runtime scheduler and allocator delays.
  • Continuous State Management: The engine handles continuous input and state retention without freeing resources mid-call.

Performance Results

On a single NVIDIA H100 SXM 80 GB, the engine achieved:

  • 56 concurrent sessions, compared to 1 for the reference stack.
  • 147.4–147.5 ms p99 latency per 160 ms beat.
  • Zero missed deadlines across 84,000 measured session-beats.
  • Approximately 98% reduction in GPU cost per conversation.

These measurements exclude network and audio playback latency and were conducted using recorded input with about two minutes of context. The engine is available for testing with Nemotron VoiceChat 11B and Nemotron 3.5 ASR Streaming 0.6B, offering free usage credits upon sign-up.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.