AI Briefing
Sign in

Quail achieves 1 billion tokens per minute on single H100 GPU by combining query planner and inference engine

·2026.09.24 09:00

Key point

The new Quail system processes over 1 billion tokens per minute on a single H100 GPU, achieving more than 10x the speed of the vLLM baseline for AI-SQL workloads.

1 / 4

Details

Charles Frye (Modal) and Shreya Shankar (CMU FSD Lab) introduced Quail (QUery-Aware Inference Layer), a system that integrates a SQL query planner with an inference engine to optimize AI-SQL workloads. Unlike traditional agentic inference engines designed for arbitrary user requests, Quail targets large-scale sequence processing within databases, such as Snowflake Cortex AI-SQL, Databricks AI Functions, and BigQuery AI functions.

Performance and Architecture

Quail demonstrates significant efficiency gains, processing over 1 billion tokens per minute on a single H100 GPU. This represents a >10x speedup compared to the vLLM baseline, with costs dropping below 6 cents per billion tokens on Modal. In benchmarks using a newly released AI-SQL benchmark, Quail achieved a geometric mean speedup of 1.84x over vLLM.

The system addresses the specific needs of AI-SQL, where LLM-based filters and joins require massive matrix multiplications that benefit from GPU Tensor Cores. Key architectural components include:

  • SQL Parser: Uses the sqlglot library to support Snowflake and BigQuery dialects.
  • Query Planner: Serializes plans via Substrait and applies custom logic for optimization.
  • Execution Engine: A fork of vLLM modified with Triton for kernel fusion.
  • Storage Engine: Utilizes pyarrow under a write-once/read-many assumption.

Key Technical Innovations

Quail’s performance stems from two primary optimizations in the query planner and execution engine:

  1. KV-Aware Join Ordering: The planner tracks Key-Value (KV) cache states from previous plans. It maintains candidate plans that are not dominated in terms of token count, attention pairs, and cached tokens, selecting the final plan based on a Speed-of-Light (SoL) cost model. This model uses hardware peak rates to estimate minimum latency.
  2. Decode Phase Elimination: For boolean classification tasks like AI.IF, Quail predicts a single token during the prefill phase, eliminating the need for the decode step, sampling, and speculative decoding. This allows for specialized kernel fusions, such as combining add-RMSNorm with fp8 quantization.

Future Directions

The authors outline several areas for future improvement, including KV Cache Tiering to utilize slower storage like tapes for latency-insensitive workloads, Cross-Query Optimization to preserve KV caches across requests, and handling Larger-than-Memory Datasets. They also plan to explore Radix Index structures for better prefix sharing and On-the-fly Fine-tuning to optimize models during execution.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.