Quail achieves 1 billion tokens per minute on single H100 GPU by combining query planner and inference engine
Key point
The new Quail system processes over 1 billion tokens per minute on a single H100 GPU, achieving more than 10x the speed of the vLLM baseline for AI-SQL workloads.
Details
Charles Frye (Modal) and Shreya Shankar (CMU FSD Lab) introduced Quail (QUery-Aware Inference Layer), a system that integrates a SQL query planner with an inference engine to optimize AI-SQL workloads. Unlike traditional agentic inference engines designed for arbitrary user requests, Quail targets large-scale sequence processing within databases, such as Snowflake Cortex AI-SQL, Databricks AI Functions, and BigQuery AI functions.
Performance and Architecture
Quail demonstrates significant efficiency gains, processing over 1 billion tokens per minute on a single H100 GPU. This represents a >10x speedup compared to the vLLM baseline, with costs dropping below 6 cents per billion tokens on Modal. In benchmarks using a newly released AI-SQL benchmark, Quail achieved a geometric mean speedup of 1.84x over vLLM.
The system addresses the specific needs of AI-SQL, where LLM-based filters and joins require massive matrix multiplications that benefit from GPU Tensor Cores. Key architectural components include:
- SQL Parser: Uses the
sqlglotlibrary to support Snowflake and BigQuery dialects. - Query Planner: Serializes plans via Substrait and applies custom logic for optimization.
- Execution Engine: A fork of vLLM modified with Triton for kernel fusion.
- Storage Engine: Utilizes
pyarrowunder a write-once/read-many assumption.
Key Technical Innovations
Quail’s performance stems from two primary optimizations in the query planner and execution engine:
- KV-Aware Join Ordering: The planner tracks Key-Value (KV) cache states from previous plans. It maintains candidate plans that are not dominated in terms of token count, attention pairs, and cached tokens, selecting the final plan based on a Speed-of-Light (SoL) cost model. This model uses hardware peak rates to estimate minimum latency.
- Decode Phase Elimination: For boolean classification tasks like
AI.IF, Quail predicts a single token during the prefill phase, eliminating the need for the decode step, sampling, and speculative decoding. This allows for specialized kernel fusions, such as combiningadd-RMSNormwithfp8quantization.
Future Directions
The authors outline several areas for future improvement, including KV Cache Tiering to utilize slower storage like tapes for latency-insensitive workloads, Cross-Query Optimization to preserve KV caches across requests, and handling Larger-than-Memory Datasets. They also plan to explore Radix Index structures for better prefix sharing and On-the-fly Fine-tuning to optimize models during execution.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.