Jev 1.13.0 Achieves 76.2% SRE Diagnosis Pass Rate at 200x Lower Cost Than LLM Agents
Key point
The Jev 1.13.0 model achieved a 76.2% pass rate on SREGym-Lite benchmarks, performing close to GPT-5.6 Sol's performance while being approximately 7 times faster and 200 times cheaper.
Details
Performance and Cost Efficiency
The Jev 1.13.0 model, operating within a structured SRE diagnosis pipeline without an LLM agent, achieved a 76.2% diagnostic pass rate on the SREGym-Lite dataset (21 fault scenarios). This performance is close to GPT-5.6 Sol (medium), which scored 77.8%, but with significant efficiency gains:
- Speed: Approximately 7 times faster than LLM agent baselines.
- Cost: Approximately 200 times cheaper per diagnosis.
- Latency: Median diagnosis time of 14.6 seconds; median Jev API latency of 0.53 seconds.
- Resource Usage: Total input tokens of 3.48 million across 105 attempts, with an estimated inference cost of $0.15.
Pipeline Architecture
The system utilizes a programmatic Collector to gather Kubernetes objects, events, Pod logs, and resource usage, generating component-specific summaries. Jev then selects suspicious components and chooses core evidence from numbered options provided by the Collector. Jev does not generate commands or final reports; it only selects from provided options. The pipeline submits a diagnosis if evidence supports the hypothesis, otherwise it reviews other candidates.
Failure Modes and Limitations
While Jev passed 16 of 21 faults completely, it failed all 5 attempts on 5 specific faults. Key failure modes include:
- Incorrect Clue Selection: In the
edge_request_filter_cpu_saturationscenario, Jev incorrectly blamed a CPU limit change rather than the actual cause (a crafted WAF request triggering expensive regex), resulting in a score of 0.67. - Missing Critical Evidence: In
search_rate_retry_collapse_hotel_reservation, Jev missed the interaction loop between queue depth, deadlines, and retries because the Collector did not include retry setting values in the snapshot. Jev focused on the rate limit backend, failing to identify the cross-service interaction.
Strategic Implications
The study suggests that Jev is a promising first-line diagnostic tool due to its speed and cost efficiency, while LLM agents remain better suited for broad, flexible investigations. The primary design challenge lies in selecting the appropriate granularity of cluster state data: too coarse hides necessary details, while too fine buries useful signals. Future work aims to integrate Jev into SRE workflows as a "System One" fast-response component, potentially using smaller specialized models like GPT-6 Luna to convert structured telemetry into natural language for Jev.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.