AI Briefing
Sign in

Datadog Releases Fine-Tuning Technique for Small LLMs for Root Cause Analysis

·2026.09.29 16:00

Key point

Datadog fine-tuned Qwen3.5-9B using GLM-5.3 investigation traces, reducing investigation costs by 1/20 while maintaining 87% performance.

1 / 5

Details

Datadog released a training pipeline for fine-tuning Qwen3.5-9B for Change Attribution analysis. The approach applies Knowledge Distillation, training the student model using investigation records from the teacher model, GLM-5.3.

Key Results and Cost Reduction

  • Performance Retention: Maintained accuracy at 87% of the teacher model's Recall@5 (0.63), achieving a score of 0.55.
  • Cost Reduction: Reduced LLM cost per investigation from $0.06 to $0.003, approximately 1/20 of the original cost.
  • Throughput: Capable of processing approximately 100,000 investigations per week on a single NVIDIA A100 GPU.

Training Pipeline Architecture

Datadog built a data flywheel based on a modified NVIDIA Data Flywheel Blueprint.

  • Proxy Label Generation: Generated labels without manual effort by extracting changes from Bits Investigation conclusions using GLM-5.3.
  • Teacher Agent Execution: Constrained GLM-5.3 to 8 turns, 5 tools, and a 65,536 token context window to perform investigations within the scope that the student model can emulate.
  • Rejection Sampling Fine-Tuning (RFT): Applied LoRA to Qwen3.5-9B using only teacher model traces that matched the proxy labels.
  • Iterative Loop: Expanded the dataset from 100 to 186 examples by adding cases where the teacher was correct but the student missed.

Performance Analysis and Generalization

  • Behavioral Change: After fine-tuning, the model shifted from broad searching to focused evidence collection. Search calls decreased, while log/span/metric collection calls increased.
  • Generalization Performance: Achieved a Recall@5 of 0.62 on customer incident data not used in training, showing significant improvement over the base model (0.52).
  • Cost Efficiency: Analysis indicated that the efficiency of investigation behavior (reduced token usage) was the primary factor in cost reduction, rather than the model size itself.

Limitations and Future Plans

  • Current evaluations are based on Bits Investigation conclusions and do not involve independent verification of actual causality.
  • Future plans include exploring Reinforcement Learning (RL) to achieve performance improvements beyond the teacher traces.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.