Building a Trace Judge with Fireworks That Cuts Costs by 100x
Key point
LangSmith partnered with Fireworks to fine-tune a Qwen model, developing a 'Perceived Error' detection model that cuts costs by 100x.
Details
As the volume of data generated by AI agents surges, the importance of trace data for understanding agent behavior is growing. LangSmith partnered with Fireworks to build a new Trace Judge that efficiently extracts critical signals from every trace.
The core of this project is detecting 'Perceived Error', which is recognized through user actions such as correcting or re-requesting an AI's response. This focuses on capturing the error users perceive in the AI's output, rather than objective correctness.
Key results include:
- Fine-tuned a Qwen model to achieve performance on par with or exceeding frontier models.
- Reduced operating costs by up to 100x compared to existing methods.
- Validated the model's generalizability using datasets from different domains, including chat-langchain and Fleet.
During data preparation, the team focused on Human and AI messages, which form the core of the information, while excluding tool calls to improve training efficiency.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.