AI Briefing
KO

ToolGrad Achieves High Pass Rate with Answer-First Approach… Small Gemma-3 Models Show Competitive Performance Against SoTA Models

·2026.09.10 09:00

Key point

ToolGrad achieves a higher pass rate using an answer-first approach, while small Gemma-3 models demonstrate competitive performance against SoTA models.

1 / 5

Details

Google researchers proposed ToolGrad, a data generation framework that introduces the answer-first paradigm to address the low pass rates and inefficiencies of existing query-first methods. This method achieves higher pass rates by first generating successful tool-use chains and then post-annotating the corresponding user prompts.

Core Mechanism and Efficiency

ToolGrad extends the TextGrad concept to construct complex API workflows from large tool libraries. Through a four-stage module structure comprising API Proposer, API Executors, API Selector, and LLM Updater, it leverages textual gradients to generate low-cost, complex long-horizon tool-use data with a single LLM step. Efficiency is significantly improved due to a reduction in optimization steps compared to existing DFS-based approaches.

Performance Results and Model Derivatives

The Gemma-3 models (1B, 4B, 12B) were fine-tuned on the ToolGrad-500 dataset, generated based on 16k+ real-world APIs from ToolBench. In the Berkeley Function Calling Leaderboard (BFCL) evaluation, ToolGrad-12B scored 83.1, showing competitive performance against SoTA proprietary models such as Gemini-2.5-Pro (83.2) and Claude-4.5 Opus (82.8). Notably, the student models achieved higher performance than the teacher model used for data generation, Gemini-2.5-flash-lite, demonstrating the potential for self-evolving.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.