AI Briefing
KO

Function Calling Backend Benchmark

·2026.05.03 22:59

Key point

The function calling harness has narrowed the frontier-local gap in backend generation.

Details

The function calling benchmark for open-source backend coding agents has been rebuilt with controlled variables and scoring criteria in place.

Unlike previous informal measurements, this version aims for reproducible comparison.

The core conclusion is that the function calling harness has effectively narrowed the frontier-local gap in backend generation.

  • For DB/API design, gpt-5.4 and qwen3.5-35b-a3b performed similarly.
  • For logic, claude-sonnet-4.6 was rated at the level of qwen3.5-27b.

Additionally, gpt-5.4 scored lower than its mini version, and deepseek-v4-pro also stayed below qwen3.5-35b-a3b. Additional reversals were observed within the Qwen family as well, but interpretation is still ongoing.

The operating scope is also changing.

  • This is the last inclusion of frontier models.
  • A single run involves 200-300 million tokens, costing $1,000-$1,500 per model at GPT-5.5 pricing.
  • Starting next month, comparisons will be limited to models running on OpenRouter endpoints priced at $0.25/M or less or on laptops with 64GB unified memory.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.