AI Briefing
KO

Introducing the SWE-Lancer Benchmark

·2025.02.18 19:00

Key point

OpenAI has released **SWE-Lancer**, a software engineering benchmark based on real Upwork freelance jobs.

Details

SWE-Lancer is a benchmark built from over 1,400 real freelance software engineering tasks from Upwork, with a total value of approximately $1 million.

The benchmark includes two types of tasks:

  • Individual contributor tasks: ranging from $50 bug fixes to $32,000 feature implementations, evaluated with end-to-end tests that are triple-verified by experienced engineers.
  • Managerial tasks: where a model chooses among proposed technical implementations, evaluated by comparing against the decisions of real engineering managers.

Evaluation results show that even the latest Frontier LLMs still fail to solve most of the tasks. To enable research, a unified Docker image and a public evaluation dataset, SWE-Lancer Diamond, are being open-sourced.

A recent update removed the internet connectivity requirement during execution, minimizing variability in model performance.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.