AI Briefing
Feed
About
Search
KO
Sign in
All
Feed, trending, and board
Feed
Latest news
Trending
Open source and releases
Board
Blogs and showcases
Tags
Browse by topic
#swe-bench
The latest AI and developer news about #swe-bench, with the original source and a short summary.
Feed
Trending
Tags
Settings
Taste-Bench: Frontier LLMs Achieve 59.7% Accuracy on Decision Forking Tasks
TLDR AI
·
2026.09.25 09:00
UiPath Reveals Performance Gap Between Claude Models Based on Agent Task Length
Reddit
·
2026.09.22 21:00
OpenAI Recommends Halting SWE-bench Reporting; Solution is 'Evaluator-Led Verification'
Reddit
·
1
·
2026.09.20 23:00
Parallel Agents Backfire on Coding Tasks
Reddit
·
2026.09.05 08:00
Agent Lightning v1.0: Toward Harnessed Agentic RL
TLDR AI
·
2026.08.20 09:00
Ramp SWE-Bench (3 min read)
TLDR AI
·
2026.08.03 09:00
Comparing 13 models and 4 agents on SWE Tasks: Go, Java, Python, Rust, TS
Hacker News
·
2026.08.01 00:00
Distil, a Context Compression Tool with No Agent Performance Degradation, Released
Reddit
·
1
·
2026.07.26 08:00
SWE-Pruner Pro: Context Pruning Using Internal Agent Representations
Reddit
·
2026.07.21 19:00
Measuring Vulnerabilities in Tool-Using LLM Agents
TLDR AI
·
2026.06.26 09:00
Microsoft releases FastContext for code agents
Reddit
·
2026.06.23 09:00
Ramp SWE-Bench: A Private Coding Benchmark Based on Real Production Environments
TLDR AI
·
2026.06.15 09:00
Qwen3.7: The Agentic Frontier
Qwen
·
1
·
2026.05.20 11:00
Adaption Unveils AutoScientist, an AI Tool That Helps Models Learn on Their Own
TLDR AI
·
2026.05.14 09:00
Laguna XS.2 and M.1: A Deep Dive
TLDR AI
·
2026.04.29 09:00
Why SWE-bench Verified No Longer Measures Frontier Coding Ability
GeekNews
·
2026.04.27 18:00
SWE-bench Contamination Confirmed
Reddit
·
2026.04.27 03:00
mini-swe-agent - A 100-line AI agent for resolving GitHub issues and command-line support
GeekNews
·
2026.04.26 09:00
Kimi Vendor Verifier - Verifying Accuracy of Inference Providers
GeekNews
·
2026.04.23 03:00
Moonshot AI launches Kimi K2.6 in Kimi Chat and API
TLDR AI
·
2026.04.21 09:00
MoE flipped the script
Reddit
·
2026.04.18 02:00
Opus 4.7 Long-Context Regression
Reddit
·
2026.04.17 00:00
Benchmark distortion caused by gold-like answers
AI21 Labs
·
2026.04.14 20:00
How AI Agent Benchmarks Were Broken and What Comes Next
GeekNews
·
1
·
2026.04.12 21:00
Previous
1
2
Next
Previous
1
2
Next
#swe-bench | AI Briefing