AI Briefing
KO

Are AI Labs Overfitting to the Pelican Benchmark?

·2026.07.23 09:00

Key point

After generating and analyzing 1,008 SVGs, no evidence was found that AI labs are overfitting to the famous 'pelican riding a bicycle' benchmark.

Details

Simon Willison's years-running "draw a pelican riding a bicycle in SVG" prompt has become one of the most famous unofficial benchmarks in the AI community. Given the massive capital at stake for AI labs, there has been speculation that they might be doing specialized training targeting this benchmark (pelicanmaxxing).

To test this, a grid of 8 animals × 6 vehicles = 48 prompts was constructed, and generations were collected from 7 models — GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro — running each prompt 3 times, yielding 1,008 SVGs. Judging was done by GPT-5.6 Luna, which scored animal, vehicle, and motion consistency on a 1–5 scale each.

The analysis found no pattern where the 'pelican' row or 'bicycle' column scored significantly higher than other entries. No model showed a particularly strong pelican-bicycle combination, and models that drew well tended to draw other animal-vehicle combinations well overall too. The conclusion for now is that there is no evidence that AI labs have overfit to this benchmark.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.