AI Briefing
KO

Qwen3.6-35B-A3B generates a better pelican image than Claude Opus 4.7

·2026.04.17 19:32

Key point

A locally run Qwen3.6-35B-A3B drew a better pelican SVG than Opus 4.7.

Details

Comparing Qwen3.6-35B-A3B and Claude Opus 4.7 by generating an image of "a pelican riding a bicycle," the Qwen result run locally turned out more polished.

Qwen used Alibaba's latest model in the Unsloth 20.9GB quantized version, run on a MacBook Pro M5 with LM Studio and the llm-lmstudio plugin. Claude Opus 4.7, on the other hand, had errors in how it rendered the bicycle frame, and using thinking_level: max brought almost no improvement.

For verification, a "flamingo riding a unicycle" was also tried, and Qwen again produced a better result. The author was impressed by details in the SVG such as <!-- Sunglasses on flamingo! -->.

However, the conclusion isn't overstated. This comparison is an extension of a joke test called the pelican benchmark, and doesn't mean Qwen is generally better than Anthropic's latest commercial model. Still, it shows that for specific SVG illustration tasks, at this point Qwen3.6-35B-A3B, which can run locally, may be more practical.

Opinions were divided in the HN comments over this result.

  • Some rated the Qwen result as more artistic, while others said Opus is better at physical realism.
  • Some thought this test might reveal overfitting to training data.
  • Other comments pointed out that such benchmarks don't really show real generalization performance, and that directly comparing a small local model to a frontier model is unfair.
  • Nevertheless, this case symbolically shows the progress of local LLMs and the narrowing gap with large commercial models.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.