AI Briefing
KO

Opus hallucination rate spikes

·2026.04.15 01:03

Key point

Across a 90-day tracking of 42 recurring tasks, Opus 4.6's hallucination rate surged.

Details

42 recurring tasks were run the same way over 90 days, logging hallucination rate changes by model.

The production setup was a RunLobster-hosted agent, with routing going Sonnet 4.6 (default), Opus 4.6 (escalation), Gemini 3 Flash (rate-limit fallback), in that order.

The key figures are as follows.

  • Jan 15 ~ Feb 14: Sonnet 0.24, Opus 0.09, Gemini 0.31
  • Feb 15 ~ Mar 14: Sonnet 0.27, Opus 0.11, Gemini 0.29
  • Mar 15 ~ Mar 31: Sonnet 0.29, Opus 0.14, Gemini 0.28
  • Apr 1 ~ Apr 13: Sonnet 0.31, Opus 0.38, Gemini 0.27

The author found that Opus 4.6's hallucinated-specific-per-briefing rate jumped by about 2.7x from mid-March into early April. Over the same period, bridgebench also showed a 83.3 → 68.3 shift, which the author noted moved in the same direction as their own logs.

Real examples cited include the following errors.

  • Quoting a nonexistent CEO statement
  • Asserting an incorrect Series B timing and amount
  • Misclassifying a Stripe payout and even fabricating a customer name

The main point is that this isn't simply tone or judgment errors, but errors that fabricate specific facts immediately verifiable against sources, which has persisted for over the last two weeks.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.