AI Briefing
KO

Anthropic Makes Safety Classifier Free

·2026.08.08 20:15

Key point

Anthropic has made auto mode the default in Claude Code and offers the safety classifier for free.

Details

An approach was proposed suggesting that layering model training, intent classifiers, and input probes for prompt injection defense can reduce success rates to near zero even against new attacks.

In a Claude Code update, Anthropic set auto mode as the default and announced that the classifier responsible for safety checks would be provided at no additional cost.

However, the post itself did not present specific experimental designs or figures to support the claim that prompt injection success rates reached zero in new attacks.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.