AI Briefing
KO

Detecting and Responding to Misuse Cases of Claude (March 2025)

·2026.05.29 12:00

Key point

Anthropic disclosed cases of detecting public opinion manipulation and malicious activity using Claude, along with its response measures.

Details

Anthropic conducts continuous monitoring to prevent misuse of its Claude models, and recently shared detected misuse cases and the response process. In particular, due to advances in agentic AI technology, malicious actors are increasingly building more sophisticated systems.

The most notable case is the operation of an 'influence-as-a-service' service. These actors attempted large-scale public opinion manipulation by using Claude not merely as a content generation tool, but as an orchestrator that decides when SNS bot accounts should comment or share content.

In addition, the following misuse cases were identified:

  • Credential stuffing involving leaked account information related to security cameras
  • Recruitment fraud campaigns targeting job seekers in Eastern Europe
  • Low-skilled individuals using AI to create sophisticated malware

Anthropic is applying research techniques such as Clio and hierarchical summarization to detect these threats. Through this, it efficiently analyzes large-scale conversation data, and responds by using classifiers to block harmful requests or sanction related accounts.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.