EmDash Integrates Cloudflare Clef for Multimodal Plugin Registry Moderation
Key point
The EmDash plugin registry now uses Cloudflare's Clef decision model to automatically screen listings for phishing, impersonation, and offensive content across both text and images.
Details
EmDash has integrated Cloudflare's Clef decision model into its official plugin registry to automatically moderate new listings. The system screens metadata, links, icons, and screenshots for risks such as phishing, impersonation, scams, and offensive content before they appear in the catalog. This change allows EmDash to maintain a secure registry while leveraging the decentralized nature of the AT Protocol, where authors publish signed releases from their own accounts.
How Clef Works in EmDash
Clef is a decision model designed to answer bounded questions with typed probabilities, supporting both text and image inputs. When a plugin is published, the EmDash labeler queries Clef with:
- Nine questions regarding text and links, covering explicit content, hate speech, phishing, and moderation manipulation.
- Eight questions for each icon and screenshot, evaluating visual assets for similar risks.
Links are analyzed based on their text and destination rather than screenshots to avoid false positives from legitimate security tools. If Clef returns a probability of 0.45 or higher for any category, the listing is flagged for human review. The model does not have the authority to block plugins outright; it only triggers manual inspection.
Performance and Evaluation
EmDash replaced a previous multi-model pipeline that required separate passes for text and images. Internal evaluations showed Clef outperformed the baseline and other tested models like Jev, which underperformed. Key metrics from the evaluation include:
- 100% accuracy on 21 public text fixtures across 63 runs, with no invalid outputs.
- 1.64 seconds end-to-end text moderation latency at p95.
- Successful screening of all 37 plugin profiles and 43 listing images currently live in the registry.
The integration simplifies the infrastructure by using a single multimodal model for all moderation tasks, reducing latency and potential points of failure while maintaining a fail-closed path to human review.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.