Meta and Google Model Safety Filters Removed in Minutes
Key point
Using the Heretic tool and abliteration technique, the safety guardrails of Meta's and Google's open-source models can be disabled within minutes.
Details
A joint test by the Financial Times and AI safety group Alice confirmed that the safety features of open-source models such as Meta's Llama 3.3 and Google's Gemma 4 can be removed in just a few minutes.
Using Heretic, a free tool on GitHub, safety filters can be stripped away through a technique called abliteration, which directly modifies the neural network's internal parameters. This approach forces the model to answer questions related to biochemical weapons or malicious code.
According to Philipp Emanuel Weidmann, the developer of the tool, more than 3,500 uncensored models have already been built, with a total of over 13 million downloads.
However, this technique cannot be applied to closed models such as OpenAI's ChatGPT or Anthropic's Claude, since their source code is not externally accessible.
Experts warn that the spread of such tools is making it difficult for governments and companies to regulate AI safety, and that society must prepare for a new type of threat.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.