AI Briefing
KO

Estimating Worst-Case Frontier Risk of Open-Weight LLMs

·2025.08.05 09:00

Key point

Malicious fine-tuning of gpt-oss was used to assess the biological and cybersecurity risks of open-weight LLMs.

Details

To study the worst-case frontier risks that could arise upon the release of gpt-oss, we introduced a technique called Malicious Fine-Tuning (MFT). This method aims to train the model to maximize its biology and cybersecurity capabilities.

To maximize risk, we built the following environments:

  • Biorisk: We curated tasks related to threat creation and trained the model in an RL (reinforcement learning) environment with web browsing capability.
  • Cybersecurity risk: We trained the model to solve CTF (Capture-the-Flag) challenges in an agentic coding environment.

The evaluation results showed that MFT gpt-oss performed lower than OpenAI o3, a frontier closed-weight model. For reference, o3 possesses capabilities below the 'Preparedness High' threshold in both biology and cybersecurity.

Compared to existing open-weight models, gpt-oss showed a marginal improvement in biological capability but did not substantially advance the frontier level. Based on these results, the research team decided to release the model, and hopes that the MFT approach can serve as a useful guide for estimating the risks of future open-weight models.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.