FreedomIntelligence Releases HuatuoGPT-3-27B Medical LLM with One-stage Policy Optimization
Key point
The model uses a single reinforcement-learning stage without prior domain-specific supervised fine-tuning.
Details
FreedomIntelligence has released HuatuoGPT-3-27B, a medical large language model built on the Qwen3.8-27B base. The model introduces One-stage Policy Optimization (OnePO), a method that adapts language models to the medical domain in a single reinforcement-learning stage, eliminating the need for preceding domain-specific supervised fine-tuning.
Training Methodology
OnePO utilizes teacher responses to provide temporary guidance during training. These responses are gradually retired as the model’s performance improves, allowing it to learn directly from reinforcement signals rather than static supervised data.
Available Resources
The release includes the following open-source components:
- Training code for the OnePO method.
- A medical RL dataset used for the reinforcement learning phase.
- An 8B rubric grader for evaluation.
This follows the previous week's release of the smaller HuatuoGPT-3-9B model.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.