FreedomIntelligence Releases Medical-Specialized LLM HuatuoGPT-3-9B
Key point
FreedomIntelligence has released HuatuoGPT-3-9B, a medical LLM based on Qwen3.5-9B that uses single-stage reinforcement learning.
Details
FreedomIntelligence has released the medical-specialized large language model HuatuoGPT-3-9B. This model is based on Qwen3.5-9B and adapts to the medical field by applying a single-stage policy optimization (OnePO) approach without prior domain-specific supervised fine-tuning (SFT).
The OnePO technique uses Teacher responses as temporary guides during the reinforcement learning stage, gradually removing them as model performance improves.
Along with this release, training code, a medical reinforcement learning dataset, and an 8B rubric evaluator have been distributed, allowing researchers and developers to verify and utilize this approach.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.