SLM Upset Victory
Key point
By specializing 3B and 7B open-source SLMs with SFT+DPO, they achieved higher scores than commercial models.
Details
Dharma-AI specialized 3B and 7B open-source SLMs using SFT + DPO.
The comparison targets were GPT-5.4, Gemini 3.1 Pro, Claude Opus 4.6, Google Document API, and open-source alternatives OlmOCR, Deepseek-OCR, GLMOCR, and Qwen3.
The results were 7B 0.925 and 3B 0.911, higher performance scores than the presented comparison targets.
- DPO used rejection examples to reduce degenerate outputs, lowering the failure rate by up to 87.6%.
- Applying AWQ reduced inference cost per page by about 22%, with almost no quality degradation.
The paper, model, and dataset are all released together, making it worth referencing for reproduction experiments and follow-up specialization research.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.