Claude Opus 4.6/4.7 Fine-Tuning Dataset
·2026.05.01 15:13
Key point
A synthetic fine-tuning dataset of 8,706 examples based on Claude Opus 4.6/4.7 has been released.
Details
A synthetic fine-tuning dataset generated with Claude Opus 4.6/4.7 has been uploaded to Hugging Face.
- A total of 8,706 examples, an estimated 17,013,533 tokens, averaging 1,954 tokens/example
- Multi-turn accounts for 3,454 (39.7%), single-turn for 5,252 (60.3%)
- Split into
full_train.jsonl(all 28 categories),instruct_train.jsonl(24 instructional categories),roleplay_train.jsonl(4 creative categories), andcode_train.jsonl(coding+math) - The top categories are coding (1,628), humanities (862), and science (737), in that order
- It states that only basic cleaning was applied, and that the data was processed in a way that suppresses refusal and safety responses
This is a synthetic fine-tuning dataset worth checking out, broadly covering coding, math, science, humanities, and creative writing.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.