OpenAI Partners with Ironclad to Train GPT-6 Astra on Complex Contracting Workflows
Key point
GPT-6 Astra scored 55.0% on Ironclad research tasks, a 32% improvement over GPT-5.6 Sol, while reducing estimated time per attempt by 48%.
Details
OpenAI has partnered with Ironclad, a leader in AI contracting, to develop training and evaluation tasks for GPT-6 Astra. The collaboration focuses on teaching agents to handle complex business workflows, such as configuring agreements, approvals, and reusable legal terms, ensuring they meet original requirements across multi-step processes.
Benchmark Performance
GPT-6 Astra is the first frontier model trained on these Ironclad-specific tasks. In a research evaluation comparing it to GPT-5.6 Sol, Astra demonstrated significant gains in both accuracy and efficiency:
- Mean rubric score: Astra achieved 55.0% compared to Sol's 41.6%.
- Estimated time per attempt: Astra averaged 19.2 minutes, down from 37.0 minutes for Sol.
- Improvement: Astra's average score was 32% higher than Sol's, with estimated time per attempt 48% lower.
An internal model used during Astra's development reached an even higher score of 63.7%, indicating potential for future model iterations.
Task Definition and Methodology
Ironclad employees and OpenAI users identified 11 tasks spanning legal, commercial, and procurement work. These tasks, such as setting up nondisclosure agreements and updating jurisdiction-specific clauses, typically take an experienced human user 30 to 40 minutes to complete. Each task was evaluated against 8 to 50 criteria to assess specific successes and failures.
To train the models, researchers used synthetic training tasks derived from contracts in the SEC’s EDGAR database, filtered to remove personal information. No OpenAI customer data or nonpublic Ironclad customer data was used. The models practiced in hosted Ironclad software environments, using reinforcement learning to improve through feedback.
Implications for Professional Work
The partnership highlights the challenge of maintaining business rules and controls while automating complex workflows. If an agent loses track of a rule, such as a spending threshold for finance approval, the entire process fails. This underscores the continued need for human oversight and robust platforms like Ironclad, even as AI agents become more capable.
OpenAI is inviting other software companies to join this research initiative. Partners must provide concrete examples of agent failures, secure testing environments, and deep domain expertise to help define success criteria for professional AI tasks.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.