OpenAI and Ironclad train agents on complex contracting workflows

OpenAI and Ironclad have teamed up to train AI agents on real-world contracting workflows, converting complex legal, commercial and procurement processes into structured tasks that models can practice and be measured against. The collaboration focuses on teaching agents to preserve business rules, execute multi-step procedures inside specialized software, and verify outcomes against detailed success criteria.

The project centered on 11 representative tasks drawn from typical contract and procurement operations managed by legal operations teams. Examples include configuring purchase approval processes that route requests above spending thresholds to Finance, directing certain requests to Security, and flagging nonstandard contract terms for Legal review. To succeed, an agent must reliably maintain these rules through every step of the workflow and handle a range of realistic scenarios.

Ironclad contributed product environments and domain expertise so OpenAI researchers could build training and evaluation problems that mirror customer-facing priorities. Subject-matter experts and Ironclad users helped translate everyday activities—such as nondisclosure agreements, procurement approval setups and updating reusable clauses tied to requester jurisdiction—into tasks that could be judged automatically. Each task was evaluated against between eight and 50 criteria, depending on complexity, enabling precise analysis of which parts of a workflow a model handled correctly and where it failed.

To give models practice inside realistic software, Ironclad provided hosted instances of its product and OpenAI created synthetic training tasks reflecting common workflows. Researchers estimated the time an experienced human would spend on each task at roughly 30 to 40 minutes. Reinforcement learning with feedback was then used so agents could iteratively improve by practicing in these controlled environments and receiving corrective guidance when they missed criteria.

OpenAI evaluated its latest GPT-6 model, Astra, against the prior GPT-5.6 Sol baseline using each model’s preferred settings. Across the 11 tasks, Astra achieved an average research evaluation score of 55.0 percent compared with 41.6 percent for GPT-5.6 Sol. Estimated average time per attempt fell materially as well: Sol averaged 37.0 minutes per attempt while Astra reduced that to 19.2 minutes. An internal development model used during Astra’s training registered 63.7 percent on the same evaluations.

The company also provided side-by-side examples showing qualitative differences in performance. In those examples Astra met roughly 94 percent of task criteria in an estimated 20 minutes, while GPT-5.6 Sol met about 85 percent of criteria in an estimated 32 minutes. OpenAI emphasized that Astra both met more task requirements and used less simulated time per attempt, highlighting progress toward agents that can operate within and across specialized business software environments.

The collaboration surfaces a persistent challenge for contracting platforms: handling exceptions without sacrificing the rules businesses rely on. If an agent loses track of required controls, that limits what a platform can safely automate. OpenAI and Ironclad underscore that human oversight and robust platform controls remain essential even as agents become more capable at handling routine and semi-structured work.

For Ironclad, the improved model performance points to an opportunity to automate more of the routine work performed by legal and business teams, provided models can preserve necessary controls and manage edge cases. Working directly with OpenAI gives Ironclad a way to surface practical failure modes to researchers and influence improvements in model behavior that matter for real customers.

OpenAI is inviting a limited number of software companies to collaborate with its research and engineering teams on professional tasks that current agents cannot reliably complete. Interested partners are asked to supply concrete examples of the tasks they want automated, evidence of where existing agents fail, and clear criteria for measuring success. They should also provide people with deep domain knowledge, a secure testing environment, and data that can be safely used for research.

The Ironclad collaboration illustrates an incremental step toward AI systems that can perform structured work inside specialized applications. By turning real contracting activities into measurable research tasks, OpenAI and Ironclad have produced a clearer picture of how far models have advanced and what remains necessary—both in modeling and in platform design—to safely expand automation in legal and procurement workflows.

As the companies continue to refine agents and invite additional software partners, the work provides a practical template for evaluating models against the kinds of multi-step, rules-driven tasks that underpin much professional work. For organizations weighing the risks and benefits of automation, the results underscore both the promise of faster, more accurate assistance and the continued need for human governance where rules and exceptions matter.

Source: Read the original source

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *