Read Ironclad’s Fine Print On OpenAI Training Agents In Your Software
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Read Ironclad’s Fine Print On OpenAI Training Agents In Your Software on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI described training GPT-6 Astra in hosted copies of contract-management software from Ironclad, using 11 legal, commercial and procurement tasks. Astra met an average 55% of task criteria, while OpenAI’s estimated completion times were simulated and the results do not establish that the agent can safely handle customer workflows without human review.

OpenAI said on October 6 that it trained its GPT-6 Astra model in hosted copies of contract-management software from Ironclad, testing whether an AI agent could carry out multi-step legal, commercial and procurement work. Across 11 tasks, Astra met an average of 55% of the listed criteria; OpenAI cautioned that its time estimates were simulated, and the results do not show that agents can reliably run such workflows without human oversight.

The tasks were selected by Ironclad staff and OpenAI employees who use the product. They included creating nondisclosure agreements, configuring procurement approval processes and changing a reusable contract clause according to a requester’s selected jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes to complete each task.

Each task was assessed against a rubric of 8 to 50 criteria, depending on its complexity. OpenAI reported that GPT-6 Astra met an average 55.0% of criteria, compared with 41.6% for GPT-5.6 Sol in a high-compute setting. An internal model used during Astra’s development reached 63.7%. On one highlighted task, Astra met about 94% of the criteria, but that example does not represent the overall average.

OpenAI said Ironclad supplied hosted product environments for model practice. The company said it created synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database, filtered to remove personal information. It stated that the work did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.

At a glance
reportWhen: Published October 6; results and partne…
The developmentOpenAI published results from a collaboration with Ironclad that used the contract-management product as a training environment for its frontier model.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Contract Workflow Accuracy Matters

The reported 55% is an average share of rubric criteria met, not the percentage of tasks completed successfully. That distinction matters in contract and procurement systems, where a missing requirement can invalidate the practical value of otherwise correct work. A workflow might need Finance approval above a spending threshold, Security review for certain requests and Legal review for nonstandard terms. Missing even one required check could send a purchase through without the intended control.

OpenAI’s post acknowledges that agents can lose track of business rules during multi-step work and says human oversight remains important. The figures do not demonstrate dependable autonomous contract processing, nor do they establish measured savings for Ironclad customers. They do show a test of a model inside specialised business software, where success depends on following detailed rules rather than simply producing plausible text.

The collaboration also has implications for software companies. Training agents on product-specific workflows could make those agents more useful to customers, while exposing where they fail. If agents increasingly become the way customers interact with software, a vendor’s value may depend less on its screens and more on the business rules, records, audit trails and controls that its product maintains. That is an implication of the development, not a result measured by the test.

Amazon

contract management AI software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Ironclad Test Was Built

OpenAI’s October 6 post was titled “Advancing computer use with Ironclad.” The source material says other AI news trackers misread Ironclad as the name of a hardened agent framework; it is instead a contract-management software company. The reported work focused on training and evaluating a model within that company’s product, rather than announcing a new standalone agent framework.

The test used a small set of 11 selected tasks and criteria tailored to each task. OpenAI estimated the time an experienced user might take, then compared that figure with a simulated agent-time estimate. It also invited a small number of other software companies to bring examples of work current agents cannot reliably complete, along with subject-matter experts, secure testing environments and data that can safely be used for research.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Results Do Not Establish

The reported averages do not identify, in the provided source material, which criteria Astra missed across all 11 tasks or how often a missed requirement would create a serious operational or legal problem. The score is not a pass rate, and the showcase result of about 94% on one task cannot establish performance across other work.

OpenAI’s comparison of 37.0 minutes for GPT-5.6 Sol and 19.2 minutes for GPT-6 Astra is based on simulated time estimates, not observed customer use. The source does not provide independently verified customer outcomes, a deployment study or evidence that the model can operate Ironclad workflows safely without review. It also does not detail the full security arrangements for partner environments or the terms under which future software partners would share research data.

Amazon

AI-powered NDA creation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Next Test for Software Agents

OpenAI said it is seeking a small number of software-company partners to identify challenging workflows for future research. Prospective partners are expected to provide concrete examples where current agents fail, people with deep knowledge of the work, secure test environments and data suitable for research. The source material does not give a schedule for new partnerships or further benchmark results.

For companies considering agents in contract, finance or customer-record systems, the reported test points to practical questions: which exact requirements did the agent miss, who reviews each result, and what controls prevent an incomplete workflow from being treated as finished? Until performance is demonstrated on real deployments with clearly reported error rates and safeguards, the Ironclad results are best read as a research-stage evaluation, not proof of ready-to-run automation.

Amazon

procurement approval workflow software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Ironclad in this report?

Ironclad is a contract-management software company, not the name of a new agent framework. OpenAI’s post describes work training and evaluating a model in hosted copies of Ironclad’s product.

What does Astra’s 55% score mean?

It is the average share of rubric criteria met across the selected tasks. It is not the share of tasks completed successfully, and it does not mean the agent delivered 55% of the value of a contract workflow.

Did OpenAI demonstrate customer time savings?

No measured customer savings are established in the source. OpenAI’s reported times—19.2 minutes for Astra and 37.0 minutes for GPT-5.6 Sol—were simulated estimates based on assumed processing and generation speeds.

Can companies use the agent to manage contracts without review?

The results do not establish that. OpenAI’s post says human oversight matters because agents can lose track of business rules, and the reported average score indicates that many listed criteria were not met.

What data did OpenAI say it used?

OpenAI said it built synthetic tasks from publicly filed contracts in the SEC’s EDGAR database and filtered out personal information. It said it did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

DojoClaw: The Engine Behind the Fleet

DojoClaw, an AI-driven content factory, now supports over 450 sites, leveraging owner hardware and provider-agnostic models to scale high-volume publishing efficiently.

The labor share. Is value really moving from labor to capital? The data isn’t on anyone’s side yet.

Examining whether AI is truly shifting value from labor to capital, with current data showing both stability and early signals of change.

Organize Estate Administration Tasks With Trust Software

An untested software concept would help law firms and financial advisors track trust funding, collect proof and flag incomplete transfers.

The conversion. What turning the largest nonprofit into a company did to charity law.

OpenAI transformed from a nonprofit into a company retaining control, diverging from standard divestiture practices. This raises legal and ethical questions.