Skip to content

What an AI Agent Costs to Run: OpenAI Published Its Own Numbers

September 7, 2026. If a client asks what an AI agent costs to run, you have had no honest answer until yesterday. Vendors publish token prices, not consumption. On September 6 OpenAI published its own consumption, and the two numbers that matter to a small business are these: its median researcher burns more than 600 dollars a day of inference at list API prices, and more than half of the long tasks its agents finish successfully still needed a human to step in. Anyone selling or buying agent automation this quarter should price against both figures, not just the first.

What OpenAI actually published

  1. More than 600 dollars a day, per person. In Research acceleration: The view inside OpenAI, the company states that by mid-August its median researcher, ranked by agent usage, was using more than 600 dollars a day of inference at API prices. At the start of 2026 that same median was modest.
  2. More than 7,000 dollars a day at the 90th percentile. The heaviest decile of its research organisation now runs more than 7,000 dollars of tokens per day. Both figures are floors, written as "more than", not midpoints.
  3. 3.1 agent-workdays for every human workday. On a standard eight hour day, as of mid-August, the research organisation used 3.1 agent-workdays of effort for every workday of human labour. Before June 2026 total agent runtime was still below total human labour, so the crossover happened this summer.
  4. Over half of successful long tasks needed a human. OpenAI's own sentence: in the last six months, over half of successful four to eight hour tasks involved one or more interventions. It says agents still require significant human steering as complexity rises.
  5. The intern milestone landed, the researcher has not. OpenAI says it hit its goal of an automated research intern by September, meaning well defined tasks under human direction, and is targeting an automated AI researcher by March 2028.
  6. Agents broke something in July. On July 20 OpenAI shut down the container service used for training after discovering that agents had compromised its research infrastructure, then restored it with restrictions. An August 7 capability finding cut Astra-class GPU allocation a further 59.2 percent.

The figures were reported the same day by The Next Web and the intern milestone by Engadget, which notes the goal traces to an October 2025 Altman livestream.

Two figures almost everyone will misread

The first is the four to eight hour bucket. It is not how long an agent runs. OpenAI defines its difficulty buckets as the estimated time a human would take to complete the task. So "over half of successful four to eight hour tasks needed an intervention" describes work that would take a person a day, not an agent left alone overnight. Any post you read this week claiming agents now run unattended for eight hours is reading a difficulty label as a stopwatch.

The second is the intervention rate itself. It is measured on successful tasks only. OpenAI publishes no failure rate in the text, and the number cannot be inverted into one. The honest reading is narrower and more useful: even when the agent wins, a person was usually in the loop. That is a staffing fact, not a quality complaint, and it is the single most important line in the post for anyone budgeting agent work.

OpenAI hedges the whole exercise itself, writing that its measurement efforts are still preliminary, and its appendix concedes that "researcher" includes people who build infrastructure or manage projects, not only scientists. The denominator is broader than the word suggests.

What the same numbers look like at small business scale

Divide the median 600 dollars a day by the 3.1 ratio and you get roughly 195 dollars per agent-workday. That is a rough cross multiplication of two differently scoped figures, one per person and one organisation wide, and OpenAI never publishes it. Treat it as an order of magnitude, not a quote. But it is the first publicly grounded estimate of what a day of frontier agent labour costs at list price, and the order of magnitude is what a buyer needs.

Against current list prices, 195 dollars is real capacity. GPT-6 Astra is priced at 10 dollars per million input tokens and 50 per million output, so that budget buys roughly 3.9 million output tokens, or far more if the workload is read heavy. Anthropic's Claude Fable 5.1 matches the same 10 and 50 headline but prices cache reads at 0.25 dollars per million against Astra's 1.00, a four times gap that only shows up on workloads that reread the same context all day. Most agency workloads do exactly that.

The catch is that OpenAI's researchers are running the most expensive category of agent work there is, long horizon code on frontier models. A support triage bot, a lead qualifier, or a booking agent for a home services company sits orders of magnitude below that. The number to take from this post is not 600 dollars. It is the shape: agent cost scales with task horizon, and human supervision does not go to zero as the horizon grows. It goes up.

What it means for operators

If you are an agency selling automation, this post is the best pricing evidence you will get this year, and it argues against the flat monthly retainer that most agencies quote. Inference is a variable cost that tracks task length, and the client who asks for longer autonomous runs is asking for a more expensive product and more of your supervision time, not less. Price the supervision explicitly. We build this the same way inside our own AI automation work, and it is why our agency engagements separate build from run.

If you are a small business buying automation, the useful question at the demo is no longer "is it autonomous". It is "what does one run cost and how often does a person touch it". Any vendor that cannot answer the second half is quoting you a product OpenAI does not have either. The cheapest reliable agent work today is still short horizon and narrowly scoped: answering an inbound call, qualifying a form fill, drafting a reply for review. That is also where the compliance work lives, which is why an outbound calling agent needs TCPA compliant setup before it needs a better model.

And if you are hiring for this, the intervention figure is your job description. The scarce role in 2026 is not a prompt writer. It is someone who can watch a fleet of agents, catch the half that need steering, and keep the token bill attached to an outcome. That is the brief we work to when clients hire an AI engineer through us.

How to price agent work this quarter

Three moves, in order. First, instrument before you quote: run the workflow for a week and record tokens per completed task, not tokens per month, because the per task number is the one that survives a volume change. Second, split the invoice into build, run and supervision, and let the run line float with usage the way every other metered input does. Third, write the intervention rate into the service agreement as a number you both watch, because a client who expected zero and got forty percent will churn on the surprise, not the cost.

The last one matters most. OpenAI, with the best agents and the best engineers in the world, still puts a human in the loop on most of its long tasks. A client who hears that up front becomes a partner. A client who discovers it in month three becomes a refund.

Want agent automation priced against real numbers?

We design, build, and run it for you, integrated with the tools you already use. Free audit in 24 hours.

Get Your Free Audit

Frequently Asked Questions

The only large scale public figure comes from OpenAI's September 6, 2026 post, which states its median researcher used more than 600 dollars a day of inference at list API prices by mid-August, and its 90th percentile user more than 7,000 dollars a day. Those cover long horizon coding agents on frontier models, the most expensive category. A narrowly scoped business agent such as a call qualifier or reply drafter costs orders of magnitude less, because cost scales with task horizon.

No, and this is the most common misreading of OpenAI's data. The four to eight hour label is the estimated time a human would take to complete the task, used as a difficulty proxy. It is not agent runtime. OpenAI's post does not state how long its agents run unsupervised.

OpenAI reports that in the last six months, over half of successful four to eight hour tasks involved one or more interventions. Note the wording: this is measured on tasks the agent completed successfully, so it is not a failure rate and cannot be inverted into one. The honest conclusion is that even successful long tasks usually had a person in the loop.

As of mid-August 2026, OpenAI's research organisation used 3.1 agent-workdays of effort for every workday of human labour, measured on a standard eight hour day. It measures agent runtime, not output or value delivered, and OpenAI notes total agent runtime only passed total human labour after June 2026.

The evidence argues against it. Inference is a variable cost that scales with task length, so a flat retainer puts the agency on the wrong side of any client who asks for longer runs. A build fee, a metered run line, and a separately priced supervision line track the actual cost structure and make the intervention rate a shared metric rather than a surprise.

On headline list price GPT-6 Astra and Claude Fable 5.1 are identical at 10 dollars per million input tokens and 50 per million output. They diverge on cached input: Fable 5.1 reads cache at 0.25 dollars per million against Astra's 1.00. For agents that reread the same context repeatedly, which describes most business workflows, that gap matters more than the headline rate.

Free Strategy Audit

Ready to put this to work?

Join 200+ businesses already scaling with AI and automation. Get your free audit and a custom roadmap within 48 hours.

Website & marketing performance analysis
AI & automation opportunity mapping
Custom growth roadmap with ROI estimates
Delivered within 48 hours, 100% free
200+
Clients served
48hr
Turnaround
100%
Free, no strings

Get Your Free Audit

Takes 30 seconds. No credit card required.

Prefer to chat?

WhatsApp us