Skip to content

OpenAI Decisions API Pricing: $0.10 per Million, No Output Fee

October 9, 2026. OpenAI's Decisions API costs $0.10 per million input tokens on gpt-6-luna, and input is the only thing it bills: there are no output-token, cache-read or cache-write charges, according to OpenAI's Decisions guide. That puts 10,000 lead-routing decisions of about 1,500 tokens each at $1.50. The endpoint entered public beta on October 6, per the API changelog, and OpenAI says it expects general availability in the coming weeks.

OpenAI Decisions API pricing: $0.10 per million input tokens on gpt-6-luna with no output or cache charges

Key numbers

ItemNumber
Decisions API input price (gpt-6-luna)$0.10 per 1M tokens
Output, cache read and cache write charges$0
Regional processing endpoints (10% uplift)$0.11 per 1M tokens
Requests over 272,000 input tokens (for the whole request)$0.20 per 1M tokens
Speed vs the Responses API (per OpenAI)about 10x faster
Models availablegpt-6-luna only
Question types (predicate, choice, score)3
10,000 routing decisions, 1,500 tokens each$1.50
Same job on gpt-6-luna through the Responses API (50-token answer)$1.75
Same job on GPT-6.1 Solabout $35
Public beta since (general availability expected in the coming weeks)October 6, 2026

Prices and terms read on 9 October 2026 from OpenAI's Decisions guide, API changelog, pricing page and gpt-6-luna model page; per-task costs at list prices.

What the Decisions API does

It reads text, images or both and returns a typed answer instead of prose, about 10x faster than the Responses API, OpenAI says. InfoQ's DevDay recap described it on October 2 as a limited preview for classification, request routing and choosing an agent's next action; OpenAI's guide now marks it public beta and prints the price.

  1. Predicate: the probability, from 0 to 1, that a condition is true, such as visible damage in a product photo.
  2. Choice: one option from a fixed set you supply, such as a department or a content category.
  3. Score: a rating against ordered levels, such as issue severity, returned as a probability-weighted average.
  4. Model and endpoint: gpt-6-luna only, on a dedicated POST /v1/decisions endpoint, with Zero Data Retention and HIPAA use for eligible customers and data residency in the US and Europe.

What it costs next to the alternatives

Take a routing job: 10,000 inbound leads or tickets, about 1,500 input tokens each, each sorted into one of a few queues. On the Decisions API that is 15 million input tokens at $0.10, or $1.50. The same job through the Responses API on gpt-6-luna, with a 50-token answer at $0.50 per million output tokens, comes to $1.75 at OpenAI's list prices. Claude Haiku 5.5 also lands at about $1.75 for that job, by the per-task math in our Haiku 5.5 pricing breakdown, while GPT-6.1 Sol at $2 input and $10 output per million would run about $35.

Two multipliers apply. Regional processing endpoints add 10%, so $0.11 per million tokens, or $1.65 for the same 10,000 decisions, and a request with more than 272,000 input tokens is billed at 2x input, $0.20 per million, for the whole request, per the gpt-6-luna model page. Against the Responses API the saving is small in dollars; the bigger gains are speed and an answer your code can branch on without parsing text.

Bar chart of the cost of 10,000 routing decisions at list prices: Decisions API 1.50 dollars, Decisions on a regional endpoint 1.65, GPT-6 Luna through the Responses API 1.75, Claude Haiku 5.5 1.75, and GPT-6.1 Sol 35.
Per-task math at list prices; Haiku 5.5 at Anthropic list prices. Source: developers.openai.com, October 2026

What it means for operators

  1. Lead routing. A choice question sends each form fill or reply to the right rep, sequence or nurture track for a fraction of a cent.
  2. Intake triage. Law firms can sort enquiries by practice area and urgency before a human reads them, the same pattern behind our AI intake for law firms.
  3. Ticket and review scoring. A score question ranks severity so the worst problems surface first.
  4. Voice agents. OpenAI's guide pairs the API with the Live API through client delegation, so a voice agent can pick an action from a spoken request.
  5. Plan for the beta. Only gpt-6-luna is available and terms can change before general availability, so keep a fallback route in production.

If you want routing like this wired into your CRM or help desk, our AI engineers build it and our AI automation team runs it.

Want AI routing that costs a fraction of a cent?

We design, build, and run it for you, integrated with the tools you already use. Free audit in 24 hours.

Get Your Free Audit

Frequently Asked Questions

$0.10 per million input tokens on gpt-6-luna, with no charges for output tokens, cache reads or cache writes. Regional processing adds 10%, or $0.11 per million, and requests over 272,000 input tokens are billed at $0.20 per million for the whole request.

Not yet. It entered public beta on October 6, 2026, after a limited preview announced at DevDay, and OpenAI says it expects general availability in the coming weeks. gpt-6-luna is the only model available.

Three question types: a predicate returns the probability that a condition is true, a choice returns one option from a fixed list you supply, and a score rates an input against ordered levels. You can ask several named questions in one request.

Slightly, for short answers. Routing 10,000 items of about 1,500 tokens costs $1.50 on the Decisions API versus about $1.75 on gpt-6-luna through the Responses API with a 50-token answer, because the Decisions API charges nothing for output. OpenAI says it is also about 10x faster.

Free Strategy Audit

Ready to put this to work?

Join 200+ businesses already scaling with AI and automation. Get your free audit and a custom roadmap within 48 hours.

Website & marketing performance analysis
AI & automation opportunity mapping
Custom growth roadmap with ROI estimates
Delivered within 48 hours, 100% free
200+
Clients served
48hr
Turnaround
100%
Free, no strings

Get Your Free Audit

Takes 30 seconds. No credit card required.

Prefer to chat?

WhatsApp us