October 9, 2026. OpenAI's Decisions API costs $0.10 per million input tokens on gpt-6-luna, and input is the only thing it bills: there are no output-token, cache-read or cache-write charges, according to OpenAI's Decisions guide. That puts 10,000 lead-routing decisions of about 1,500 tokens each at $1.50. The endpoint entered public beta on October 6, per the API changelog, and OpenAI says it expects general availability in the coming weeks.

Key numbers
| Item | Number |
|---|---|
| Decisions API input price (gpt-6-luna) | $0.10 per 1M tokens |
| Output, cache read and cache write charges | $0 |
| Regional processing endpoints (10% uplift) | $0.11 per 1M tokens |
| Requests over 272,000 input tokens (for the whole request) | $0.20 per 1M tokens |
| Speed vs the Responses API (per OpenAI) | about 10x faster |
| Models available | gpt-6-luna only |
| Question types (predicate, choice, score) | 3 |
| 10,000 routing decisions, 1,500 tokens each | $1.50 |
| Same job on gpt-6-luna through the Responses API (50-token answer) | $1.75 |
| Same job on GPT-6.1 Sol | about $35 |
| Public beta since (general availability expected in the coming weeks) | October 6, 2026 |
Prices and terms read on 9 October 2026 from OpenAI's Decisions guide, API changelog, pricing page and gpt-6-luna model page; per-task costs at list prices.
What the Decisions API does
It reads text, images or both and returns a typed answer instead of prose, about 10x faster than the Responses API, OpenAI says. InfoQ's DevDay recap described it on October 2 as a limited preview for classification, request routing and choosing an agent's next action; OpenAI's guide now marks it public beta and prints the price.
- Predicate: the probability, from 0 to 1, that a condition is true, such as visible damage in a product photo.
- Choice: one option from a fixed set you supply, such as a department or a content category.
- Score: a rating against ordered levels, such as issue severity, returned as a probability-weighted average.
- Model and endpoint: gpt-6-luna only, on a dedicated POST /v1/decisions endpoint, with Zero Data Retention and HIPAA use for eligible customers and data residency in the US and Europe.
What it costs next to the alternatives
Take a routing job: 10,000 inbound leads or tickets, about 1,500 input tokens each, each sorted into one of a few queues. On the Decisions API that is 15 million input tokens at $0.10, or $1.50. The same job through the Responses API on gpt-6-luna, with a 50-token answer at $0.50 per million output tokens, comes to $1.75 at OpenAI's list prices. Claude Haiku 5.5 also lands at about $1.75 for that job, by the per-task math in our Haiku 5.5 pricing breakdown, while GPT-6.1 Sol at $2 input and $10 output per million would run about $35.
Two multipliers apply. Regional processing endpoints add 10%, so $0.11 per million tokens, or $1.65 for the same 10,000 decisions, and a request with more than 272,000 input tokens is billed at 2x input, $0.20 per million, for the whole request, per the gpt-6-luna model page. Against the Responses API the saving is small in dollars; the bigger gains are speed and an answer your code can branch on without parsing text.

What it means for operators
- Lead routing. A choice question sends each form fill or reply to the right rep, sequence or nurture track for a fraction of a cent.
- Intake triage. Law firms can sort enquiries by practice area and urgency before a human reads them, the same pattern behind our AI intake for law firms.
- Ticket and review scoring. A score question ranks severity so the worst problems surface first.
- Voice agents. OpenAI's guide pairs the API with the Live API through client delegation, so a voice agent can pick an action from a spoken request.
- Plan for the beta. Only gpt-6-luna is available and terms can change before general availability, so keep a fallback route in production.
If you want routing like this wired into your CRM or help desk, our AI engineers build it and our AI automation team runs it.
Frequently Asked Questions
$0.10 per million input tokens on gpt-6-luna, with no charges for output tokens, cache reads or cache writes. Regional processing adds 10%, or $0.11 per million, and requests over 272,000 input tokens are billed at $0.20 per million for the whole request.
Not yet. It entered public beta on October 6, 2026, after a limited preview announced at DevDay, and OpenAI says it expects general availability in the coming weeks. gpt-6-luna is the only model available.
Three question types: a predicate returns the probability that a condition is true, a choice returns one option from a fixed list you supply, and a score rates an input against ordered levels. You can ask several named questions in one request.
Slightly, for short answers. Routing 10,000 items of about 1,500 tokens costs $1.50 on the Decisions API versus about $1.75 on gpt-6-luna through the Responses API with a 50-token answer, because the Decisions API charges nothing for output. OpenAI says it is also about 10x faster.