OpenAI Decisions API
Use OpenAI Decisions API for ticket routing, labels, and action gates. Three guide requests, refusal handling, pricing, limits, and when to keep your LLM.
Published

OpenAI's Decisions API lets you route tickets, label records, and assess proposed agent actions with answers your code can use directly. If your current LLM call exists only to return a category or rating, this is worth a trial. Replace that call only after it matches your routing quality and improves the workflow enough to justify the migration.
Start With the Decision Your Code Needs
Think of Decisions as a dispatcher with a defined set of destinations. You supply the evidence and the question; your application decides what happens after the answer arrives.
As of October 11, 2026, the API is in public beta, with general availability expected “in the coming weeks.” That is OpenAI's expectation, not a release date. gpt-6-luna is the only supported model, and the endpoint is POST /v1/decisions. OpenAI describes it as “about 10x faster than the Responses API.” Treat that as OpenAI's claim, not a measured result for your application. OpenAI's Decisions guide
Choose the answer shape before writing the prompt:
For a score, the levels start at index 0. The result can fall between levels: it summarizes uncertainty across them. Use a choice when your code needs one category. Question types

Make Each of the Three Requests
These are the guide's three cURL request samples, copied with formatting normalized and comments added to separate them. Set OPENAI_API_KEY in your shell environment; the predicate example also needs a local product.png. Each command makes its own request. These examples were checked against the public guide, not run against a signed-in account. Original request samples
The common pieces are input, the evidence, and questions, the decisions to make about it. A question's name identifies its answer in the returned answers array. Request and response reference
# Predicate: inspect product.png for visible damage
IMAGE_BASE64="$(base64 < product.png | tr -d '\r\n')"
curl https://api.openai.com/v1/decisions \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
--data-binary @- <<JSON
{
"model": "gpt-6-luna",
"input": [{
"role": "user",
"content": [
{"type": "input_text", "text": "Inspect the product in this photo."},
{"type": "input_image", "image_url": "data:image/png;base64,$IMAGE_BASE64"}
]
}],
"questions": [{
"type": "predicate",
"name": "visible_damage",
"instructions": "Does the product have visible damage, such as a crack, tear, or dent? Ignore shadows and damage to the packaging."
}]
}
JSON
# Choice: route a customer complaint
curl https://api.openai.com/v1/decisions \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6-luna",
"input": "I was charged twice for my order.",
"questions": [{
"type": "choice",
"name": "department",
"instructions": "Which department should handle this complaint?",
"choices": [
{"value": "billing", "description": "Payments, invoices, and refunds."},
{"value": "technical", "description": "Problems using the product."},
{"value": "shipping", "description": "Delivery and tracking."},
{"value": "other", "description": "Requests outside these categories."}
]
}]
}'
# Score: assess issue severity
curl https://api.openai.com/v1/decisions \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6-luna",
"input": "Export fails in Safari but works in Chrome.",
"questions": [{
"type": "score",
"name": "severity",
"instructions": "How severe is this issue?",
"levels": [
{"label": "Cosmetic", "description": "Appearance only; no lost functionality."},
{"label": "Workaround available", "description": "A task fails, but another way works."},
{"label": "Fully blocked", "description": "A task fails with no workaround."}
]
}]
}'The choice sample is already a useful business starting point: a duplicate-charge complaint must go to one of the fixed queues. The score sample asks a different question: how disruptive is a failure when another browser still works? Keep queue ownership and severity separate so a billing issue can be urgent without becoming a technical ticket.
Handle refusal before reading the value. The guide's SDK samples check answer.type == "refusal" before accessing probability, choice, or score. A refusal is a separate result, not a low-confidence answer and not your other category. Refusal behavior
Turn the Choice Into Support-Ticket Triage
Build the first version around queue assignment. A support team could send the ticket subject and relevant customer message, get a department choice, then let ordinary application code apply its routing policy.
Use the guide's queue definitions as a starting point and replace them with your team's actual ownership boundaries. Preserve an other option for requests the named departments do not cover. A refund request belongs to billing; that does not authorize a refund.
This small application adapter illustrates the policy after you parse a successful JSON response and select the department answer. thresholds must contain cutoffs you established on labelled tickets; a missing cutoff leaves the ticket for review.
QUEUES = {
"billing": "billing",
"technical": "technical",
"shipping": "shipping",
}
def queue_for(answer, thresholds):
if answer.get("type") == "refusal":
return "manual_review"
if answer.get("type") != "choice":
return "manual_review"
department = answer.get("choice")
if department not in QUEUES: # Includes the guide's "other" choice.
return "manual_review"
cutoff = thresholds.get(department)
confidence = answer.get("confidence")
if cutoff is None or confidence is None or confidence < cutoff:
return "manual_review"
return QUEUES[department]For this proposed workflow, API errors, timeouts, and missing answers also leave tickets in manual review. Record the question version, proposed queue, confidence, final queue, and any staff correction. Make the queue update safe to retry so a retried request cannot create duplicate assignments.
You can put independent questions about the same ticket in one request. A follow-up question whose meaning depends on an earlier answer belongs in a later request. Multiple-question guidance

Set Thresholds on Labelled Work
Choose cutoffs by measuring the mistakes your business can tolerate. OpenAI publishes no calibration numbers in the guide and directs builders to labelled application data. A confidence value is not a promise that the answer will be correct that often on your tickets. Interpreting answers
Start with past tickets whose correct destination a support lead has checked. Include short requests, mixed issues, missing context, and complaints containing instructions aimed at the model. Separate the examples used to tune questions and cutoffs from a held-out set you reserve for the final comparison.
Measure wrong-queue assignments, the share sent to review, and the time staff spend correcting assignments. Check each queue separately. Mistaking shipping for billing may have a different operational cost from missing a report of account compromise.
Run the candidate alongside the existing classifier without changing live assignments first. Promote it only when the results meet a written acceptance rule. Keep the old route available for rollback. These are proposed deployment steps, not results from a test of this API.
Six Jobs Worth Trying, Ranked
The best early uses have stable categories, observable mistakes, and a person who already owns exceptions. This ranking is an implementation judgment, not an accuracy leaderboard.
For labelling, decide whether records can have multiple themes. A single choice selects one category; separate questions may suit overlapping labels. For incident severity, define impact in operational terms such as lost functionality and available workarounds. Words like “serious” leave too much interpretation to the model.
What the Price Changes in Practice
The guide's price is input $0.10 per 1M tokens on gpt-6-luna, with no output, cache-read, or cache-write charges. Regional processing premiums and long-context input multipliers may apply. These are Decisions rates; do not apply this billing rule to ordinary calls using the same model. Decisions pricing
Here is a calculated batch example, not a usage measurement or flat per-decision price. Assume 100,000 classification calls, each with 1,000 uncached input tokens including instructions and options. For the existing Responses call, also assume 50 total billed output tokens per call. Use standard short-context base pricing, with no cache writes, regional premiums, retries, or other charges.
The ordinary model rates come from OpenAI's standard pricing table. In this example, removing output billing saves $2.50 across the whole batch. That alone is a weak reason to rewrite a working integration.
The stronger case is operational: less waiting in a serial workflow, less response-handling code, or less manual sorting at the same error rate. Compare against your actual invoice and review workload. A more expensive incumbent model or longer generated answers change the math; so do existing cache discounts. Engineering time and misrouted tickets still belong in the migration budget.
Two Things Worth Building
First Pick: A Support Router With Review and Override History
A support operations lead could pay for a connector that suggests fixed queues, holds uncertain cases, and turns staff corrections into evaluation data. The useful product is the complete routing workflow, including its maintenance.
DataForSEO estimates 170 US Google searches per month for “ticket triage,” checked October 11, 2026. That is a narrow, informational demand signal, not a buyer count. Existing helpdesk products also address this job: Zendesk offers intelligent triage classifications and requires its Copilot add-on to use them in workflows. Zendesk's triage guide
The smallest useful version could support one helpdesk, import historical tickets, preview queue assignments, and offer a review inbox with overrides. Its strongest sales evidence would be the customer's reduction in avoidable handoffs. The catch is incumbent coverage: if the helpdesk already handles the queues well, another router adds maintenance. Win on a specific ownership problem or cross-system handoff.
Second Pick: A Review Workbench for Fixed Labels
A research or data team could pay for a workbench that suggests labels, collects corrections, and shows which categories keep causing disagreement. DataForSEO estimates 90 US Google searches per month for “automated data labeling,” on the same check. This indicates interest in the job, not proven willingness to buy this implementation.
A first version could take a CSV, apply a versioned label set, present uncertain rows for review, and export corrected results. Keep a held-out evaluation set so label changes can be compared fairly. The catch is that a model can cheaply reproduce an unclear taxonomy. The product needs good review and category-management tools; wrapping an API call is easy to copy.
The support router is the stronger first build. Queue ownership gives you a visible mistake, an operator who can correct it, and a recurring workflow against which to demonstrate value. Interview that operator before building a general decision platform.
When to Keep Your Existing Approach
Keep the Responses API when you need extracted fields in your own JSON schema, a written explanation, or a model-requested tool call with arguments. Decisions serves the narrower answer types above. OpenAI's guidance on choosing the interface
Keep deterministic rules when the answer is already in an account field or explicit policy. A model adds little to “route customers in this region to this team.” Keep human authorization and application permission checks for consequential actions. A routing judgment can inform a refund workflow; it cannot establish the customer's entitlement or the operator's authority.
Check these integration constraints before scheduling a migration:
- Images: the guide permits only inline base64 data URLs, meaning the image bytes are encoded inside the request. Hosted HTTP or HTTPS image URLs and
file_id, a handle to a previously uploaded file, are not supported under the guide's instructions. Image input requirements - Data controls: the guide states support for Zero Data Retention, or ZDR, and HIPAA use for eligible customers. Residency and regional processing are supported in the US and Europe, specifically EEA + Switzerland. Eligibility, agreements, configuration, and limitations apply; these are not automatic account defaults. Decisions availability, OpenAI data controls
- Release maturity: public beta is a reason to preserve a rollback path. If your existing classifier meets its targets and changing it has no measurable benefit, leave it in place.
Alternatives: Jev, Clef, and Microsoft-Decision-1
Compare the same labelled workload across candidates before switching providers. TypeSafe's Jev accepts state and typed questions through its System One API. Cloudflare Clef offers typed decisions in Workers AI. Microsoft-Decision-1 is available in Microsoft Foundry for jobs including classification, routing, and prioritization. TypeSafe quick start, Cloudflare Clef documentation, Microsoft's announcement
Existing integrations, hosting requirements, evaluation results, and exception handling should drive the shortlist. Our Jev support-ticket routing guide covers the routing pattern; Is Jev Router Free? explains the distinction between router software and hosted inference. Treat each provider's request format and confidence behavior separately.
Should Decisions replace my existing classification call?
Trial it when the output is a fixed category, condition estimate, or rubric score. Compare routing errors, manual-review workload, cost, and timing on the same labelled examples. Keep the current call if the improvement does not justify migration.
What should my application do when a decision is refused?
Check the answer type before reading its value. In a support router, send refusal to manual review and retain enough context for staff to handle it. Do not treat refusal as permission to choose a default action.
Which confidence threshold should I use?
Set it on labelled data from your own workflow. Measure mistakes and review volume for each candidate cutoff, preferably by queue. There is no universal cutoff supplied by this article or calibration table in OpenAI's guide.
Can I pass a hosted image URL or an uploaded file ID?
Follow the guide's inline base64 data URL format. It explicitly excludes hosted image URLs and file_id inputs for this endpoint. The predicate request above shows the supported guide pattern.
Your Monday Move
Pick the classification call that sends tickets to an established set of queues. Have its owner check a representative labelled sample, write down the acceptable error and review rates, and compare Decisions alongside the current call without changing assignments. Switch only after it earns that place in the workflow.
If you want a routing workflow built around your existing tools, we build production AI systems.
- Published
- Category
- Build
- Language







