OpenAI Decisions API Is Public: When to Use It Instead of Responses

OpenAI’s Decisions API is now in public beta with fast predicate, choice and score outputs. Here is how it differs from Responses, what it costs, and how to test it safely.
OpenAI presents the Decisions API onstage at DevDay 2026
OpenAI introduced the Decisions API at DevDay 2026 before opening the public beta on October 6. Image: OpenAI DevDay presentation, via ITmedia.

OpenAI opened its Decisions API to all developers in public beta late Tuesday, turning a DevDay preview into a production-facing endpoint for applications that need a fast answer from a constrained set of possibilities.

The API accepts text, images or both through POST /v1/decisions. Instead of generating a free-form response, it can estimate whether a condition is true, select from choices supplied by the developer, or score an input against an ordered rubric. OpenAI says the endpoint returns those answers about 10 times faster than running GPT-6 Luna through the Responses API. The company expects the service to reach general availability in the coming weeks, according to its public-beta documentation.

This is a deliberately narrow API. It is designed for jobs such as routing a support ticket, deciding whether a product photo shows damage, assigning an issue-severity score or choosing which model or tool should handle a request. It does not replace the Responses API when an application needs prose, an arbitrary JSON object, tool arguments or a multi-step answer.

What the Decisions API returns

Every request contains a model, shared input and one or more questions. GPT-6 Luna is the only supported model in the beta. Each question has a type and instructions; choice and score questions also define the permitted options.

Question type Use Returned value
predicate Test whether a condition is true A probability from 0 to 1
choice Select one item from a fixed, unordered list The selected value, probabilities for every option and a confidence value
score Rate input against ordered levels A probability-weighted numeric score, per-level probabilities and confidence

A support system, for example, could pass a complaint as the input and offer billing, technical, shipping and other as choices. The result is machine-ready: the selected value plus a distribution across the available options. That distribution lets an application send uncertain cases to a general queue instead of treating every classification as equally reliable.

Score questions behave differently. Developers define ordered levels such as cosmetic, workaround available and fully blocked. The API returns a weighted average of the level indices, so the result can fall between two categories. That is useful for prioritization, but applications should preserve the underlying probability distribution rather than treating a fractional score as a precisely measured severity.

A minimal request

{
  "model": "gpt-6-luna",
  "input": "The customer says checkout fails for every user.",
  "questions": [
    {
      "type": "choice",
      "name": "route",
      "instructions": "Which team should receive this report?",
      "choices": [
        {"value": "payments", "description": "Payment processing failures"},
        {"value": "storefront", "description": "Shopping and checkout interface"},
        {"value": "account", "description": "Sign-in and account access"},
        {"value": "other", "description": "None of the listed teams"}
      ]
    },
    {
      "type": "score",
      "name": "severity",
      "instructions": "How severe is the reported service impact?",
      "levels": [
        {"label": "Low", "description": "No lost functionality"},
        {"label": "Medium", "description": "A workaround is available"},
        {"label": "High", "description": "A core task is blocked with no workaround"}
      ]
    }
  ]
}

Independent questions can share one input in a single call. Questions that depend on an earlier answer require separate requests, so developers still need ordinary application logic for branching workflows.

Where it differs from Responses and Structured Outputs

The dividing line is the shape of the answer. Decisions is appropriate when the acceptable result is known before the request: yes or no, one of several routes, or a position on an ordered scale. OpenAI directs developers to Structured Outputs through the Responses API when they need an object that follows a custom JSON schema, and to function calling when the model must propose a tool and its arguments.

That distinction matters for agent systems. A Decisions call can choose among a set of approved actions, but the endpoint does not execute those actions, call tools or reason through a workflow on the application’s behalf. It can act as a fast routing layer in front of a larger model or conventional service. The application remains responsible for validating the returned value, enforcing permissions and carrying out the next step.

The beta also has input constraints. It accepts text and inline images in user messages, with as many as 128 images across a request. Images must be supplied as base64 data URLs. External image URLs, file IDs, audio, function calls, tool results and other message roles are not supported, according to the API reference. A question may return a refusal while other questions in the same request still receive answers, so callers need to handle mixed result types.

The price is built for high-volume routing

Decisions requests using GPT-6 Luna cost $0.10 per million input tokens. OpenAI charges no cache-read, cache-write or output-token fees on this endpoint. Regional-processing premiums and the long-context input multiplier still apply.

That pricing is materially different from using Luna through a general generation endpoint, where standard short-context rates are $0.10 per million input tokens and $0.50 per million output tokens. Decisions avoids output-token billing because its responses are fixed, typed values rather than generated text. The comparison is not purely financial, however: teams give up open-ended outputs, built-in tools and arbitrary schemas in exchange for the narrower interface.

OpenAI says eligible customers can use the endpoint with Zero Data Retention and HIPAA support. Data residency and regional processing are available in the United States and Europe, including the European Economic Area and Switzerland. Eligibility and contractual requirements still apply.

How to test it before automating real work

Confidence is not a universal approval threshold. A value that is acceptable for sorting low-risk feedback may be unacceptable for fraud review, medical intake or an action that changes a customer’s account. OpenAI recommends using labeled examples from the target application and choosing thresholds based on the relative cost of false positives and false negatives.

A sensible rollout should include five checks:

  • Build a representative test set. Include routine cases, ambiguous language, misspellings, multilingual inputs and examples that do not fit the available choices.
  • Add an explicit fallback. An other or manual-review option prevents the model from being forced into a bad category when the taxonomy is incomplete.
  • Measure each class separately. Overall accuracy can hide a route that fails frequently or a rare high-impact class the system misses.
  • Set action-specific thresholds. Low-confidence results can go to a human queue or a slower model instead of triggering an irreversible action.
  • Re-run evaluations after changes. Editing instructions, choices, level descriptions or upstream input formatting can move the probability distribution even when application code stays the same.

Teams should also log the model name, question definitions, returned probabilities, final action and any human correction. That creates the evidence needed to recalibrate thresholds and detect drift. Sensitive logs need the same access controls and retention rules as the source data.

The useful part is the constraint

General-purpose models are often used for classification because they are already available, not because free-form generation is the best interface for the job. Decisions gives developers a dedicated endpoint for the many moments when software needs a bounded judgment rather than a conversation.

The public beta makes that pattern inexpensive and easier to integrate, but it does not make the model’s probabilities ground truth. The teams most likely to benefit are those that can define clear options, measure mistakes and preserve a deterministic fallback. Where the categories are subjective, incomplete or tied to high-stakes actions, human review and conventional policy controls remain part of the system.

Previous Post
Network cables connected to server hardware in a data center rack, illustrating DNS and HTTPS infrastructure

Fake Google HTTPS Certificates Came From Hijacked Country Domains

Related Posts