OpenAI Decisions API Hits Public Beta With 10x Faster Typed Answers

OpenAI Decisions API Hits Public Beta With 10x Faster Typed Answers


OpenAI has released its Decisions API in public beta, an endpoint that converts text and images into typed answers your code can branch on directly. According to OpenAI, the Decisions API runs roughly 10x faster than the Responses API. It targets a common developer pattern: prompt a large language model, then parse its text output into a usable label.


TL;DR


  • Model size: The GPT-6 Luna parameter count is not disclosed. Its model card lists a 1,050,000-token context window.
  • Deployment: OpenAI-hosted API only, via POST /v1/decisions. No open weights and no self-hosting.
  • Performance: About 10x faster than the Responses API, per OpenAI.
  • Pricing: $0.10 per 1M input tokens, with no output, cache-read, or cache-write charges.
  • Bottom line:
  • Best for: Fast, typed decisions with probabilities.
  • Worst for: Teams needing multiple model choices, a stable non-beta API, or independent benchmark data (none are available yet).

What Is the OpenAI Decisions API?


The Decisions API is an OpenAI endpoint that evaluates text, images, or both, and returns typed answers instead of free-form prose. Rather than generating a paragraph you must then parse, the API returns structured values your application can consume directly. This makes it well suited to classification, routing, gating, and other decision logic inside AI agents and automated workflows.


The timing reflects a broader shift in 2026: as agentic systems move into production, developers increasingly need deterministic, low-latency decision primitives rather than full conversational generation. The Decisions API is OpenAI's answer to that demand, trading general-purpose text generation for speed and type safety.


Why It Matters


Parsing LLM output into labels is a longstanding source of fragility in AI pipelines. The Decisions API removes that step by returning typed results natively, which can reduce latency, cut token costs, and simplify downstream error handling. For teams building high-volume agents, that combination may be more valuable than raw model capability.


Caveats


The API remains in public beta, runs on a single model, and has not yet been independently evaluated. Teams with strict stability or multi-model requirements should weigh those constraints before adopting it for production-critical paths.

via MarkTechPost

Related