Oct 7 edition/Reporting & analysis
ModelsAgentsBusiness

ModelsArchitectures & capability

OpenAI opens a public beta of its Decisions API on GPT-6 Luna, which returns typed answers instead of text

OpenAI has opened a public beta of its Decisions API, which uses gpt-6-luna to return probabilities, choices or scores instead of text. It follows TypeSafe AI's Jev, costs more per input token, but adds image inputs.

THE CORE IDEAS3 TAKEAWAYS
01

The new endpoint is POST /v1/decisions, and gpt-6-luna is the only model it runs on. It answers three kinds of question: a yes/no probability between 0 and 1, a pick from a fixed list with a confidence value, or a probability-weighted score across ordered levels. Inputs can be text, inline base64 images, or both. Input costs $0.10 per million tokens, and output and cache tokens are free. OpenAI says general availability will come in the coming weeks but has not given a date. [3] [7]

02

The launch came three weeks after TypeSafe AI introduced Jev, a model that makes typed decisions and generates no text. OpenAI previewed its version at DevDay. Jev costs $0.042 per million input tokens, so it is about 2.4 times cheaper per input token, but it accepts only text and JSON. OpenAI's API also accepts images. [4] [5] [6]

03

On launch day, Simon Willison released an alpha plugin for his llm tool by adapting his earlier plugin for Jev. Because the two APIs share the same three question types, the work carried over easily, which suggests switching between providers could be fairly simple. The plugin returns JSON with a probability for each option and each level. [1] [8] [11]

WHY IT MATTERS

OpenAI's documented pricing charges only for input tokens, so high-volume routing, moderation and agent-step decisions could become cheap. Whether the service is fast and accurate is unproven.

Read the full assessment

The latency figures come from vendors, and there are no independent calibration or head-to-head results.

Sources are gathered and checked against the OpenAI docs; writing the dossier now.

Executive brief

OpenAI has released a decision-only model just three weeks after TypeSafe AI's Jev made that category visible. The Decisions API runs only on gpt-6-luna. It returns typed answers instead of text: a probability, a choice from a fixed list, or a score. It costs $0.10 per million input tokens, and output is free. That is more than twice Jev's reported $0.042, but OpenAI's version also takes images. On launch day, Simon Willison released the llm-openai-decisions 0.1a0 plugin. He says he had GPT-6 Astra build it from OpenAI's documentation. OpenAI's latency claim has not been independently benchmarked.

What changed and event timeline

  1. TypeSafe launches Jev

    Jev is a "System One" model that generates no text and answers typed choice, score and yes/no ("Noul") questions. TypeSafe lists the price at $0.042 per million input tokens (;).

  2. DevDay preview

    OpenAI previewed a Decisions API that picks from predefined options in a fraction of a second. Willison read it as OpenAI's response to Jev ().

  3. Pricing still unknown

    At this point the Decisions API was in limited preview and OpenAI had not published its pricing ().

  4. Public beta

    POST /v1/decisions opened in public beta on gpt-6-luna at $0.10 per million input tokens. OpenAI expects general availability "in the coming weeks" (;).

  5. Plugin release

    Willison released llm-openai-decisions 0.1a0, based on his earlier llm-typesafe plugin for Jev (;).

Capabilities and access

  • Model: gpt-6-luna is the only supported model. The endpoint is POST /v1/decisions (OpenAI).
  • Question types:
  • predicate: a probability from 0 to 1.
Read the full section
  • Model: gpt-6-luna is the only supported model. The endpoint is POST /v1/decisions (OpenAI).
  • Question types:
  • predicate: a probability from 0 to 1.
  • choice: one option from a fixed list, plus a confidence value.
  • score: a probability-weighted score across ordered levels.
  • Inputs: Text, images, or both. Images must be inline base64. Hosted URLs and file_id inputs are rejected.
  • Pricing: $0.10 per million input tokens. Output and cache tokens are not charged.
  • Compliance: Supports Zero Data Retention, HIPAA eligibility, and US/EEA data residency.
  • Plugin: Install with llm install llm-openai-decisions. It accepts up to 128 images per request (README).

Technical analysis for researchers and developers

  • API design: It is a classifier-style interface with no text decoding. Several independent questions can go in one request.
  • Calibration: OpenAI recommends setting thresholds from labeled examples, weighing the cost of false positives against false negatives.
  • Undocumented internals: OpenAI has not published the architecture or how a probability is extracted from Luna.
Read the full section
  • API design: It is a classifier-style interface with no text decoding. Several independent questions can go in one request. Decisions that depend on earlier answers need separate calls (OpenAI).
  • Calibration: OpenAI recommends setting thresholds from labeled examples, weighing the cost of false positives against false negatives. This means teams need their own evaluation sets.
  • Undocumented internals: OpenAI has not published the architecture or how a probability is extracted from Luna.
  • Plugin behaviour: The plugin's example passes an HTTPS image URL with -a, and the README says it accepts file paths and raw bytes. This suggests the plugin converts images to the base64 format the API requires. That is an inference; the README does not say so.
  • Output format: Results come back as JSON with per-option and per-level probabilities (README).

Claims and evidence

  • "About 10x faster than the Responses API": Vendor-reported (OpenAI). The MIXED report notes that no benchmark was published.
  • ~150 ms end to end: Reported as coming from early demonstrations (AI Weekly). Not independently verified.
  • Jev claims (70–500 ms, "193x faster, 444x cheaper"): Vendor-reported (Futura-Sciences).
Read the full section
  • "About 10x faster than the Responses API": Vendor-reported (OpenAI). The MIXED report notes that no benchmark was published. The figure compares against Luna's own Responses API, not against Jev (Valyu).
  • ~150 ms end to end: Reported as coming from early demonstrations (AI Weekly). Not independently verified.
  • Jev claims (70–500 ms, "193x faster, 444x cheaper"): Vendor-reported (Futura-Sciences).
  • Head-to-head comparison: No matched independent test of the two services exists in the reviewed sources (Valyu).
  • Plugin example: Asked whether a photo of two pelicans contains mammals, the model returned probability 0.0 (Willison). Pelicans are birds, so the answer is right, but it is one example.

Context and prior work

  • Jev, from TypeSafe AI (founded by ex-OpenAI researcher Diogo Almeida), introduced the "System One" framing: fast, typed decisions with a confidence value instead of generated text (Futura-Sciences).
  • Willison's `llm-typesafe` plugin already exposed Jev's three question types.
  • The two products differ on input: Jev takes text and JSON only, while OpenAI's API also takes images (Valyu).
Read the full section
  • Jev, from TypeSafe AI (founded by ex-OpenAI researcher Diogo Almeida), introduced the "System One" framing: fast, typed decisions with a confidence value instead of generated text (Futura-Sciences).
  • Willison's `llm-typesafe` plugin already exposed Jev's three question types. OpenAI matches that structure conceptually (Willison).
  • The two products differ on input: Jev takes text and JSON only, while OpenAI's API also takes images (Valyu).

Limitations, safety and contested findings

  • Documentation gap: The gpt-6-luna model page has an 18-row table of supported endpoints, and the Decisions endpoint is not in it (MIXED).
  • Pricing differs from standard Luna: Elsewhere Luna charges $0.125 for cache writes and $0.50 per million output tokens.
  • Image handling: Images must be inline base64, up to 128 per request. Processing runs in US and EU regions only.
Read the full section
  • Documentation gap: The gpt-6-luna model page has an 18-row table of supported endpoints, and the Decisions endpoint is not in it (MIXED).
  • Pricing differs from standard Luna: Elsewhere Luna charges $0.125 for cache writes and $0.50 per million output tokens. The Decisions endpoint charges for neither (same source).
  • Image handling: Images must be inline base64, up to 128 per request. Processing runs in US and EU regions only.
  • Beta status: The API is still in beta, and OpenAI has given no date for general availability.
  • Accuracy: No independent accuracy or calibration results are available for either vendor.

Business and practitioner implications

  • Cost: Because output is free, the bill depends only on how much context goes in. Jev remains about 2.4x cheaper per input token.
  • Images: Pipelines that need image decisions currently have only OpenAI's option.
  • Switching providers: The two APIs share the same three question types, so switching should be fairly easy.
Read the full section
  • Cost: Because output is free, the bill depends only on how much context goes in. At $0.10 per million input tokens, high-volume work like routing, moderation and choosing an agent's next action becomes cheap. Jev remains about 2.4x cheaper per input token.
  • Images: Pipelines that need image decisions currently have only OpenAI's option.
  • Switching providers: The two APIs share the same three question types, so switching should be fairly easy. Willison built the new plugin by adapting the Jev one.
  • Before production: Build labeled evaluation sets to set thresholds, and measure latency yourself rather than relying on vendor figures.

Sources

Read the full section
FOLLOW THE EVIDENCE

The source trail.

Sources (11)
A LITTLE LESS NOISE. A LOT MORE CONTEXT.

Stay curious.
Follow the evidence.

Independent perspectives, the original sources, and room for the questions that don't have easy answers.

How we build the brief