I recently covered TypesafeAI’s JEV in TDS (link at the end). JEV makes fast decisions, such as classifying content or choosing the next step in a workflow, and returns typed answers that software can use directly. The accompanying probabilities help an application decide whether to act on an answer or pass it to a person for review.
The internet went a bit crazy over JEV, so it was no surprise to see rival products appear after its release. Among the most prominent was the Decisions API that OpenAI previewed at its recent DevDay and made available as a public beta shortly after.
In this article, I’ll look at the new API, how to get it, and some practical use-case examples, and I’ll let you know whether I think it’s a JEV killer or whether TypesafeAI can sleep soundly at night.
The background to JEV-like models
Suppose you run an online shop. A customer writes to say that their keyboard arrived with three broken keys and asks for a replacement. Before anyone replies, your application needs to decide which team should handle the message.
You could ask a language model to classify it and return JSON. But the application only needs a category and enough information to decide whether to trust that category. That process could also be relatively slow and costly if it had to handle 1000s of requests.
OpenAI’s Decisions API gives that operation its own endpoint. Additionally, OpenAI claims its Decisions API can make decisions up to ten times faster than GPT-6 Luna, the LLM that Decisions is based on, through the Responses API. To explore how it works, we’ll walk through three Python examples and compare OpenAI’s approach with TypesafeAI’s JEV model.
One major advantage of the Decisions API over JEV is that it can handle images. We’ll see an example of that later.
What the OpenAI Decisions API returns
A request supplies a model, some input and a list of questions. The current beta supports the gpt-6-luna LLM. Questions share the input, which can contain text, images or both.
There are three question types:
These fields are defined in the SDK’s response types. A predicate doesn’t return a Boolean: your code chooses the threshold that turns its probability into an action.
The API returns answers in question order. It can also refuse individual questions, so check an answer’s type before reading its other fields.
Defining the available answers is part of designing the application. Suppose your choices are only billing and delivery, but the customer asks about opening hours. Neither response is appropriate, but adding a general category gives the system a sensible place to route that request. Test the categories against real customer messages, including requests that don’t fit any of them.
With the Decisions API, you choose from three question types: predicate, choice and score. Each returns a defined answer format, including probabilities. If you need to extract an invoice number, a customer name and a list of purchased items into your own JSON structure, use Structured Outputs with the Responses API.
When to use the Decisions API and what it costs
A support system processing thousands of messages may only need a queue name before passing each request to the right team. The Responses API can produce that answer, but the Decisions API provides a dedicated interface for questions with defined answers: a yes/no probability, a choice from a list or a score against a rubric.
You send those questions to /v1/decisions and use the returned answers in your application logic. That makes it a useful option for routing messages, selecting a search index or choosing an agent’s next action from a permitted list.
It also has separate pricing. At launch, OpenAI lists $0.10 per million input tokens, with no charges for output tokens, cache reads, or cache writes. Long-context multipliers and regional processing premiums still apply.
As a simple calculation, one million requests averaging 1,000 billable input tokens would cost $100 at that base rate. Include the questions and their descriptions when estimating input size.
A task that needs a written explanation or an extracted object with arbitrary fields still needs a generation interface. It doesn’t make sense to squeeze an invoice extraction problem into twenty classification questions just because the endpoint is fast.
Equally, keep simple rules in Python. If priority depends only on an order total exceeding £500, for example, compare the number directly in Python code itself. A model becomes useful when the input expresses something your rules can’t easily recognise, such as a customer describing the same fault in several different ways. Even then, the extra network call must save enough downstream work to justify its latency.
Set up Python with the API
The official openai-python repository added Decisions support in version 3.26.0. Install that version in your environment:
You’ll need an OpenAI API key. If you don’t already have one, you need to make sure you have registered with OpenAI and added a payment method with some credit to your account. Afterwards, go to https://platform.openai.com/home. On the left side of the screen, you’ll see an API keys link. Click on that and follow the instructions to create a new secret key.
Set the OPENAI_API_KEY environment variable to your API key. Do that in PowerShell like this:
Each example below is a complete, standalone Python program with its own imports and client setup.
The SDK implementation exposes client.decisions.create() and sends the request to the /v1/decisions endpoint. This is the client implementation; inference runs on OpenAI’s service.
Save each program using the filename shown in its section, then run it directly with Python. The three files work independently. Each program makes a billable API request when you run it.
Code example 1: Routing a support request
Our shop has three specialist queues and a general queue for anything that doesn’t fit. Create a file named decisions_route.py and add this code.
The (correct) output to my question about a broken keyboard was:
When I asked a different question in the code,
I got this output, which is correct again.
The choice objects contain values and descriptions, as specified in the SDK’s request types. Descriptions let us distinguish damaged goods from a delivery that never arrived.
The 0.8 threshold is illustrative. It isn’t an OpenAI recommendation, and it doesn’t establish an 80% success rate. Start with messages people have already classified, and measure mistakes at several thresholds before you decide on your threshold.
Code Example 2: Check a document for missing information
Suppose staff write return instructions in several formats. We want to flag instructions that omit either a deadline or a postal address. Create the file decisions_document.py with this content.
My output was:
The wording is important. Being told to request an address isn’t the same as being given one. A search for the word address would miss that distinction.
Both questions can share one request because neither depends on the other’s answer. If a later question needs an earlier result, make another request after inspecting that result.
This checks whether information appears in the text. Checking whether an address exists or whether a deadline matches your business rules needs separate validation.
Code Example 3: Inspect a product photograph
Our last example uses three images of a parcel, one heavily damaged, one with slight damage and the other completely undamaged. We’ll see if the model can distinguish between damaged/undamaged. Here are the images I used. All were in .PNG format.


Place all three images in the same location as the Python scripts, called, say, undamaged.png, heavy_damage.png and slight_damage.png. Next, create a file called decisions_image.py with this code.
Images must use inline data URLs; ordinary web URLs and file IDs aren’t accepted. The endpoint supports up to 128 images per request, and its message input supports only the user role with text and image parts.
Here are my outputs:
That last result surprised me. The package was only slightly damaged, but the model identified it with high probability. That’s pretty impressive, though I accept that even slight damage can be visible and may lead to a high score.
Summary: How Decisions compares with JEV
JEV addresses a similar programming problem to the Decisions API. Its interface evaluates shared state using three primitives: Noul, Choice, and Score. Noul returns the probability of a yes/no answer, making it the closest equivalent to OpenAI’s predicate, which we used in examples 2 and 3. JEV’s choice primitive is largely the same as OpenAI’s, which we used in our first example.
The products behind those interfaces differ. OpenAI exposes GPT-6 Luna through a specialised endpoint. TypeSafeAI describes JEV as a model built for decisions, with a parallel sampler and Reinforcement Learning for Calibrated Decisions, or RLCD.
The practical comparison, as of early October 2026, looks like this:
JEV’s listed input token cost is 58% lower, although different tokenisation and request sizes affect the bill. Its current context limits are 64,000 tokens for the complete request and 32,000 for the state plus the longest question.
TypeSafeAI reports 70–500 millisecond response times in its launch post. During my testing of the OpenAI Decisions API, response times often ran into multiple seconds, but I didn’t time a direct comparison between the two products.
For both systems, Python still controls what happens next after the models return their results. Keep arithmetic and fixed business rules in ordinary code, and use a generative model when the task is more ambiguous and needs a written explanation.
So, finally, do I think JEV should be worried? No, not yet, at least. In my experience, JEV has the advantage of speed; Decisions API has the advantage of interpreting images. But how long do you think it will be before JEV can process images?
You can read my original TDS article on JEV here.

