OpenAI's Decisions API enters Jev's territory
Finite answers, confidence scores, image input and the bill: what OpenAI announced, what Jev documents, and where the sales pitch outruns the evidence.
A surprising amount of production AI ends up answering a rather small question. Which queue should this ticket enter? Does this document contain the required evidence? Should the agent try another search or hand the problem to a person?
We have been asking models capable of writing an essay to pick an enum. Sometimes they write the essay anyway. Then another engineer gets to maintain the parser.
OpenAI's Decisions API moves directly into this territory. It also puts pressure on the pitch behind Jev, TypeSafe's model for structured decisions. There is a useful idea here, and a less useful race to declare ordinary software obsolete because an API can now return billing.
What OpenAI actually announced
The DevDay recap describes a Luna-based API that answers developer-defined questions from a finite set of permitted answers. Context can be text or images. Suggested uses include classification, request routing and selecting an agent's next action. Access is in limited preview, with a broader release planned.
That is the confirmed scope. The announcement does not publish an endpoint, request schema, price, latency measurements or a probability contract. I did not find a public Decisions API reference during this check. An SDK example with client.decisions.create() would look convincing and would also be something I made up.
OpenAI already offers Structured Outputs, including schema-constrained enum values. Bounded answers alone are therefore not a new invention. A dedicated decision interface could improve the economics or behaviour of this work. The announcement gives us a reason to investigate that, not a measurement of how much it improves.
Jev has a more specific public contract
Jev is designed around typed questions over shared input, called state. Its documentation defines three primitives: Choice for selecting an option, Score for a rubric, and Noul for a yes/no probability. Questions are evaluated independently against the same state, rather than as a conversation in which one answer becomes the next question's context.
For a Choice, the documented response includes the selected option, a probability distribution over the options and a confidence statistic. Those values give application code more to work with than a bare label. Whether they are reliable enough for your workload remains an evaluation question.
A note about the supplied link: jevai.net advertises Jev, but its example uses /v1/decide and quotes $0.084 per million input tokens. TypeSafe's quickstart documents /v1/systemone; its model reference lists $0.042. I use TypeSafe's documentation for the integration and pricing below. Those discrepancies are a reason to check the provider, rather than merge two different pages into one imaginary API.
| Question | OpenAI Decisions API | Jev, through TypeSafe |
|---|---|---|
| Core interface | Questions with predefined answers | Typed questions over shared state |
| Input | Text and images | Text, including structured text data |
| Uncertainty | Not specified in the recap | Choice and Score expose probabilities and confidence |
| Public integration details | No reference found during this check | HTTP API and SDK examples |
| Price | Not specified in the recap | $0.042 per million input tokens; output free |
| Performance evidence | No measurements in the recap | Vendor results, requiring workload-specific validation |
The OpenAI column comes from its announcement. The Jev column comes from the linked TypeSafe documentation and its launch report. Blank details are not evidence of missing capability. They are details we cannot compare yet.
A ticket that belongs to two teams
Consider this fictional support request:
I have been charged twice for our team plan. None of us can open the workspace. We need access back before this afternoon's demo.
An exclusive billing or access classification throws away part of the request. Ask which team owns the immediate blocker, then separately ask whether billing work is needed. For this example, the routing policy says restoring workspace access takes priority when both problems appear.
That policy has to be written down. A model cannot recover a company rule nobody supplied.
Here is how the routing question can be expressed with TypeSafe's documented Python SDK. The ticket and criteria are original example data. The snippet makes a real request if you supply a TYPESAFE_API_KEY; it was not executed for this article.
from typesafe_sdk import Choice, TypeSafeClient
ticket = (
"I have been charged twice for our team plan. "
"None of us can open the workspace. "
"We need access back before this afternoon's demo."
)
with TypeSafeClient(model="jev-1.13.0") as client:
result = client.system_one(
state={"ticket": ticket},
questions={
"primary_queue": Choice(
instructions="Which queue owns the immediate blocker?",
criteria={
"access": "Workspace access is blocked, even if billing is also mentioned.",
"billing": "A charge or invoice issue without blocked workspace access.",
"manual_review": "The request is unclear or neither queue fits.",
},
),
},
)
answer = result.answers["primary_queue"]
print(answer.choice, answer.probabilities, answer.confidence)
The call follows the quickstart and Choice reference. A production version would add the separate billing question, retain the original ticket and handle API failures. An other or review option matters whenever your taxonomy fails to cover the input.
The same ticket is a reasonable evaluation case for Decisions API when its contract is available. The comparison should use equivalent questions and routing rules. Otherwise we are benchmarking one prompt author's interpretation of "urgent" against another's.
A confidence field is not a warranty
Imagine an illustrative distribution of access: 0.72, billing: 0.23, manual_review: 0.05. These are invented numbers, not an observed Jev response.
The winning label is access. That does not settle whether the system should act. Your application might require stronger evidence before routing automatically, or retain a billing flag while a person checks the blocker.
Jev's confidence documentation says confidence is derived from how concentrated the probability distribution is. It is not simply another name for the selected option's probability. A sharp distribution also does not prove the choice is correct on your data.
Calibration has to be checked across labelled cases. Among decisions assigned roughly 0.9 probability, do about nine in ten turn out right? Does that still hold for short angry tickets, unfamiliar product names and messages containing two issues? Those are proposed tests, not results either vendor has delivered for this example.
The useful outcome is knowing how much traffic you can automate at an acceptable error rate. A service that gets impressive accuracy by sending almost everything to manual review may be correct and still fail to reduce much work.
Images change the comparison
A second hypothetical workflow: a customer uploads a parcel photo and asks for a replacement. You need to distinguish visible damage, an intact parcel and an image too unclear to assess.
Native image input is an announced Decisions API capability. Jev's current model documentation explicitly says text only. A Jev pipeline needs another component to describe the image or extract relevant features first.
That extra component adds cost, latency and another place to lose evidence. A description saying "box damaged" may omit that the product inside looks intact. OCR cannot recover that visual distinction by reading the shipping label harder.
This gives OpenAI a practical advantage for trying the photo workflow. It does not establish better visual accuracy. Evaluate the complete pipelines, including unclear images. The application still checks the order, replacement eligibility and whether a replacement was already issued. A model's opinion about a photograph is not your returns policy.
The "no hallucinations" claim needs a smaller font
The Jev marketing pitch includes "no hallucinations". Restricting a model to permitted answers prevents it from inventing an extra category. It does not prevent choosing the wrong permitted category.
Sending a billing issue to shipping is still a mistake, even when the JSON is immaculate. Congratulations on the valid enum. The customer remains in the wrong queue.
TypeSafe's own Jev 1.13 limitations are more useful than that slogan. They discuss literal interpretation, unreliable arithmetic and date comparisons, distraction from irrelevant context, and adversarial content. The suggested remedies include clearer criteria, smaller relevant inputs and keeping deterministic work in code.
Read that page before the launch graphics. It tells you where the system may need help. The same separation applies when evaluating OpenAI: a constrained answer and a correct answer are different properties.
The bill is interesting. The speedup needs context
At TypeSafe's published $0.042 per million input tokens, one million calls averaging 1,000 billed input tokens would cost $42 in model input charges. That is arithmetic based on the listed rate, not a quote for a production deployment. Include the question definitions in the token budget. Image preprocessing, retries and the rest of your infrastructure sit outside that calculation.
There is no comparable Decisions API price in the recap, so a claim that Jev is cheaper than it would be premature.
TypeSafe's launch report advertises 70-500 ms responses and large speedups over LLM workflows. It also explains that many measurements ran from West Coast laptops, that short input favours its demo, and that workflow reference probabilities came from other frontier models. The comparison wrappers request probabilities too, adding work relative to returning a label alone.
Those disclosures matter. Agreement with a reference model is not the same measurement as correct routing against independently labelled tickets. The report also does not test OpenAI's newly announced Decisions API.
A 200x speedup over a different workflow makes a lovely headline. It does not answer how long this ticket waits in your queue under Tuesday's traffic.
For an actual comparison, use the same held-out cases and record wrong automatic decisions, manual-review rate, median and p95 latency, failures and total cost. Test from the deployment region at realistic concurrency. Log model versions so a moving alias does not quietly invalidate the threshold you tuned.
Where this leaves Jev
My reading is that OpenAI validates the demand while making the broad pitch less distinctive. Developers want models that fit inside ordinary code and answer bounded questions. Jev now has to compete on the quality and usefulness of its probabilities, its price, and its performance on those questions. Merely pointing out that chat models generate strings will carry less weight.
OpenAI has an announced route for image-based decisions, but too little published detail here to award it the whole category. Jev exposes enough of its contract to build a serious text-workflow evaluation. Neither deserves a production migration on the strength of a launch paragraph.
The design direction is sensible: narrow the model's job and let code own the consequences. Give it the ticket classification. Keep duplicate-payment checks, permissions and replacement limits in the application.
The enum still needs an adult in the room. Usually that adult is a few boring lines of code.