Skip to content
All articles
  • Applied AI
  • OpenAI
  • Decision models

OpenAI's Decisions API enters Jev's territory

Finite answers, confidence scores, image input and the bill: what OpenAI announced, what Jev documents, and where the sales pitch outruns the evidence.

Eduard Smirnov9 min read

A surprising amount of production AI ends up answering a rather small question. Which queue should this ticket enter? Does this document contain the required evidence? Should the agent try another search or hand the problem to a person?

We have been asking models capable of writing an essay to pick an enum. Sometimes they write the essay anyway. Then another engineer gets to maintain the parser.

OpenAI's Decisions API moves directly into this territory. It also puts pressure on the pitch behind Jev, TypeSafe's model for structured decisions. There is a useful idea here, and a less useful race to declare ordinary software obsolete because an API can now return billing.

What OpenAI actually announced

The DevDay recap describes a Luna-based API that answers developer-defined questions from a finite set of permitted answers. Context can be text or images. Suggested uses include classification, request routing and selecting an agent's next action. Access is in limited preview, with a broader release planned.

That is the confirmed scope. The announcement does not publish an endpoint, request schema, price, latency measurements or a probability contract. I did not find a public Decisions API reference during this check. An SDK example with client.decisions.create() would look convincing and would also be something I made up.

OpenAI already offers Structured Outputs, including schema-constrained enum values. Bounded answers alone are therefore not a new invention. A dedicated decision interface could improve the economics or behaviour of this work. The announcement gives us a reason to investigate that, not a measurement of how much it improves.

Jev has a more specific public contract

Jev is designed around typed questions over shared input, called state. Its documentation defines three primitives: Choice for selecting an option, Score for a rubric, and Noul for a yes/no probability. Questions are evaluated independently against the same state, rather than as a conversation in which one answer becomes the next question's context.

For a Choice, the documented response includes the selected option, a probability distribution over the options and a confidence statistic. Those values give application code more to work with than a bare label. Whether they are reliable enough for your workload remains an evaluation question.

A note about the supplied link: jevai.net advertises Jev, but its example uses /v1/decide and quotes $0.084 per million input tokens. TypeSafe's quickstart documents /v1/systemone; its model reference lists $0.042. I use TypeSafe's documentation for the integration and pricing below. Those discrepancies are a reason to check the provider, rather than merge two different pages into one imaginary API.

Question OpenAI Decisions API Jev, through TypeSafe
Core interface Questions with predefined answers Typed questions over shared state
Input Text and images Text, including structured text data
Uncertainty Not specified in the recap Choice and Score expose probabilities and confidence
Public integration details No reference found during this check HTTP API and SDK examples
Price Not specified in the recap $0.042 per million input tokens; output free
Performance evidence No measurements in the recap Vendor results, requiring workload-specific validation

The OpenAI column comes from its announcement. The Jev column comes from the linked TypeSafe documentation and its launch report. Blank details are not evidence of missing capability. They are details we cannot compare yet.

A ticket that belongs to two teams

Consider this fictional support request:

I have been charged twice for our team plan. None of us can open the workspace. We need access back before this afternoon's demo.

An exclusive billing or access classification throws away part of the request. Ask which team owns the immediate blocker, then separately ask whether billing work is needed. For this example, the routing policy says restoring workspace access takes priority when both problems appear.

That policy has to be written down. A model cannot recover a company rule nobody supplied.

A fictional ticket mentions a duplicate charge and blocked workspace access. Separate questions identify the primary blocker and the billing issue. Application code then routes the access problem, preserves the billing follow-up, or sends an uncertain case to manual review.
A proposed workflow for the fictional ticket. These are application decisions, not measured responses from either provider. Open the diagram to view it at full size.

Here is how the routing question can be expressed with TypeSafe's documented Python SDK. The ticket and criteria are original example data. The snippet makes a real request if you supply a TYPESAFE_API_KEY; it was not executed for this article.

from typesafe_sdk import Choice, TypeSafeClient

ticket = (
    "I have been charged twice for our team plan. "
    "None of us can open the workspace. "
    "We need access back before this afternoon's demo."
)

with TypeSafeClient(model="jev-1.13.0") as client:
    result = client.system_one(
        state={"ticket": ticket},
        questions={
            "primary_queue": Choice(
                instructions="Which queue owns the immediate blocker?",
                criteria={
                    "access": "Workspace access is blocked, even if billing is also mentioned.",
                    "billing": "A charge or invoice issue without blocked workspace access.",
                    "manual_review": "The request is unclear or neither queue fits.",
                },
            ),
        },
    )

answer = result.answers["primary_queue"]
print(answer.choice, answer.probabilities, answer.confidence)

The call follows the quickstart and Choice reference. A production version would add the separate billing question, retain the original ticket and handle API failures. An other or review option matters whenever your taxonomy fails to cover the input.

The same ticket is a reasonable evaluation case for Decisions API when its contract is available. The comparison should use equivalent questions and routing rules. Otherwise we are benchmarking one prompt author's interpretation of "urgent" against another's.

A confidence field is not a warranty

Imagine an illustrative distribution of access: 0.72, billing: 0.23, manual_review: 0.05. These are invented numbers, not an observed Jev response.

The winning label is access. That does not settle whether the system should act. Your application might require stronger evidence before routing automatically, or retain a billing flag while a person checks the blocker.

Jev's confidence documentation says confidence is derived from how concentrated the probability distribution is. It is not simply another name for the selected option's probability. A sharp distribution also does not prove the choice is correct on your data.

Calibration has to be checked across labelled cases. Among decisions assigned roughly 0.9 probability, do about nine in ten turn out right? Does that still hold for short angry tickets, unfamiliar product names and messages containing two issues? Those are proposed tests, not results either vendor has delivered for this example.

The useful outcome is knowing how much traffic you can automate at an acceptable error rate. A service that gets impressive accuracy by sending almost everything to manual review may be correct and still fail to reduce much work.

Images change the comparison

A second hypothetical workflow: a customer uploads a parcel photo and asks for a replacement. You need to distinguish visible damage, an intact parcel and an image too unclear to assess.

Native image input is an announced Decisions API capability. Jev's current model documentation explicitly says text only. A Jev pipeline needs another component to describe the image or extract relevant features first.

Two proposed parcel-photo workflows. OpenAI's announced image input allows an image to enter the decision call. Jev currently needs a separate vision component to produce text before its decision call. Both still need application checks before a replacement is approved.
Proposed pipelines based on OpenAI's announced input support and Jev's documented text-only input. The diagram does not claim either pipeline classified a real photograph.

That extra component adds cost, latency and another place to lose evidence. A description saying "box damaged" may omit that the product inside looks intact. OCR cannot recover that visual distinction by reading the shipping label harder.

This gives OpenAI a practical advantage for trying the photo workflow. It does not establish better visual accuracy. Evaluate the complete pipelines, including unclear images. The application still checks the order, replacement eligibility and whether a replacement was already issued. A model's opinion about a photograph is not your returns policy.

The "no hallucinations" claim needs a smaller font

The Jev marketing pitch includes "no hallucinations". Restricting a model to permitted answers prevents it from inventing an extra category. It does not prevent choosing the wrong permitted category.

Sending a billing issue to shipping is still a mistake, even when the JSON is immaculate. Congratulations on the valid enum. The customer remains in the wrong queue.

TypeSafe's own Jev 1.13 limitations are more useful than that slogan. They discuss literal interpretation, unreliable arithmetic and date comparisons, distraction from irrelevant context, and adversarial content. The suggested remedies include clearer criteria, smaller relevant inputs and keeping deterministic work in code.

Read that page before the launch graphics. It tells you where the system may need help. The same separation applies when evaluating OpenAI: a constrained answer and a correct answer are different properties.

The bill is interesting. The speedup needs context

At TypeSafe's published $0.042 per million input tokens, one million calls averaging 1,000 billed input tokens would cost $42 in model input charges. That is arithmetic based on the listed rate, not a quote for a production deployment. Include the question definitions in the token budget. Image preprocessing, retries and the rest of your infrastructure sit outside that calculation.

There is no comparable Decisions API price in the recap, so a claim that Jev is cheaper than it would be premature.

TypeSafe's launch report advertises 70-500 ms responses and large speedups over LLM workflows. It also explains that many measurements ran from West Coast laptops, that short input favours its demo, and that workflow reference probabilities came from other frontier models. The comparison wrappers request probabilities too, adding work relative to returning a label alone.

Those disclosures matter. Agreement with a reference model is not the same measurement as correct routing against independently labelled tickets. The report also does not test OpenAI's newly announced Decisions API.

A 200x speedup over a different workflow makes a lovely headline. It does not answer how long this ticket waits in your queue under Tuesday's traffic.

For an actual comparison, use the same held-out cases and record wrong automatic decisions, manual-review rate, median and p95 latency, failures and total cost. Test from the deployment region at realistic concurrency. Log model versions so a moving alias does not quietly invalidate the threshold you tuned.

Where this leaves Jev

My reading is that OpenAI validates the demand while making the broad pitch less distinctive. Developers want models that fit inside ordinary code and answer bounded questions. Jev now has to compete on the quality and usefulness of its probabilities, its price, and its performance on those questions. Merely pointing out that chat models generate strings will carry less weight.

OpenAI has an announced route for image-based decisions, but too little published detail here to award it the whole category. Jev exposes enough of its contract to build a serious text-workflow evaluation. Neither deserves a production migration on the strength of a launch paragraph.

The design direction is sensible: narrow the model's job and let code own the consequences. Give it the ticket classification. Keep duplicate-payment checks, permissions and replacement limits in the application.

The enum still needs an adult in the room. Usually that adult is a few boring lines of code.