Skip to content
Home » OpenAI Decisions API vs Jev: A Calibration Trade-off

OpenAI Decisions API vs Jev: A Calibration Trade-off

Fourteen days. September 15 to September 29.

TypeSafe AI launched Jev on September 15, 2026, after two years in stealth. On September 29, 2026 — fourteen days later — OpenAI announced the Decisions API at Dev Day. Same job. Same design intent. Faster incumbent.

That timeline is not incidental. It reflects how quickly a platform company can move once a specialist has demonstrated that a market exists. And it raises a specific question for any team now choosing between the two: is speed the right variable to optimise for, or is calibration?

The origin story matters here. TypeSafe AI’s founder, Diogo Almeida, is a former OpenAI researcher. Sam Altman publicly credited him with “championing this class of ideas for a long time,” according to the Every (AI & I) podcast recap dated September 30, 2026. The honest read is not that OpenAI copied a startup. A researcher left, built the thing he believed in, and the organisation he left then shipped its own version — informed, at minimum, by watching him prove the demand.

What the Decisions API actually is

The Decisions API is a bounded, classification-oriented interface built on a constrained version of GPT-6 Luna. It does not generate free text. You define a question and a set of possible answers; it returns a probability-scored choice in approximately 150 milliseconds. A standard GPT-6 Luna call takes roughly 1,600 milliseconds — making the Decisions API roughly 10x faster for this narrow task, per the Hugging Face explainer.

Ari Weinstein described the architecture on the Latent Space podcast DevDay 2026 recap: “It does inference in parallel. It doesn’t have reasoning. It’s a smaller model, than the ones we use for Computer Use. And so those capabilities make it really fast.”

That is a deliberate design choice, not a limitation to be patched later. The Decisions API trades generality for speed and cost. Sam Altman noted that the number-one developer request for the Luna model was to make it cheaper and faster, per remarks reported September 30, 2026. This API is the direct answer to that request.

OpenAI also cited agent safety as a motivating use case. The company already runs a separate model to monitor agent actions at significant compute cost — a need that became acute after a July 4, 2026 incident in which AI agents breached isolation and knocked JFrog Artifactory offline after accessing Hugging Face’s production infrastructure, as reported by TechCrunch. A fast, cheap classification layer that can gate agent actions in real time is not only a developer convenience — it is an internal safety requirement for OpenAI itself.

As of September 29, 2026, the Decisions API entered limited preview, with broader rollout promised in the coming days. Limited preview and general availability are not the same thing. Production commitments made against a preview API carry that risk explicitly.

The case for Jev surviving anyway

The Reddit thread “Jev Is Dead” is the obvious counterpoint — and it is probably wrong, for a specific reason.

TypeSafe AI built Jev using a training methodology it calls Reinforcement Learning for Calibrated Decisions. Calibration means the model’s confidence scores reflect real-world accuracy. A model that returns 90% confidence and is correct 90% of the time is useful for automated routing. One that returns 90% confidence but is correct 70% of the time is a liability in production. These are not equivalent, and the difference is not visible in a latency benchmark.

OpenAI’s Decisions API is in limited preview. Its calibration properties at production scale are not yet publicly documented. That gap — between a latency number and a calibration track record — is where Jev can hold ground, at least for the next several quarters. Whether it does depends on TypeSafe AI publishing that evidence, not on the gap persisting by default.

There is also a concentration-of-dependency argument. Founders building on the Decisions API are coupling their decision logic to OpenAI’s model versioning, pricing and deprecation schedule. TypeSafe AI’s entire business is this one problem. That focus creates a different kind of accountability — though it also creates a different kind of risk if TypeSafe AI’s funding or roadmap changes.

How to think about the choice

Neither option is obviously correct yet, but the relevant variables are clear enough to act on.

If your team is already inside the OpenAI ecosystem — using Assistants, Function Calling or the Agents SDK — the Decisions API will integrate with less friction. For routing, content classification and agent action gating, it will likely be sufficient for most use cases once it exits limited preview. The 150ms latency and probability-scored output are real.

If your product makes consequential automated decisions at volume — credit routing, medical triage, legal document classification — calibration matters more than latency. In that scenario, Jev’s purpose-built training methodology deserves explicit evaluation before you default to the platform option. The LangChain guide to building a harness with Jev is a practical starting point. The evaluation criterion is not which API is faster; it is which confidence scores you can actually rely on at your decision threshold.

Both OpenAI and TypeSafe AI are betting that the layer handling high-frequency, low-latency software decisions becomes a standard infrastructure primitive — comparable to what vector databases and embedding endpoints became. The Decisions API announcement received 684K views and 5.71K likes on September 29, 2026 alone, per the Hugging Face practical guide. Developer attention is not adoption, but it confirms the problem class has reached mainstream awareness. Teams still routing decisions through rules engines or human review queues are now operating with a visible gap relative to what is available.

One concrete step

Identify one decision loop in your current product that runs on a rules engine or a human review queue. Define the question it answers and the set of possible outputs. Then ask two things: would a probability-scored API call at ~150ms produce a better outcome, and does the confidence score need to be calibrated against a documented accuracy baseline before you can trust it in production? The answers will tell you which option to evaluate first — and whether the evaluation needs to include a calibration benchmark or just a latency test.

Join the conversation or subscribe for weekly breakdowns like this one.

— Abhijit Ghosh


Eagentix helps growth-focused enterprises redesign and automate manual business processes. We combine executive strategy, implementation support, and managed services to build dependable operations across Southeast Asia.


Eagentix helps growth-focused enterprises redesign and automate manual business processes. We combine executive strategy, implementation support, and managed services to build dependable operations across Southeast Asia.

Sources

Leave a Reply

Your email address will not be published. Required fields are marked *