# An analysis of Jev: Faster and cheaper, but watch the accuracy

Decision models show a lot of promise. We reviewed benchmark tests and ran our own to better understand their strengths, shortcomings, and how to improve their accuracy.

Source: https://amplitude.com/en-us/blog/jev-analysis

---

<!--$-->

<!--/$-->

[Vinay Goel and Ram Soma](/blog/author/vinay-and-ram)

[Staff AI Engineers](/blog/author/vinay-and-ram)

[](https://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Famplitude.com%2Fblog%2F%2Fblog%2Fjev-analysis)[](https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Famplitude.com%2Fblog%2F%2Fblog%2Fjev-analysis)[](https://twitter.com/intent/tweet?url=https%3A%2F%2Famplitude.com%2Fblog%2F%2Fblog%2Fjev-analysis\&text=An%20analysis%20of%20Jev%3A%20Faster%20and%20cheaper%2C%20but%20watch%20the%20accuracy)[](mailto:?subject=Checkout%20this%20Amplitude%20Article\&body=Check%20this%20out%3A%20https%3A%2F%2Famplitude.com%2Fblog%2F%2Fblog%2Fjev-analysis)

Decision models are moving into the branch points of the agent loop. A typed judgment routes the ticket, gates the tool call, picks the specialist, passes the answer. Jev, TypeSafe's decision model, runs 70 to 500 milliseconds a call. It’s fast and cheap enough to put at every branch point, so many teams are. We’re working on deploying it ourselves.

However, in reviewing independent benchmark tests and running our own, we’ve found that decision models like Jev share the same pitfalls in accuracy as their slower, more expensive LLM cousins. While the gains in speed and cost-efficiency are impressive, if teams implement a decision model without also having a way to measure the model’s judgment, they’ll just be on a faster and cheaper road to potential slop.

In this post, we share analysis of previous decision model tests, our findings from our own AI analytics benchmark test, and our recommendations for how to use decision models more effectively by adding a measurement layer to your stack. While Jev is our primary point of analysis, our argument is about the class of decision models, not just Jev’s merits.

### What is Jev? What are decision models?

Jev is TypeSafe AI’s flagship [System One model](https://docs.typesafe.ai/concepts/system-one), which the industry has taken to calling generically a “decision model.” Like an LLM, a decision model’s input is text, both natural language and code. Unlike an LLM, it returns typed decisions or probabilities instead of text. Because the scope of its output is small and structured, a decision model works faster and uses fewer tokens.

For example, give an LLM your analytics data and ask, “What stage of my customer journey has the most drop-off?” It may respond with, “Your final checkout step seems to have the greatest drop-off. Want me to dig deeper?” A decision model, with its answer space configured to match your stages, will simply respond with “checkout”.

You can point a decision model at two different kinds of jobs: it can do the work, routing a ticket or gating a tool call in production, or it can grade the work, scoring sessions as a cheap judge. The second is an eval that happens to be fast. The first is new behavior your agent performs all day, and it is the harder one to measure.

Jev is the decision model most people are familiar with, but [open rebuilds are already close behind](https://benchmarkheaven.com/jev-models) on composite rankings.

### Prior benchmark reviews

To implement decision models, the cascade is the pattern nearly everyone has converged on. Run the cheap model across everything, act on the confident answers, escalate the uncertain tail. An [independent study of 791 labeled decisions](https://www.ayautomate.com/blog/jev-vs-llm-benchmark) from AY Automate found that a cascade gated at 0.80 matched frontier accuracy at roughly a quarter of the cost and half the latency.

However, that result only measures how similar the models are, not how often either is right. The same 791-decision study found the decision model agreed with the frontier model 6 to 8 points more often than it matched the ground-truth labels.

The decision model's confident errors clustered on intents with similar meanings, like "Direct debit payment not recognised" versus "card payment not recognised." Five confident answers were wrong that way, and the frontier model also got the same five wrong. Admittedly, this is a small sample, but it is also exactly the failure the cascade exists to catch, and the cascade did not catch it.

Anthony Maio, AI Engineer at Pieces, has [the cleanest expression of this problem](https://anthonymaio.substack.com/p/jev-the-language-model-that-wont): the type system constrains the shape of the output, not the judgment. Clean types, high confidence, no exception, no failed check, and the ticket still went to the wrong queue.

To improve judgment, decision models can be calibrated. But even though models like Jev come pre-calibrated, a single calibration does not fit all data sets or all use cases. [jevbench.xyz](http://jevbench.xyz) records [91.7% on a 60-case tool-call set](https://webofmike.com/jev-benchmark/), and a range of 62.6% to 95.4% across everything it has run. It declines to publish one headline accuracy for exactly that reason. A [77-case BANKING77 pilot](https://sanand0.github.io/llmevals/jev/) found 87.4% mean confidence against 75.3% actual accuracy, about 12 points of overconfidence.

Calibration does not resolve at the node level either. As Maio points out, individually calibrated judgments do not compose into a calibrated workflow once you add thresholds, weights, and branches. Correlated mistakes survive composition. You can verify every gate, router, and verifier in your harness and still ship a system that is systematically wrong.

### Amplitude’s analytics benchmark on Jev

[Ram Soma published this benchmark](https://x.com/ramsoma/status/2101851201684042083) on September 21.

One of the fundamental challenges in analytics is abundance: there are far more charts, measurements, segments, and time windows than a person can continuously monitor. AI shows the potential to solve this not by replacing analysts but by “managing by exception.” It can continuously identify the small number of charts with genuinely interesting behavior, so the analyst can focus on those and ignore ordinary noise.

To test AI models for this application, we created a controlled synthetic benchmark representing common product-analytics work:

- Event segmentation
- Funnels
- Retention
- Customer journeys

The benchmark contained 64 charts and 190 time series. We injected meaningful recent changes into 24 charts. About 15% of the underlying series included spikes, drops, sustained level shifts, segment divergence, volatility changes, and possible instrumentation breaks. The remaining 40 charts were controls designed to look realistic but not require escalation.

For each chart, the model received chart metadata and CSV-formatted time-series data, then answers: *Is there a meaningful recent exception worth escalating to a deeper analytics investigation?*

We compared Jev 1.13 with several strong, inexpensive general-purpose models: Luna, DeepSeek V4.1 Flash, GLM-5.3 Flash, and Sol. (This was not inexpensive.)

A caveat: This is a synthetic benchmark, not a claim that these numbers will transfer unchanged to every organization or analytics implementation. The data was designed to resemble real product-analytics charts, but production evaluation should use a blinded sample of real charts labeled by analysts or chart owners.

Rows of synthetic charts with model responses.

Overall results with precision, recall, cost, and median latency.

Jev caught every meaningful chart in this benchmark. It did so at under half the cost and in under a quarter of the time of DeepSeek V4.1 Flash, the next cheapest and the next fastest model in this test.

Its tradeoff was precision: it escalated more benign charts than the best general-purpose models. That might be a sensible tradeoff for this task, depending on how teams design the follow up step. In a manage-by-exception workflow, missing an important data-quality break, conversion drop, or sudden segment divergence may be more costly than reviewing some extra charts.

While not a “slam-dunk” result, Jev seems like a useful tool in the agent toolbox. At roughly $0.06 per 1,000 charts in this test, Jev makes broad, continuous screening economically feasible. However, Jev’s low precision becomes a concern if every flagged anomaly triggers an expensive root-cause-analysis step: too many false positives can erase the savings from cheap first-pass detection.

### Where a decision model could fit in your stack

Because of their extreme speed and low cost, we think decision models merit inclusion in an agent stack. Balancing those pluses with their moderate accuracy, we think they fit best in the first layer of a system, with higher-cost models used to verify and orchestrate off output.

For an analytics application, a useful production workflow could look like this:

- Run Jev frequently across the full chart inventory
- Filter to charts with a high probability of a meaningful exception
- Send those candidates to a more expensive model, an automated deep-dive workflow, or an analyst
- Tune the threshold based on available review capacity and tolerance for false positives

On its own, though, we don’t believe this kind of workflow is enough to guarantee accuracy. A second stage catches uncertain decisions, but confident errors are not uncertain. They cluster where the schema is ambiguous, and the expensive model inherits the ambiguity too. And although in our test, the expensive LLMs did perform more accurately than the decision model, in other industry tests, the LLMs and decision models both missed against the labeled data set.

### How to best improve decision model accuracy

Every proposed fix for decision model accuracy we’ve seen is more judgment. A bigger model, a reviewer, a hand-labeled set. This is probably why the public evidence for decision models is a few dozen cases in one study and a few hundred in another. That is about what any of us can hand-label in a sitting.

Hand labels are still the right spot check, but they are the wrong primary instrument to improve model accuracy at scale. The right instrument is already sitting in your product data.

We hit a version of this at Amplitude before Jev existed, on our own agent. Earlier this year, we ran [our Agent Analytics product](https://amplitude.com/agent-analytics) over [27,000 Amplitude Global Agent sessions](https://amplitude.com/blog/how-people-use-agents), scoring every one with LLM judges. In 98% of those sessions, the user gave no indication that a judge could use to evaluate if the agent did its job right. No thanks, no complaint, nothing for a judge to read. Similarly, there would be no way for a human, reviewing the transcript, to hand-label the results.

But even though a transcript itself didn’t have a marker for quality, the interaction’s quality could still be inferred from a user’s subsequent behavior. Users whose first session tripped no failure flag, and who saved what the agent produced, retained at 3x the rate of everyone whose first session tripped one. This is correlation, not a causal claim. But no judge, human or AI, could have surfaced it just from the set, because half of what mattered was a *save* event that happened after the session ended.

This is the same gap decision models have now. Trace tooling, like the [Learn Jev harness guide](https://learnjev.com/tutorials/agent-harness) recommends, holds the decision and never sees the outcome.

Agent Analytics bridges that gap by connecting traces with product analytics to unify accuracy signals. Say a decision agent routes a ticket to billing at 0.94 certainty. Was it right? Nobody knows on Tuesday; but the ticket reopens on Thursday. Agent Analytics lets you connect that signal back to the routing decision to improve your model.

Behavioral signals are inherently noisy. In this example, the user may have reopened the ticket for reasons completely unrelated to routing. But by combining signal volume, user segmentation, randomized holdouts, and more, Agent Analytics can help you close that loop effectively.

### Deciding on decision models

The initial results we’ve seen from decision models are encouraging. Jev is a moderately credible and extremely low-cost candidate for the first layer of a cascade system.

However, a decision model’s accuracy pitfalls mean you need additional tools to fine-tune it, just like you do with other models. This can’t be done with a more expensive model alone, since expensive models may have similar accuracy issues, and may also not be able to catch confident errors from a decision model.

In our opinion, the best solution for this is a tool like [Agent Analytics](https://amplitude.com/agent-analytics). Every decision your agent makes gets graded by what the user does next, and that grade is already in your product data. These are observed outcomes, not another model's opinion. They are events, in the same stream, on the same user. A reopened ticket is not a second guess about the decision; it reopened, or it did not. And it keeps arriving, which is the only way to track a number that will not sit still.

We’re excited to continue investigating decision models and seeing how best to implement them. After all, Jev is only a little over a week old. If you’re building with decision models too, let us know what you figure out.

##### Ready to use AI to transform your product?

Amplitude Agents help you understand your users more easily than ever.

[Get started now](/signup?source=blog-ai\&topic=ai\&siteLocation=blog-inline-cta)

About the author

Vinay Goel and Ram Soma

Staff AI Engineers

[More from ](/blog/author/vinay-and-ram)

<!-- -->

Vinay Goel and Ram Soma are both staff AI engineers at Amplitude, working at the cutting edge of AI implementation on the Amplitude platform. They occasionally author blog articles together. For pieces written just by Vinay Goel, see: \[https\://www\.amplitude.com/blog/author/vinay-goel]. For pieces written just by Ram Soma, see: \[https\://www\.amplitude.com/blog/author/ram-soma].

Topics

[AI](/blog/tag/artificial-intelligence)

[Engineering](/blog/tag/engineering)

#### Recommended Reading

[Read ](/blog/ai-context-configuration)

[Product](/blog/ai-context-configuration)

###### [Your agent isn’t broken. It’s guessing.](/blog/ai-context-configuration)

[Sep 23, 2026](/blog/ai-context-configuration)

[6 min read](/blog/ai-context-configuration)

[Read ](/blog/headless-amplitude)

[Company](/blog/headless-amplitude)

###### [Bring headless Amplitude where you work](/blog/headless-amplitude)

[Sep 16, 2026](/blog/headless-amplitude)

[3 min read](/blog/headless-amplitude)

[Read ](/blog/amplitude-sdk-or-not)

[Insights](/blog/amplitude-sdk-or-not)

###### [Should I install the Amplitude SDK?](/blog/amplitude-sdk-or-not)

[Sep 15, 2026](/blog/amplitude-sdk-or-not)

[4 min read](/blog/amplitude-sdk-or-not)

[Read ](/blog/g2-fall-2026-digital-analytics)

[Customers](/blog/g2-fall-2026-digital-analytics)

###### [Amplitude climbs to #3 in G2 Digital Analytics Momentum](/blog/g2-fall-2026-digital-analytics)

[Sep 15, 2026](/blog/g2-fall-2026-digital-analytics)

[10 min read](/blog/g2-fall-2026-digital-analytics)
