Jev: Independent Benchmarks and Real Costs
· 13 min read
TypeSafe AI launched Jev with an arresting promise: typed decisions for software at a fraction of the latency and cost of a large language model. The homepage still advertises very large speed and cost multipliers. Independent tests published since launch make the picture more useful. Jev can indeed be unusually cheap and quick, but the largest multipliers are vendor benchmark results, not a general exchange rate between Jev and every LLM.
The practical question is narrower: when does a specialized decision model improve a real system, and what does it cost after the surrounding engineering is included?
Sources and product documentation in this article were checked on 6 October 2026. Every figure it relies on sits in the table below, with its source and reading date.
What Jev actually is
Jev is not a chatbot. Its interface takes some state and one or more predefined questions, then returns typed answers. A question may ask the model to choose among named options, score an item against ordered levels, or estimate a yes/no probability. The documented primitives make it a candidate for routing support tickets, assigning a risk band, scoring a lead, or choosing the next branch in a workflow. It cannot write the customer reply that follows, and the vendor caps how many options a single Choice can carry (see the table).
TypeSafe's launch post, listed in the table's sources, describes Jev as outputting all of its probabilities in parallel instead of generating a string token by token, with a parallel sampler built for efficiency. The company also says it trained the model with a method called Reinforcement Learning for Calibrated Decisions, or RLCD. Those are the vendor's architectural and training descriptions; there is not yet enough public technical material to audit the mechanism independently. The observable product contract matters more today: constrained questions go in, schema-valid decisions and probability fields come out.
Imagine a returns system with three permitted routes: refund, replacement, and manual_review. An illustrative Jev response might put most of the probability on refund, a little on replacement, and the remainder on manual_review, then select refund. That response is invented to explain the interface. It is neither a measured result nor evidence that a high probability means the decision will be correct that often.
The distinction is central. The native decision contract prevents an unexpected fourth label or another value outside the declared schema. It does not make the selected label true, and TypeSafe's own System One overview notes that calibration is measured across groups of predictions and does not guarantee that an individual answer is correct. A schema-valid wrong answer remains wrong; transport serialization is a separate layer.
The launch claims need their denominator
The same launch post is more nuanced than the homepage headline. The speed and cost multipliers come from the company's own workflow evaluations. Those workflows compare Jev with large models wrapped to emit compatible structured decisions and probabilities. The reference answers are averages from powerful LLMs, rather than independently established ground truth. TypeSafe also acknowledges that its model capabilities team created the workflows and that the reported gains are likely at the high end of real deployments.
That setup is useful for testing the product TypeSafe wants to build, but it cannot establish universal superiority. It favors tasks shaped like Jev's interface, evaluates agreement with LLM probabilities, and compares against systems doing more expensive generative work. The company's published response-time range is similarly a per-call latency claim whose published evals the vendor says are generally run from the West Coast, where its service is based. It is not a throughput guarantee.
Two independent evaluations now provide better anchors. AY Automate ran a labeled study in September 2026 in which several LLMs received the same instructions and options through OpenRouter chat completions with strict JSON schemas, while Jev ran through OpenRouter's decisions endpoint. On eight-way routing, Jev was not statistically distinguishable from Gemini 3.5 Flash-Lite or Claude Haiku 4.5, but it was significantly behind GPT-5.4 nano and GPT-5.6 Terra.
A preregistered evaluation published on GitHub returned "ambiguous" on both of its experiments. On one zero-shot dataset Jev beat the nano model and trailed the frontier model, and its recorded median call was roughly half as long as the nano model's. Those timings describe the tested client and service paths, not model inference generally.
These studies are encouraging precisely because they are less spectacular. They suggest a specialized, low-cost router worth testing. They do not show production reliability, sustained throughput, or accuracy on a private domain.
| Measure | Value | Sources |
|---|---|---|
| Jev list price, input | $0.042 per million input tokens ($42 per billion) | [1] [2] |
| Jev list price, output | Output tokens are not charged | [1] [2] |
| Jev 1.13 default rate limits | 100K tokens per second / 80 requests per second (documented as dynamic, can change without notice) | [2] |
| Maximum options in one Choice question (vendor) | 255 | [1] |
| Vendor claim from its own workflow evaluations | 193.6x faster, 444.6x cheaper than LLM workflows | [1] |
| Vendor claim: end-to-end response time | 70 ms to 500 ms; the vendor's published evals are generally run from its laptops on the West Coast, where the service is based | [1] |
| AY Automate: labeled decisions in the study | 791, run on 19 September 2026, Jev and four LLMs through OpenRouter | [3] |
| AY Automate, eight-way routing (160 items): median call latency | Jev 1.13 0.327 s; Gemini 3.5 Flash-Lite 0.667 s; Claude Haiku 4.5 0.992 s | [4] |
| AY Automate, eight-way routing: measured bill per 1,000 decisions | Jev 1.13 $0.0151; Gemini 3.5 Flash-Lite $0.0869; Claude Haiku 4.5 $0.3565 | [4] |
| Computed: that measured bill extrapolated to 1 million decisions (× 1,000) | Jev 1.13 $15.11; Gemini 3.5 Flash-Lite $86.94; Claude Haiku 4.5 $356.47 | [4] |
| Computed: measured bill relative to Jev | Gemini 3.5 Flash-Lite 5.8x; Claude Haiku 4.5 23.6x | [4] |
| Computed: Jev API cost for 1 million decisions at list price, for two assumed input sizes | 200 input tokens each: $8.40; 1,000 input tokens each: $42.00 | [1] [2] |
| Computed: minimum time for 1 million one-decision calls at the default request limit | 12,500 s (3 h 28 min 20 s), before errors and retries | [2] |
| Preregistered evaluation: median call duration, same 30 CLINC150 items | Jev through Vercel 0.423 s; gpt-5.4-nano through Lightning AI 0.924 s | [5] |
| Preregistered evaluation: Banking77 accuracy (208 items) | Local bge-small-en-v1.5 + logistic regression trained on 10,003 labeled examples 0.933; Jev zero-shot 0.832 | [5] |
| Preregistered evaluation: CLINC150 confidence (200 items) | Jev returned confidence exactly 1.0 on 102 items, 6 of them wrong | [5] |
Vendor rows come from TypeSafe's own workflow evaluations: the vendor says its model capabilities team built the workflows, that some bias could exist, and that these gains are on the higher end of real-world gains.
Computed rows are arithmetic on sourced values, not measurements: a per-1,000 bill multiplied by 1,000, input tokens multiplied by the list price, or one million calls divided by the request limit, assuming one decision per call, no failures and no retries. They exclude taxes, gateway markups, logging, escalation and engineering time.
The AY Automate rows describe one run per system on public data through OpenRouter. The authors report 95 percent intervals 6 to 13 points wide and say they have not used Jev on client projects.
The preregistered evaluation compares client-and-service configurations, not models in isolation. Its supervised baseline used labels that the zero-shot systems never saw, and its Banking77 subset was cut from 300 to 208 items because of Jev's rate limit, without proof that the retained items are representative.
The input sizes in the API cost row (200 and 1,000 tokens per decision) are assumptions chosen to bracket short and long inputs, not measured sizes.
The minimum-time row holds only while the token limit does not bind: at the default limits, that means calls averaging no more than 1,250 tokens each (100,000 tokens per second divided by 80 requests per second). Longer calls make the token limit the binding one and the floor longer.
TypeSafe documents its rate limits as dynamic and offers higher limits on custom and enterprise plans: the default limits can change between two readings.
Sources
- Introducing System One models and Jev, TypeSafe AI, read on 2026-10-06
- Models (Jev 1.13: price, rate limits), TypeSafe AI documentation, read on 2026-10-06
- Jev vs LLM benchmark, AY Automate, read on 2026-10-06
- Jev vs LLM benchmark, raw summary (summary.json, cohort intent8), AY Automate, read on 2026-10-06
- jev-baselines-eval, preregistered evaluation of Jev (README), ickma2311 (GitHub), read on 2026-10-06
What a million decisions cost
Jev's list price charges input tokens only, with no metered output charge. Both the state and the question definitions, including option text, count as input. Token counts must be measured with the relevant tokenizer; equal-looking prompts can produce different bills across providers.
The table gives two computed Jev bills for a million decisions, one with short inputs and one with longer inputs. The arithmetic is simply input tokens multiplied by the list price. It excludes taxes, retries, gateway markups, rejected requests, logging, and any escalation to another model or a human.
AY Automate's raw summary gives the measured OpenRouter bills for the same eight-way routing messages, and the table extrapolates each one to a million decisions. This is arithmetic, not a million-request load test. On those items Jev was the cheapest and the fastest of the five systems AY measured; of the three in the table, Gemini 3.5 Flash-Lite came second and Claude Haiku 4.5 third on both counts, while the accuracy gaps among those three were too uncertain to rank them reliably. Different prompts, reasoning, caching, providers, or tasks can change those ratios, and none of these bills includes integration or escalation.
Latency is not capacity
A fast individual response does not tell an operator how many responses the service accepts. A single call reveals a latency, not a rate the service will sustain: a system that must finish a large batch of independent decisions in seconds needs a high request rate that only the provider's quotas can grant, and Little's Law ties that rate to the number of requests in flight: at a given latency, a higher target rate means proportionally more concurrent requests.
TypeSafe's model documentation (see the table's sources) supplies the decisive default limit, and describes it as dynamic and subject to change without notice. The default limits have already changed since this article was first published. At the current default request limit, a million one-decision calls still take hours, even before errors and retries; the table gives the computed floor. TypeSafe invites customers to request higher enterprise limits, so this does not prove a large workload impossible. It proves that latency alone cannot establish capacity. Jev can also batch multiple questions about the same state, which may improve useful decisions per call, but that is different from proving a high rate of independent requests. Buyers should secure quotas and test tail latency under representative load.
Vercel made Jev available through AI Gateway, confirming a third-party access path. A gateway can simplify procurement and testing, but adds its own routing, limits, price, data path, and failure modes. Treat gateway latency as a property of the end-to-end route.
TypeSafe's current quickstart separately documents direct API keys and the /v1/systemone endpoint without stating a waitlist requirement.
The local baseline is different, not free
The preregistered evaluation also tested local bge-small-en-v1.5 embeddings plus logistic regression trained on the full labeled training split of Banking77. On a retained Banking77 subset it scored clearly above zero-shot Jev (see the table). This is supervised versus zero-shot, on a small subset whose author could not show it was representative. It shows what labels can buy, not universal superiority.
"Local costs nothing per API call" is true but incomplete. A local classifier trades a per-call bill for fixed costs: the instance that serves it, sized for peak traffic, plus annotation, training, validation, serving redundancy, monitoring, updates, and engineering time. The table does not price that deployment, because its cost depends on hardware and traffic that this article cannot measure for you.
A local classifier becomes attractive when the label set is stable, representative labeled examples exist, data must stay inside a controlled environment, and traffic makes fixed infrastructure efficient. Jev is attractive when a team needs zero-shot or lightly configured decisions, has changing natural-language state, and values a managed typed interface. An LLM remains preferable when the task needs generation, broad reasoning, tools, or explanations.
Confidence needs validation
Jev returns probabilities by design; Choice and Score also return a separate confidence field. That is operationally valuable because software can route low-confidence cases elsewhere. It is not an explanation of the decision, and it is not proof of calibration after a domain shift.
In the preregistered study, Jev assigned the maximum confidence to about half of the zero-shot items and still missed some of them (see the table). Cascade results changed with threshold choices, so the data did not establish a better escalation gate. AUROC ranks cases; it does not prove calibrated probabilities.
Teams should therefore build a held-out validation set from their own traffic, plot reliability by confidence bucket, choose thresholds before looking at the final test set, and monitor drift after launch. High-impact decisions also need rules outside the model: permissions, audit logs, reversible actions, and human review where errors carry material consequences. Workflows that must justify a decision may also need a separate evidence or explanation layer; a probability distribution alone does not satisfy that need.
A practical buying checklist
Before adopting Jev, run a shadow evaluation that preserves the real state, option descriptions, class imbalance, and error costs. Compare at least three systems: Jev, a small structured-output LLM, and a supervised or rule-based baseline when labeled data exists. Record end-to-end latency distributions, billed cost, invalid responses, retries, accuracy by class, calibration, and escalation volume.
Then ask the operational questions the public benchmarks cannot settle: sustained and burst quotas, regional processing guarantees, retention and training policies, service-level commitments, outage behavior, and price stability. Public documentation describes a proprietary hosted model; no public model weights or self-hosting package were found during this review. Publicly documented region guarantees were also not found. Those are gaps in available evidence, not proof that no contractual options exist. The concrete lock-in risk is replacement: a team can preserve its typed questions and evaluation set, but it cannot transfer Jev's weights to its own infrastructure or another provider.
Jev's strongest case is modest and consequential. It turns a common software need, choosing among known actions under uncertainty, into a first-class model interface. Independent evidence supports testing it as a fast, inexpensive decision component. The evidence does not support replacing every classifier or LLM, accepting confidence as truth, or planning capacity from a latency headline. The right benchmark is your decision, your distribution, and your full operating cost.