Highlights
Most teams use large language models (LLMs) for every AI task, even yes-or-no calls. Jev, launched by TypeSafe AI on September 15, 2026, is a System One model that returns typed, probabilistic decisions instead of text. But the useful question is not which one wins. It is how the two work together. This blog explains how Jev and LLMs differ, what each does best, and how each can make the other better: Jev can route, gate and check LLM work, while LLMs can plan, explain and handle the hard cases Jev passes on. It also covers a worked example, AI model routing and guardrails, alternatives, and how to test safely. The core idea: give each decision to the smallest reliable model, and generate only when generation is the job.
What if one of the more interesting AI models of 2026 is interesting precisely because it doesn’t want to generate a sentence?
We now have models capable of writing code, reviewing architecture, researching topics and reasoning through complex problems.
And then we use those same models to answer:
- Is this ticket urgent?
- Which department owns it?
- Is this message spam?
- Should this action require approval?
It is a little like hiring a senior consultant and spending half the day asking them to operate an if statement.
That is the problem Jev, TypeSafe AI’s first public System One Model, is trying to address.
TypeSafe launched Jev on September 15, 2026. Its mental model is simple:
unstructured state in → typed probabilistic decisions out.

Fig: AI Decision Workflow
Jev is not trying to replace GPT, Claude or Gemini. It targets a different part of software: the millions of small judgment calls that happen before, after and around generation. That “before, after and around” is exactly where the two kinds of model can work together, and it is the real subject of this blog.
So instead of asking “Jev or an LLM?”, ask a better question: where does each one belong in the same system? A large language model is your writer, coder and thinker. A decision model is your fast, cheap gatekeeper. Used together, each does the job it is best at, and the whole system gets faster, cheaper and safer.
Why AI Decision Making Is Trending Now
Before we look at how Jev works, it helps to see why this topic is getting so much attention. For the last few years, most teams reached for a large language model (LLM) whenever they needed AI. That worked well for writing, coding and research. However, as companies moved from demos to real products, they noticed something: a big share of their AI calls were not writing tasks at all. They were small judgment calls.
This is where AI decision-making models come in. Today, people talk about two kinds of AI work. Generative AI creates content, while decision AI chooses between options. In the same spirit, Nitor Infotech’s blog on Generative AI: From Prompt to Production explains that production AI needs guardrails, validation and measurement, not just a good prompt. AI decision models are one practical way to add those layers. Importantly, the two kinds of AI are partners, not rivals: one produces, the other steers and checks.
So what is driving the shift? Three simple reasons stand out:
- Cost: Using a very large model for a yes-or-no answer gets expensive when you make millions of calls.
- Speed: Live products and agent loops need fast answers, not long generated responses.
- Control: Teams want fixed options with probabilities, so normal code can decide what happens next.
To see the difference clearly, the diagram below compares a generative LLM with a decision model. Notice how the decision model starts with your data and ends with scores, not sentences.

Fig: Generative AI vs decision AI at a glance.
Jev and LLMs: what is the actual difference?
Before combining two tools, it helps to be clear about how they differ. The table below sets them side by side.
| Aspect | LLM (generative model) | Jev (System One model) |
|---|---|---|
| Core job | Generate language, code and reasoning | Evaluate decisions you define in advance |
| Thinking style | System 2: slower, deliberate, step by step | System 1: fast, intuitive judgment |
| Output | Tokens written one after another; a schema can constrain the format | Typed probabilities over your options; nothing outside the schema |
| Speed | Seconds, or longer for deep reasoning | Vendor-reported 70 to 500 ms end to end |
| Cost | Input and output tokens are both metered | Vendor-listed $0.042 per million input tokens; output not charged |
| Uncertainty | A stated “92% confident” may just be generated text | Trained for calibrated probabilities (RLCD); still verify on your data |
| Can it explain itself? | Yes, in natural language | No. It returns numbers, not reasons |
| Best at | Explanation, code, summaries, research, conversation, planning | Classification, scoring, routing, validation, real-time loops |
| Weak at | High-volume small decisions: slow and costly | Anything that needs written output or open-ended reasoning |
Read the table as a division of labour, not a scoreboard. There is also an accuracy point worth knowing. On TypeSafe’s own four-workflow benchmark, as reported by DataCamp, Jev scored about 67.8%. That is roughly level with one GPT-5.6 tier and a few points behind the strongest models tested (74.1% and 73.1%), at a small fraction of the cost and latency. These are vendor-reported figures on vendor-chosen tasks, so test on your own data. But they explain the sensible design: let Jev handle the volume, and keep a stronger LLM for the hard cases.
What does Jev actually do?
Suppose a customer writes:
“The export button crashes in Safari. Chrome works. Finance needs the report in 20 minutes.”
A generative model might summarize the problem, explain its reasoning and return JSON.
Jev takes a narrower approach.
You define the exact decisions you want it to make.
| Primitive | What you ask | Example |
|---|---|---|
| Noul | Is this true? | Is this customer-impacting? |
| Choice | Which option best fits? | Bug / Billing / Feature |
| Score | Where does this sit on an ordered scale? | Severity 0–3 |
You might receive:
category bug 0.96 billing 0.01 feature 0.03 severity low 0.04 medium 0.79 high 0.17
Then normal software takes over:
if p_bug > 0.90 and severity > threshold: escalate() elif confidence < review_threshold: human_review()
That separation matters:
the model makes a fuzzy judgment; your software decides what happens next.
TypeSafe says multiple questions over the same application state can be evaluated together and in parallel rather than being generated as a sequential response.
Isn’t this just structured output from an LLM?
This is probably the most important question.
Modern LLMs already support schemas:
class Decision: team: Literal["payments", "risk", "support"] priority: int
OpenAI Structured Outputs and similar systems can constrain a generated result to a developer-defined schema.
That solves the format problem.
But the underlying job is still generative.
A conventional autoregressive LLM broadly works like this:
context ↓ token ↓ next token ↓ next token ↓ structured response
Jev instead exposes a bounded decision interface:
application state ↓ allowed decisions ↓ payments 0.82 risk 0.13 support 0.05
So I think of the distinction this way:
Structured-output LLM: Generate an answer, but stay inside this schema.
Jev: Evaluate these predefined decisions.
That does not automatically make Jev smarter. It makes the model specialized for a different workload.
TypeSafe currently prices Jev at $0.042 per million input tokens, with outputs not separately charged, and reports end-to-end latency commonly in roughly the 70–500 ms range. Those are vendor-published figures, not guarantees for every workload. (Check the content and okay we can put) TypeSafe has also said its early pricing may be subsidized, so confirm the live rate before you plan a budget.
The more interesting part: uncertainty
Imagine an LLM responds:
Risk: HIGH Confidence: 92%
That 92% may itself simply be generated text.
Probability calibration is a different problem.
A well-calibrated model should behave roughly like this:
predictions given ~80% probability ↓ correct roughly 80% of the time across a sufficiently large comparable set
TypeSafe says Jev is trained using Reinforcement Learning for Calibrated Decisions (RLCD), where useful probability estimates are part of the objective.
But this needs an important warning:
0.97 probability does not mean truth.
Your application still decides how much confidence is enough:
> 0.95 automate 0.70–0.95 verify < 0.70 review manually
Those thresholds are business and risk decisions.
Why the best results come from using both
Each model is strong exactly where the other is weak, which is why combining them beats choosing one.
- LLMs are strong at language, reasoning and open-ended problems. But calling one for every small decision is slow and costly, and its own confidence score is only text.
- Jev is fast and cheap, and it returns probabilities inside a fixed schema. But it cannot explain, draft or plan, and on vendor-reported tests it does not match the strongest LLMs on peak accuracy.
Put them together and the weaknesses largely cancel out. Jev absorbs the volume; the LLM absorbs the difficulty. A simple rule of thumb: Jev decides what should happen next, and the LLM does the work that needs language.
A quick way to split the work is to ask four questions about each step:
- Is the answer exact (a role check, a limit, a date)? Use normal code.
- Is the answer one of a fixed set of options that needs judgment? Use a decision model like Jev.
- Does the step need writing, explaining, coding or open-ended reasoning? Use an LLM.
- Is the cost of being wrong high, or is confidence low? Add a human.
Where does Jev actually fit?
Rather than thinking about 15 separate use cases, it is easier to group Jev into five kinds of work.
1. Classify
This is the most natural fit.
input ↓ bug / billing / fraud / spam / other
This includes:
- Classification
- Intent detection
- Document classification
- Anomaly classification
- Moderation
Support triage is an obvious example, but the same pattern can classify documents, content, events or records.
2. Score

Sometimes we do not need a category. We need a judgment on a scale.
low → medium → high → critical
That can support:
- Severity scoring
- Risk detection
- Prioritization
- Transaction screening
- Customer frustration
- Code risk
- Semantic feature extraction
For example:
transaction ↓ unusual behaviour? 0.84 possible fraud? 0.71 needs review? 0.88
I would not use Jev as a standalone fraud system. A real fraud architecture would normally combine deterministic rules, historical features, specialized models and human investigation.
Jev becomes one semantic signal in that system.
It can also turn fuzzy language into numeric or categorical features:
customer message ↓ urgency 0.82 frustration 0.67 purchase intent 0.91
Those features could feed analytics, ranking, rules or another ML model.
3. Route
Once you can classify and score, routing becomes very useful.
request ↓ decision model ├── simple → retrieval / small model ├── coding → coding model ├── complex → frontier model └── uncertain → human
This applies to:
- model routing
- Workflow routing
- Agent routing
- Escalation
- Human-review decisions
Instead of making Jev do the expensive work, use it to decide which resource should do the expensive work.
4. Validate
This may be one of the most important enterprise patterns.
Suppose an agent proposes:
DELETE FROM customer_records;
A decision layer could evaluate:
destructive? 0.99 sensitive operation? 0.94 requires approval? 0.98
That covers:
- Guardrails
- Agent action validation
- Policy interpretation
- LLM output evaluation
- Semantic risk checks
But the distinction matters:
Jev should support policy enforcement, not replace it.
If the rule is exact:
if user_role != "admin": deny()
keep it in code.
Use a decision model where the question itself is ambiguous:
Does this requested action appear destructive?
The same pattern can wrap an LLM:
LLM response ↓ Jev ↓ relevant? 0.94 policy-safe? 0.97 needs review? 0.12
That is better described as LLM output evaluation, not guaranteed fact-checking.
5. Decide in a loop
Because Jev is designed for bounded decisions rather than long-form generation, it can potentially participate in latency-sensitive loops:
current state ↓ decision ↓ action ↓ new state
This covers:
- Real-time decisions
- Browser actions
- Tool selection
- Games
- High-volume event processing
A browser agent, for example, could use an LLM to understand the high-level goal while Jev repeatedly chooses from allowed actions:
Goal ↓ LLM planner ↓ browser state ↓ Jev ↓ click / select / wait / stop
One public Browser Use experiment reported a flight-search workflow completing in about seven seconds, although this should be treated as a reported experiment rather than a general benchmark. (madewithjev.com)
The same principle applies to high-volume systems:
millions of events ↓ classify score flag route ↓ only difficult cases go downstream
This could apply to log triage, email classification, listing moderation, stream enrichment or event prioritization.
How Jev makes your LLM better
Jev works best as a layer around an LLM. Here are the five most useful ways it can improve the LLM you already run.
1. Before the LLM: route and triage
Let Jev read each request first and choose the path: a cheap retrieval step or small model for simple questions, a coding model for code, a frontier model for hard problems, and a person for uncertain cases. The expensive model then only sees the work that needs it. This is usually the easiest first project.
2. Before the LLM: decide what context it needs
Answer quality depends on what goes into the prompt. Jev can quickly decide whether a question needs retrieval, which knowledge source to search, or whether it is out of scope entirely. Sending the LLM less irrelevant material means better answers and lower cost. Our blog on Context Engineering for Agentic LLM Systems explores this in depth.
3. After the LLM: check the output
Before a drafted answer reaches a customer, Jev can score it: is it relevant, is it policy-safe, does it need review? Think of this as a fast second opinion, not a guarantee of factual accuracy. Anything below your threshold goes to a person.
4. Around an agent: gate its actions
When an LLM agent proposes an action, Jev can judge whether it looks destructive, touches sensitive data or needs approval. Exact rules such as role checks still belong in code; Jev handles the ambiguous “does this look risky?” questions.
5. In place of the LLM: handle the small, high-volume calls
Tagging, spam checks, urgency scoring and similar tasks do not need a writing model. Moving them to a decision model frees your LLM budget for the tasks where generation really matters.
How your LLM makes Jev better
The help flows the other way too. Jev cannot write, plan or explain, and that is exactly where an LLM steps in.
1. The LLM plans; Jev acts
In an agent loop, an LLM can understand the goal and make the plan once. Jev then picks the next allowed action at each step (click, select, wait, stop) at a speed and cost that a generative model cannot match in a tight loop.
2. The LLM turns messy input into clean state
Long email threads, transcripts and documents can be condensed by an LLM into short, structured state. Jev is priced on input, so a concise, clear state is both cheaper to judge and easier to judge well.
3. The LLM handles what Jev is unsure about
When Jev’s probabilities fall in the uncertain middle, route the case to a stronger LLM, or to a person with an LLM-prepared summary. Jev keeps the easy 80% cheap, and the LLM spends its effort on the hard 20%.
4. The LLM explains the decision
Jev returns numbers, not reasons. When a customer, auditor or teammate needs a plain-language explanation, an LLM can write one from the decision and the input. One caution: that explanation is written after the fact. It is not a view into how Jev actually decided, so label it accordingly.
5. The LLM helps you build and test the decision layer
An LLM can help you draft clear, literal decision criteria, generate realistic test cases, and review a sample of cases where Jev and your current system disagree. It can speed up the work, but people should still label the ground truth you measure against.
A worked example: a support ticket, end to end
Here is how the pieces fit in one flow. Jev and the LLM alternate, and your own code holds the thresholds.
Customer ticket ↓ LLM (if the thread is long): condense into a short state ↓ Jev: category? severity? needs a human? sensitive data? ↓ Your code applies thresholds ├── spam / low severity → template reply ├── routine bug → small LLM drafts a reply ├── high severity, high confidence → escalate; LLM writes the engineer summary └── low confidence → human review (LLM prepares context) ↓ Jev: is the drafted reply on-policy? needs review? ↓ Send
Notice that the LLM never decides who gets what, and Jev never writes a word. Each does one job, and your software stays in charge.
How AI Decision Intelligence Supports Agentic AI
Now that you have seen the five kinds of work, let us connect them to one of the biggest trends in enterprise AI: agentic AI. An agent plans a task, picks tools, takes an action and checks the result. Nitor Infotech explains this bigger picture in Agentic AI and the Next Generation of AI Assistants.
In such systems, small choices happen at every step. Which tool should the agent call? Is this action safe? Should a person approve it first? These are perfect jobs for AI decision intelligence. A large model can plan the goal, while a fast decision model picks from the allowed actions. This is the LLM-plus-Jev partnership in its most natural form.
Of course, a decision is only as good as the information it sees. For this reason, context matters a lot. Our blog on Context Engineering for Agentic LLM Systems shows how to manage instructions, memory and tools so that models get the right input at the right time.
The same thinking works outside chat apps too. For example, in DevOps, an agent must judge whether a release looks risky before it deploys. You can read how this works in Agentic AI in DevOps: From CI/CD to CA/CD.
AI model routing: an easy place to start
To make this practical, let us look at AI model routing, which is often the simplest first project. Instead of sending every request to the biggest model, a small decision model first reads the request and chooses the best path.

Fig: A decision model can route each request to the right resource.
As the diagram shows, simple questions go to retrieval or a small model, coding tasks go to a coding model, and hard problems go to a frontier LLM. Meanwhile, uncertain cases go to a person, which protects quality. Here is what teams usually gain:
- Lower cost: the expensive model is used only when needed.
- Faster answers: simple requests skip heavy processing.
- Better safety: unclear cases are sent for human review.
Where should you not use Jev on its own?
Jev is a poor fit when the output itself is the product.
Use a generative LLM when you need:
- Explanation
- Code generation
- Summarization
- Research
- Creative writing
- Open-ended reasoning
- Conversation
Use normal code when the answer is deterministic:
amount > limit date < expiry_date role == "admin"
Use Jev-like models between those two:
the answer requires judgment, but not generation.
That is probably the cleanest mental model. And notice that everything in the first list is work for the LLM you already have. Jev does not compete for those jobs. It clears the small decisions out of the way so the LLM can focus on them.
What are the alternatives?
Jev is not the only solution.
Structured-output LLMs
If your task genuinely needs reasoning and you simply need predictable JSON at the end, keep using the LLM and constrain the schema.
Traditional classifiers
BERT-style classifiers remain excellent when categories are stable and you have enough labelled data.
Small or distilled LLMs
For routing and classification, a cheap small model may provide enough semantic understanding without introducing another provider.
Rerankers
Search systems have used specialized non-generative models for years. Their job is not to write an answer; it is to score relevance.
Open implementations
The community is already experimenting with Jev-like patterns.
OpenDecision implements Choice, Noul and Score using a ModernBERT-based model, while explicitly noting that its probabilities should currently be considered uncalibrated. (github.com)
That matters because the architectural idea may eventually become more important than the vendor. The same pairing works with any of these: a decision layer in front of and behind an LLM.
Open-Source Alternatives
| Project | Architecture | Size | Jev-style decisions | Best suited for |
|---|---|---|---|---|
| Laya | ModernBERT/mmBERT + decision head | 322M–421M | Yes | Small multilingual decisions |
| Nimble | Qwen3.5 + LoRA candidate scoring | 9B | Similar | Training and research |
| Kev | Qwen + LoRA + pointer head | 0.8B–9B | Yes | Jev-like local deployment |
| SemIf | Direct logits from open LLMs | Model-dependent | Similar | Direct-logit experiments |
| Rizzo Flow | Spark + direct option scoring | 1.7B–4B | Yes | Easy local serving |
| Von | ModernBERT + OptionMarker | ~395M | Yes | Compact decision inference |
| NanoJev | Qwen3 + shared decision heads | 0.6B | Yes | Small trainable Jev replica |
Where AI Classification Models Help in Real Businesses
So far, we have talked mostly about ideas and tools. Next, let us see how the same ideas can look in real business settings. The examples below are common patterns, not customer results.
- Software product companies (ISVs): tag and route support tickets, spot risky user actions, and score feature requests.
- Healthcare technology: sort incoming documents and flag records for review. Because this is a sensitive area, clinical validation, privacy checks and human oversight stay essential.
- Retail and supply chain: score customer intent, prioritize order exceptions and flag unusual activity.
- Sales teams: score lead quality and route leads to the right owner. For a wider view, see How AI Agents Are Transforming Traditional Sales Workflows.
- Data engineering: rank pipeline alerts by severity so that only serious ones reach engineers. Our blog on Agentic Data Engineering and Self-Healing Pipelines covers this direction.
In these examples, the decision model does not write anything itself. It turns messy input into a clear signal that normal software can use, and where text is needed, an LLM takes over from that signal. That is why AI decision intelligence often works quietly in the background while the visible product feels smooth and fast.

Engineers GenAI-based solutions and GenAI-enabled software products for ISVs and platform companies. Download the factsheet to plan your next AI step.
Is Jev genuinely new, or just a better-packaged classifier?
My answer today is: there is something interesting here, but we do not know enough yet to treat every claim as settled.
TypeSafe says Jev uses a new architecture, parallel sampling and RLCD training.
The company has not publicly disclosed enough internal architectural detail for outsiders to independently understand every part of its implementation.
So from an application developer’s perspective, it is still largely a black box.
What has improved is the contract around the black box.
If I declare:
payments risk support
Jev cannot suddenly answer:
“send it to Bob”
But it can still choose payments when risk was correct.
Which gives us two useful rules:
Type safety is not truth safety.
and:
Calibration is not correctness.
How to Start with Jev and LLMs Together: A Simple 6-Step Plan
Given these limits, the next question is where to begin. The good news is that you do not need a big rebuild. A small, careful pilot is enough to learn a lot.
- List your decisions. Review your current LLM calls and mark the ones that are only classification, scoring, routing or checking.
- Separate rules from judgment. Keep exact rules, such as role checks, in code. Use a decision model only when the question is fuzzy.
- Design the hand-offs. Decide what the decision model passes to the LLM, what the LLM passes back, and where a person steps in. Start with one flow, such as route first, then generate.
- Run side by side. Let the decision model work in shadow mode next to your current system for a few weeks.
- Measure everything. Track accuracy, calibration, latency, cost, false positives, false negatives and human escalation rate, for the combined system as well as each model on its own.
- Set thresholds and keep people in the loop. Choose confidence levels for automate, verify and review based on business risk.
Step five deserves extra attention, so here is a simple picture of how confidence thresholds can work in practice.
Please treat these numbers as examples only. Your own testing and risk level should set the real values. Along the same lines, Nitor Infotech’s blog What is Multimodal Generative AI? recommends least-privilege access and human approval for high-impact actions, and the same care applies to decision models.
After the pilot, it helps to see the bigger picture. A healthy production stack does not depend on one model for everything. Instead, each layer does the job it is best at.
Common mistakes to avoid
Finally, a few mistakes show up again and again, so watch out for these:
- Treating a high probability as proof that the answer is true.
- Using a decision model for exact rules that normal code can handle.
- Skipping calibration checks before automating decisions.
- Replacing the whole architecture at once instead of testing one decision first.
- Asking the LLM to redo a decision Jev has already made, which throws away the cost and speed savings.
- Treating an LLM’s after-the-fact explanation of a Jev decision as the model’s real reasoning.
What matters more than Jev: the whole stack
For the last few years, many AI architectures have effectively been:
everything ↓ giant LLM ↓ everything
I think production AI systems are moving toward:
rules + classifiers + decision models + retrieval + generative LLMs + human judgment
Each component does less.
The system does more.
For developers, that means asking:
Do I need generation here, or do I only need a decision?
For companies, the question is economic as much as technical:
How much of our “LLM workload” is really classification, routing, scoring or verification wearing a chat-completion costume?
I would not rebuild an architecture around Jev tomorrow.
I would take one narrow decision, run Jev—or an alternative—beside the existing system, and measure:
- Accuracy
- Calibration
- Latency
- Cost
- False positives
- False negatives
- Human escalation rate
Then automate based on evidence. Once the decision layer has proved itself, wire it in front of and behind your LLM: route before, validate after, and hand the hard cases back.
The bigger idea is not:
Jev replaces LLMs.
It is also not that LLMs make Jev unnecessary. It is:
use the smallest reliable intelligence for the decision in front of you, and generate only when generation is actually the job.
Key Takeaways
To wrap up, here are the main points to remember:
- Jev and LLMs are partners. Jev is a System One model that returns typed, probabilistic decisions. LLMs generate language, code and reasoning. The best systems use both.
- Jev makes LLMs better by routing requests, choosing context, checking outputs and gating agent actions.
- LLMs make Jev better by planning, cleaning up messy input, handling uncertain cases, explaining decisions and helping you build tests.
- Split the work by question type. Exact answers go to code, fixed-option judgments to Jev, writing and reasoning to the LLM, and high-risk or low-confidence cases to people.
- AI model routing is an easy first project. A small decision model can send each request to the right resource and save cost and time.
- Your software stays in charge. The model makes a fuzzy judgment, and normal code decides what happens next.
- Probability is not truth. Type safety is not truth safety, and calibration is not correctness.
- Keep exact rules in code. Use decision models only where the question is ambiguous.
- Test before you trust. Run a pilot side by side, measure accuracy, calibration, latency, cost and errors, then set thresholds with humans in the loop.
Ready to Bring AI Decision Making into Your Product?
Talk to Nitor Infotech, an Ascendion company. Our teams help ISVs and enterprises design, test and scale AI solutions, from LLM applications to decision layers that keep your systems fast, safe and cost-effective.
Contact us at Nitor Infotech →