What is Jev and how to apply it: TypeSafe's judgment model

In a brick office with a large industrial black-framed window on the left and a rubber plant in a terracotta pot beside it, a woman with brown hair tied up, wearing a beige knitted cardigan over a white t-shirt, sits alone at a solid wood table. On the table are three identical piles of cream card tokens, and she holds a loose token in her right hand, hovering above the piles, looking at it with a focused expression. In the background, a wooden bookshelf with books.
Escrito por
Eduardo Garcia Garzón Seguir

Head of AI en Shakers. Diseña e implementa sistemas LLM para matching de talento a escala en entornos enterprise. Especialista en IA en producción, evaluación de modelos y aplicaciones de alto riesgo, con experiencia liderando programas de cumplimiento del AI Act.

• ...

Jev is TypeSafe's judgment model, what the company calls its System One model: an AI that does not generate text but answers closed questions and returns probabilities. It exists to settle the question that shows up in almost every AI project that stalls and that almost nobody says out loud: who decides. Not who writes the prompt, nor who picks the model. Who decides whether this ticket escalates, whether this invoice goes through, whether this request moves on.

For three years the default answer has been to ask a generative model to return a JSON with the decision inside and to trust that the format holds. It usually holds. What does not hold is the audit.

Jev is the opposite proposition, and the name is not decorative: System One refers to fast thinking, the answer a knowledgeable person gives in a second without needing to deliberate. Instead of prose it returns a probability distribution over the options you defined in advance. The current version is jev-1.13 and the primitives documentation is the source for everything that follows.

Jev is TypeSafe's judgment model. It does not generate text and it does not reason in steps: it reads a state, answers in parallel every closed question you send it and returns a probability distribution over the options you defined. Code keeps the control flow, the arithmetic and the policy. The model only contributes the judgment.

Three response types, and none more

That is the entire surface of the model. The brevity is deliberate.

  • Choice: the answer is one of a known set with no internal ordering, such as a ticket's category. Code acts with one branch per option.
  • Score: the answer is a position on a spectrum you can describe in levels, such as an incident's severity. Code acts with a threshold, a rank or a weight.
  • Noul: the answer is a clean yes or no, and what matters is the probability, not the label. Code acts with an if.

All three always return something your code consumes without having to interpret prose, which is the property that genuinely matters here, because it removes in one stroke the entire defensive layer of parsing, retries and format validation that accompanies any serious integration with a generative model. A Choice returns the chosen option, the probabilities for all of them and a confidence. A Score returns the weighted mean of the levels. A Noul returns a number between zero and one.

The design rule that orders everything else fits in one sentence: a good question for Jev is one a knowledgeable person would answer in a second if they had the right context in front of them. Does this message convey urgency? Fine. Analyse this message and decide what to do, no. The second is not a question. It is a brief, and it has to be split into small questions and reassembled in code.

The difference from asking an LLM for the decision

It lies in where the policy lives: it sounds like an architecture nuance, and it is what decides whether the system can be audited six months later.

When you ask a generative model to classify and decide, the business rule ends up written inside the prompt, mixed with the instruction, the output format and the exceptions that were added over time. Changing a threshold then means rewriting a paragraph in prose and re-evaluating everything from scratch, with no guarantee that the change has not quietly shifted the behaviour of three other cases nobody was watching, and the upshot is that nobody can say for certain what the system does, because the system is a text.

Jev's split is explicit: the model contributes the judgment, code keeps the control flow, the arithmetic and the policy. The weights and the thresholds live in the code. The documentation insists on a point worth underlining because it contradicts almost everybody's instinct: when the final decision is wrong but every individual answer is right, what is wrong is the policy, so you change the weights and leave the questions alone.

There is a second contrast, less philosophical and more operational. Jev does not count. Nor does it sum, or compare dates. The documentation says so without roundabout language and offers the replacement pattern: to count items you fire one Noul per item in the same request and sum in code the answers that clear the threshold. Anyone who has watched a generative model fail a seven-line count will recognise the problem at once.

Four patterns that cover almost everything

The first is the speculative fan-out. Every question that shares the same state travels in a single request, including the ones that only matter on some branches. They are answered in parallel, so asking extra adds very little latency and very few tokens to the request's total cost, and the code simply ignores the answers it does not need on that particular branch. A second request is only justified when the first request's answer determines which data has to be fetched.

The second is confidence-gated routing, and it is the one that most resembles directing a team. The answer says what. Confidence says whether to act.

ConfidenceWhat the system does
Below the floor (0.5 to 0.6)Runs nothing
Above the floorActs, or flags for confirmation
High-risk action (0.85 to 0.9)Only acts with high confidence; otherwise hands off to a person

Those are the values the confidence documentation uses as a starting point, with the explicit instruction to start conservative and tune them against your own data.

The third is composite scoring. A complex judgment, the kind settled in a meeting by saying it depends, is split into one Score per dimension that genuinely weighs on the decision, each is normalised by its number of levels, and they are then combined with the weights you set in code. When the business priorities change, you touch a weight. You do not rewrite a question.

The fourth is the taxonomy walk: one Choice per level of the tree, traversed in code, giving each option as its criteria what hangs beneath it so the model sees what sits under each branch before choosing it, and following more than one branch at a time when the probabilities come out close.

Where that state comes from is a separate decision, and some teams settle it by putting Model Context Protocol in front, so that the context arrives pre-filtered from the systems that hold it.

It is also worth saying what none of the four fixes: if accuracy falls as the input data grows, the problem is that the state carries irrelevant detail and it has to be filtered earlier in code. If a reworded question swaps one error for another, that question is weighing two properties at once and has to be split. And the state is not treated as hostile by default: the text you send can steer the answer, so adversarial cases are tested before deploying.

What changes for tech talent working on projects

A new capability shows up, and it is not knowing how to use a tool: that distinction separates those who bill by the hour from those who bill by judgment.

Designing a Jev program is five steps, and none of them is solved by writing longer prompts:

  1. List the decisions the system has to make, each one as a branch, a threshold or a ranking.
  2. Write one atomic question per judgment, and split any question that weighs two properties.
  3. Pick the primitive according to what the code will do with the answer.
  4. Build the smallest state that answers all of them, computing in code whatever code can compute.
  5. Compose the answers in code: branches, weights and confidence gates.

That is decision analysis and system design, an engineering muscle rather than a writing one, and it continues directly from what building an AI agent from end to end already demands: the difference is that here the judgment part stops hiding inside a prompt and becomes a component you can show, measure and defend.

The most underestimated part is evaluation. The documentation is categorical: you revise one or two questions per iteration, never more, because the probabilities move in ways that are hard to anticipate, and no revision counts as good without labelled data behind it. Higher confidence, on its own, does not prove the question improved. Whoever knows how to build that test bench, pick the cases that genuinely discriminate and defend the results in front of a committee that asks about risk before technology has a commercial argument that almost nobody can yet support with evidence on the table.

The UK market offers a hint of where this is heading, although it does not name Jev yet. According to Shakers' market analysis, 954 of 42,396 UK tech postings with tagged skills name LangChain, against 17,184 that name Python (n=48,596 active postings, 42,396 with tagged skills, June to October 2026). Model orchestration already has declared demand and a name of its own inside the wording of the postings, which is the first place a new capability becomes visible before anyone gives it a formal label in a role catalogue. The judgment layer inside that orchestration is the piece that still has no label.

That figure measures mentions in the wording of the postings, not requirements checked one by one. And the honest denominator is the postings with tagged skills, not the whole corpus, because the rest has no skills assigned and counting them would inflate the percentage.

What changes for a company

The board question stops being which model is best and becomes a different one: which decisions are we delegating, and with what threshold. That is governance, not procurement, and it is the conversation missing from most deployments of AI agents in companies.

Three practical consequences. The policy becomes readable, because the thresholds sit in the code and are versioned like any other engineering decision, so you can answer precisely, with the history in front of you, what was automated, under what condition and from what exact date. A designed, not improvised, exit path appears, because confidence-gated routing forces you to define from the start what happens when the model is unsure, which is precisely the scenario demonstrations never show. And the cost of changing your mind falls: if tomorrow the business decides that severity weighs more than urgency, that is a weight on one line of code.

It is worth placing this against the background problem. BCG's study of more than 1,250 companies sizes the gap bluntly: only 5% obtain value at scale and 60% get no material value despite substantial investment (BCG, September 2025). A good part of that implementation gap is not a model problem. It is that nobody defined which decision was being automated, nor with what minimum confidence, nor who answers when the answer is doubtful.

A judgment model does not solve that on its own. It forces you to write it down, which is considerably more than most tools achieve, and the pattern repeats in any system that reaches real production: the same is true of taking a RAG architecture to production, where retrieval is rarely the bottleneck and what fails is not having decided what happens with what was retrieved.

The limits, before you discover them in production

  • It does not generate text. If you need a written answer, that is another model. Jev sits in front to decide which one, or behind to verify what came out.
  • It does not reason in steps. A problem that requires chaining three inferences is split into three literal questions and joined in code.
  • It does not do arithmetic, counts or date comparison. For dates, the documented pattern is to extract each part with a Choice over enumerated options, including one for the case where it is not stated, and to assemble and compare in code.
  • Its scale is not a continuous magnitude. A Score of 1.0 can mean certainty at level 1 or a tie between levels 0 and 2, so you read the probabilities alongside the number and use the result to apply a threshold or to rank, never to infer a quantity.
  • Context has a cap. The state and every question share 64,000 tokens, and the state plus the longest question must fit in 32,000.

All of this is documented for jev-1.13, and the documentation itself warns that several of these limits will change in later versions, so the model jaggedness page gets re-read every time the version changes. Not just once at the start.

Where to start this week

Pick a decision that a person makes today in a repetitive way, many times a day and almost always the same, and whose criteria you can describe in a single sentence without resorting to a diagram or an exception that only the person with five years on the team knows. Write it down as a closed question. Choose the primitive according to what your code will do with the answer, which is the selection criterion, not the elegance of the format. Gather thirty already-solved cases with their correct answer, run them and look at the probabilities of the misses.

If the model fails with high confidence, almost always it read your instruction strictly literally, attending to every word you wrote and to none of the ones you took for granted, while you meant something slightly different that you never actually wrote. The explanation you would give when you see that error is, literally, the missing half of the instruction. Put it inside.

Frequently asked questions about Jev

Does Jev replace a generative LLM?

No. It covers a different function: deciding between options you define, not producing text. The usual arrangement is to combine them, with Jev classifying and routing in front and a generative model writing behind when writing is needed.

What is a Noul?

One of Jev's three response primitives. It returns a number between 0 and 1 for a condition that admits a clean yes or no, and the signal is the probability itself. A 0.5 means the model is not sure, not that the answer is intermediate. To measure a degree you use a Score.

Can a decision made with Jev be audited?

More than one made inside a prompt. The questions, each option's criteria, the thresholds and the weights are versioned code artefacts, and every answer carries its complete probability distribution.

How many questions is it sensible to send together?

Every question that shares the same state, in a single request, including the ones only used on some branches. They are answered in parallel, so asking extra adds little latency and few tokens.

Is there demand for this in the UK market?

For the orchestration layer, yes, and by name: 954 of 42,396 UK postings with tagged skills name LangChain, according to Shakers' market analysis (n=48,596 active, June to October 2026). Of Jev specifically, not yet: it is a recent model and no posting names it.

Recursos relacionados