Those of us who are in the industry & have read Gartner’s 2026 Hype Cycle for Agentic AI, acknowledge AI is at peak of inflated expectations. Less than 20% of enterprises have deployed agents, yet more than 60% expect to in the next year horizon. And, most of these deployments are quite narrowly scoped & fully autonomous agents are not yet ready for majority of the use cases.
The more telling detail is what else sits on the curve. Governance, security and FinOps for agentic AI appear as their own profiles, next to agent management platforms, orchestration, the agent development life cycle and context graphs. Gartner’s message is that adoption depends as much on foundations as on agent intelligence.
I want to add one architectural idea to that list, because a model released this month makes it concrete.
The default pattern has a cost problem
Most enterprise agents today run every step through a large language model. Routing a ticket, checking whether an answer is grounded, deciding whether to escalate, judging whether a task is finished: each is a full generative call. The model writes prose that another piece of code then has to parse back into a decision.
Most of these steps are not reasoning problems. They are classification problems dressed up as conversation. That is expensive, slow and hard to audit. It also explains why cost and control are climbing the Gartner curve so early.
System One models (Thanks Mr. Kahneman)
Jev from TypeSafe AI does not generate natural language. It returns typed values together with probability estimates and confidence scores, meant to be consumed directly by other software. As we know, the company calls this class System One, after Kahneman’s fast and intuitive mode of thinking.
You send a state and typed questions, and every question is evaluated in parallel and in isolation against that state in one call. One developer describes it as a smart if statement, a place where code needs a single judgment that ordinary logic cannot compute.
The vendor numbers deserve attention and scepticism in equal measure. TypeSafe claims up to 100 times faster and 100 times cheaper than conventional LLMs for certain tasks. On TypeSafe’s own four workflow evaluation, Jev scored 67.8%, level with one frontier model and behind two others, at roughly 1/200th the cost. TypeSafe itself acknowledges that its team created those workflows, notes possible bias and says the reported gains likely sit at the high end of real-world results. Treat these as hypotheses to test, not procurement inputs.
The layer nobody budgets for: The judge
Every serious agent programme eventually needs a way to grade what its agents do. Governance, security and reliability all depend on it. The default method is LLM-as-a-judge: hand the output to another model with a rubric and ask for a verdict.
It works, but it stacks a second generative bill on top of the first. Evaluating agent output is repeated work across every task, every trace and every model revision, so the judge often costs as much as the agent or forces teams to grade only a thin sample. There is also a subtler weakness. Judges are usually measured by accuracy against a gold set, which tells you how many errors to expect but not where they will be. A verdict without a trustworthy confidence signal cannot be automated, because a person still has to review everything.
This is where the System One idea gets interesting for CIOs. A decision model is a natural judge.
One recent paper describes a two step approach. A first judge returns a verdict along with probabilities over the allowed labels. When its confidence is high the verdict stands, and when it is uncertain the case goes to a stronger LLM. In a frozen offline test this policy kept roughly 99% of the stronger judge’s accuracy at a fraction of its fee, though the authors stress that the result is specific to the tested workload and is not a guarantee. Against sixteen generative and reward model judges, Jev landed within three percentage points of the strongest comparator on ordinary preference and evidence grounded factuality at 0.36% of its fee. Larger gaps appeared when a judgment required checking a derivation or resisting an elaborately written wrong answer.
The ecosystem is already moving. Langfuse now lets teams use Jev as an evaluator alongside their LLM judges and code evaluators. DeepEval supports it as a judge for its metrics. LangChain’s team framed the opportunity as a possible third form of agent evaluator, next to code based checks and LLM judges, and called its own results promising but early.
The emerging pattern is a division of labour. Run the decision model on every trace for coverage, then sample failures and uncertain scores through an LLM judge when you need a written explanation to improve the system. Jev does not replace LLM judges. It has no rationale to give and only answers questions whose possible answers you define upfront.
Overlaying this on the Gartner curve
On cost, FinOps for agentic AI becomes easier to practise when the highest volume work, including evaluation, runs on a model priced for volume. The question shifts from how many tokens an agent burns to how many decisions and audits the enterprise can afford.
On governance, calibration is the feature that matters. Every Jev decision comes with a confidence estimate, so software can act when confidence is high and escalate when it is not. Combine that with a cheap judge on every trace and governance stops being a quarterly sample review. It becomes a continuous control that can trigger alerts and human review in real time.
On security, a typed output has a small blast radius. A model that can only return a choice, a score or a probability cannot be talked into writing a rogue instruction. Prompt injection does not vanish, but the surface shrinks.
On the agent development life cycle, decisions and evaluations become testable components with defined inputs and outputs, not paragraphs of prompt.
How to look at agentic workloads from here
Stop treating an agent as one model doing everything. Treat it as a system with three roles.
Doers handle open ended work such as planning, synthesis, drafting and novel problem solving. This still belongs to frontier LLMs, used where their cost is justified.
Deciders handle high frequency judgments inside the loop: routing, triage, gating, stop conditions and escalation. These need speed, consistency and calibrated confidence.
Judges grade the work of both. They run on every trace, feed governance and cost dashboards and route uncertain cases to a stronger model or a person.
Five practical steps follow.
Inventory the decisions inside your agent loops and tag each as classification or reasoning. Most teams find the first group is larger than expected.
Batch and scope. TypeSafe’s own cookbook reports that batching thirteen questions into one call was 12.2 times cheaper and 10 times faster with identical answers. If that holds in your environment, it changes how you design agent steps and evaluators.
Route on confidence. Let the fast tier act on high confidence outcomes and pass the rest upward. Log both paths so audit and cost reporting come from the same trace.
Judge everything cheaply and some things deeply. Score every trace with a decision model and reserve strong LLM judges for samples and flagged cases.
Validate against humans and pin your versions. HoneyHive advises using the versioned model ID you validated, since the latest alias moves when a release ships, and checking any upgrade against human labelled examples. It also recommends binary questions where possible, exact calculations in code and a generative model wherever a written rationale is needed.
What to watch before you commit
Jev is early.
It launched in limited early access and it is a hosted model called over an API, with the tooling around it open but the model itself running on TypeSafe’s servers. For regulated enterprises that raises data residency, vendor concentration and continuity questions that a seed stage lab cannot yet answer. A judge you cannot audit is its own governance risk, so the human labelled validation set matters more here, not less. Scores on vendor authored benchmarks say little about your own ticket queue.
Have a clear vision, run a bounded pilot on your own data before anything else. (Be clear to evolve from Pilot. Why? Read Here.)
The takeaway
Gartner is right that the hype is peaking and that foundations decide who scales. Two of those foundations are matching the model to the job and knowing when the model is wrong. Agentic AI will not be won by putting the biggest model at every step or by grading every step with another one. It will be won by enterprises that know which steps need thinking, which only need deciding and which need a fast, calibrated judge watching all of it.

