All postsInside the lab

Jev by TypeSafe: the case for small, fast AI decisions

Written by
Jake Worsham
Published
October 9, 2026
Read time
6 min read

What TypeSafe's Jev model does well, where it struggles, and how Sync is testing it for fast, low-cost AI decisions inside business workflows.

Look closely at any AI workflow we build and most of the steps are small decisions. Is this email a sales lead or a support request? Which department owns this ticket? Does this invoice have a PO number? Did the agent finish the job it was given?

Each of those questions has a short list of valid answers you could write down in advance. For the last few years the default way to answer them has been to send the question to a large language model, wait a second or two, and parse the text that comes back. It works. It is also slow, it costs more than it should, and every so often the model replies with something that is not on the list.

Jev is a model from TypeSafe AI that is built for exactly these questions. We have been studying it closely, and it is worth explaining what it is and why we think it matters for the businesses we work with.

What Jev is

TypeSafe calls Jev a "System One" model. The name borrows from the idea of fast, instinctive thinking as opposed to slow, deliberate reasoning. Jev does not write text. You hand it some state (an email, a form submission, a log, a transcript) and a set of questions, and it hands back probabilities.

The questions come in three shapes:

  • Choice. Pick one option from a list you define, such as "sales, support, billing, or spam."
  • Score. Place the input on an ordered scale, such as urgency from low to critical.
  • Yes or no. A single probability between 0 and 1.

You can ask several questions in one request. Jev answers them in parallel against the same input, so checking five things about a document takes about as long as checking one.

Why that shape matters

It answers inside the lines. Jev can only return an answer from the list you gave it. There is no stray paragraph to parse and no invented category to handle. That does not mean every answer is right, and we will come back to that, but it removes a whole class of failure that shows up in production with text-generating models.

It tells you how sure it is. Every Choice and Score answer comes with a confidence value. That gives you a dial you can set per decision. When confidence is high, the workflow acts on its own. In the middle, it asks a person to confirm. When it is low, it escalates. A routing decision can run on a looser threshold than an action that sends an email to a client.

It is fast. TypeSafe and early users report answers in the range of tens to a few hundred milliseconds. An analysis of launch-week reports, covered by InfoQ, put the median speedup users saw over the model they replaced at about 7 times. TypeSafe's own homepage shows much larger numbers on a single workflow. We would plan around the median.

It is cheap. TypeSafe lists Jev at $0.042 per million input tokens, with output free. The same launch-week analysis put the median cost saving users reported at about 30 times. At that price, running a check on every single email, ticket, or agent run stops being a budget conversation.

The logic stays in your code. Because each answer is a typed value, the branching that happens next lives in ordinary code you can read and test. TypeSafe's own guidance is to break a complicated judgment into several simple questions and combine the answers in code. That fits how we already build: small, checkable steps with a clear owner for each one.

Where it struggles

TypeSafe is unusually direct about Jev's weak spots, and we think anyone adopting it should read that list before writing a line of code. According to their documentation, Jev:

  • takes questions at face value, so vague wording gets vague results
  • is unreliable at counting, arithmetic, and dates
  • loses accuracy on double negatives and multi-step reasoning
  • gets worse when the input is full of unrelated detail
  • can be steered by adversarial content inside the input
  • leans slightly toward the first option in a list

There is also no published accuracy benchmark yet. A model that can only return valid answers can still return a confident, valid, wrong one. The confidence value helps you catch the uncertain cases. It does not replace testing on your own data.

Our working rule follows from that list. Anything that involves math, dates, or a fixed lookup stays in code. Jev gets the judgment calls a person could settle at a glance.

How we are testing it

We do not put a new model in front of a client because its launch post looked good. We test it against a decision we already make every day.

Every scheduled AI Employee we run ends with a check: did this run do what it was supposed to? We built a test set of 64 hand-labelled runs, half good and half bad, with a portion held back so we cannot tune our way to a good score. Our current check, built on a general language model, passes zero of the bad runs, at a median of about 2.4 seconds per check.

That zero is the bar. Jev gets one yes-or-no question per acceptance criterion, all in a single call, and the run passes only if every answer is a confident yes. If Jev matches that record at a fraction of the time and cost, we adopt it, with the language model kept as a fallback for the low-confidence cases. If it lets a bad run through, it does not ship.

Building the test set paid off before Jev answered a single question. Writing out the bad cases exposed a gap in our existing check, and the fix is already in review. Testing against your own data tends to do that.

Where it fits for a business

The pattern that interests us most for agentic workflows is the pairing: Jev handles the high-volume, clear-cut decisions in milliseconds, and a larger model or a person handles whatever it is unsure about. A few places that pattern fits:

  • Support routing. Pick the department and score urgency on every incoming ticket. Unclear tickets go to a person.
  • Inbox and lead triage. Sort every inbound message into reply now, research first, or no action, and flag the ones that need a human eye.
  • Document checks. Ask a short list of yes-or-no questions about every invoice or form, one per requirement, and keep the totals and due dates in code.
  • Guardrails for AI agents. Before an agent takes an action, ask whether that action is safe to run.

A quick test for whether a decision belongs with a model like Jev: can you write every valid answer down in advance, could a person settle it with a glance, and does it happen often enough that seconds and cents add up? If the answer is yes three times, it is a candidate.

We will share real numbers from our own trial once we have them. If your team has a workflow full of small decisions that a person is making by hand, or that an expensive model is making slowly, that is the conversation we want to have.

Back to all postsJake Worsham

Want this kind of work running inside your business?

Every post on this blog comes out of the work happening inside Sync. If the thinking here sounds like the kind of partner you want, the discovery call is the right next step.

Direct CEO access · Replies within one business day