Skip to content

Jev is a typed decision model. Route the ticket, then read the probability.

jev blog nucleusbox

If the next step in your system is a closed judgment, route this ticket, score this trace, and gate this action; Jev is worth a measured trial. If that step needs sentences, source code, or a count a parser can already do, leave Jev off the path.

I have not sent a request to this API. Early access opened on 15 September 2026. The call below is the official quickstart, with the model pinned. The probabilities in the sample response are the ones the docs print. They are not a run of mine. Where a figure is a vendor chart, I say so. Where the model card is silent, I do not fill the gap.

The open weights that aim at the same job are in Open models for a closed decision. Laya is a real Hugging Face model. Laia is not.

The problem it solves

A lot of production code already knows the shape of the answer. The team is billing, technical, or sales. The trace should page a person, or it should not. The claim is a meal, travel, or equipment. What the code does not know is which of those labels the unstructured text supports.

A chat model is trained to continue a string. You can force that string into JSON, then parse it, then hope the parse matches the type your code branched on. A hand-written rule breaks when the wording moves and the meaning does not. Jev is the third option TypeSafe is selling: you send a state and a map of questions whose options you wrote, and the service returns one typed answer per question, with a distribution.

That is also the limit. The how-to-build page says System One “does not generate code or choose its own next action.” Code owns control flow, side effects, and anything a deterministic rule can do. The model is inserted only where the input is unstructured and the judgment is narrow. If you wanted a planner, this is the wrong primitive. If you wanted the if inside the planner, this is the pitch. (how to build)

Their own expense-claim cartoon is the pattern I would trust.

  1. Code does the exact work. Days overdue, amount over a limit, whether a receipt file exists, whether a regex already found the invoice number.
  2. The model answers one atomic question a person could make in a second: is this a meal, does the description match the receipt, is the customer asking for a refund.
  3. Code applies the policy. A meal over $75 whose description does not match goes to a manager. Anything the model is unsure about goes to a person. The threshold is a constant in review, not a sentence buried in the instruction.

Why people are trying it

On 15 September 2026, Diogo Almeida published “Introducing System One Models & Jev” and opened early access. TechCrunch, on 18 September 2026, reported that he left OpenAI two years earlier to start TypeSafe, and that demand after launch was high enough that the company briefly could not serve API users. (launch post, TechCrunch)

The name is literal. TypeSafe named the model Jev, after the economist William Stanley Jevons, on the idea that a large drop in the cost of a resource can raise demand for it. The class name, System One, is their nod to Daniel Kahneman’s fast judgments. It is not a JEPA, a Jamba, or a speech-to-text ghost of Gemma.

The practical attractions, as they publish them, are a schema the service fills in, a distribution on every Choice and Score, and a bill that does not charge output tokens. The models page prices input at $0.042 per million tokens, which they also print as $42 per billion. Output tokens show up in usage and are not billed. Rate limits on that page: 250,000 tokens per second and 1,200 requests per minute, with 429 past either cap. The same page says those limits “can change without notice.” English is the primary training language. Other languages, including CJK scripts, “are handled but not equally well.” There is no per-customer fine-tune. The same weights serve every account. (models)

The homepage multipliers, 193.6x faster and 444.6x cheaper, are called out in the launch post as coming from the workflow evals, and as the high end of what they expect in the real world. They also say they cannot prove the price is unsubsidized. I would budget the published rate. I would not budget 193x, and I would not budget the hope. The table further down is the number I would keep.

A call you can run

The no-code path is the playground. The quickstart says to open it and log in, paste a state, and add questions. I have not used it.

The code path is the Python SDK. It wants Python 3.10 or newer. TypeSafeClient() reads TYPESAFE_API_KEY from the environment. The quickstart omits model and the client then sends jev-latest. I pass jev-1.13.0, which is still the only stable ID on the models page. jev-latest and jev-preview both point at that ID today. An alias moves when a new release ships. (quickstart, models, client)

pip install typesafe-sdk
import os
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient

client = TypeSafeClient(api_key=os.environ["TYPESAFE_API_KEY"])

ticket = (
    "Hi, I've been trying to connect my Stripe account for 3 days "
    "and the integration keeps failing. I'm losing sales. Please help ASAP."
)

response = client.system_one(
    state=ticket,
    model="jev-1.13.0",
    questions={
        "department": Choice(
            instructions="Which team should handle this",
            criteria={
                "billing": "Payment or subscription issues",
                "technical": "Bugs or integration problems",
                "sales": "Pricing or account questions",
            },
        ),
        "frustration": Score(
            instructions="How frustrated the customer appears",
            criteria=[
                "Calm, just stating facts",
                "Frustrated but civil",
                "Very angry, strong language",
            ],
        ),
        "is_urgent": Noul(
            instructions="The message conveys urgency or time-sensitivity",
        ),
    },
)

I have not executed this. The HTTP equivalent is POST https://api.typesafe.ai/v1/systemone with Authorization: Bearer and the same JSON body. The question key is for your code. The docs say the key is not sent to the model, so a clever key is not a prompt. The instruction text is the prompt. (primitives)

What comes back, and one routing decision

The quickstart prints this response for that payload. Read it as documentation of the schema.

{
  "model": "jev-1.13.0",
  "answers": {
    "department": {
      "type": "choice",
      "choice": "technical",
      "confidence": 0.78,
      "probabilities": {
        "technical": 0.85,
        "sales": 0.0,
        "billing": 0.15
      }
    },
    "frustration": {
      "type": "score",
      "score": 1.0,
      "confidence": 1.0,
      "legend": {
        "0": "Calm, just stating facts",
        "1": "Frustrated but civil",
        "2": "Very angry, strong language"
      },
      "probabilities": {
        "0": 0.0,
        "1": 1.0,
        "2": 0.0
      }
    },
    "is_urgent": {
      "type": "noul",
      "noul": 1.0
    }
  },
  "usage": {
    "input_tokens": 392,
    "output_tokens": 65
  }
}

model is the versioned ID that answered, including when you sent an alias. answers is keyed by the names you chose.

On a Choice, choice is the option with the highest probability. probabilities is a distribution that sums to 1. confidence is a statistic of that distribution, from 0 to 1. A peak reads as high confidence, a flat spread as low. It is not a second opinion about correctness. The confidence page’s demo, for three options, uses (3 * largest probability - 1) / 2, which is the special case of (n * peak - 1) / (n - 1). TypeSafe calls this a convenient default. A fuller note on other computations is still a cookbook they have not linked. Do not ship a threshold as if that demo formula were a calibration curve. (confidence)

On a Score, criteria is an array of level descriptions, at least two and at most ten. The model sees the words, not the index of the neighbour. score is the probability-weighted mean of the level numbers. The current docs’ worked example is 0 * 0.0 + 1 * 0.57 + 2 * 0.43 = 1.43. A score of 1.0 can mean all of the probability sits on level 1, or half sits on 0 and half on 2. Read probabilities and confidence with the score. legend maps the level number back to the text you sent. Levels should describe situations. On the misaligned-button report, levels written as "0", "1", "2" come back in the docs at score 0.55, confidence 0.33. The same report with descriptive levels scores 0.0 at confidence 1.0. Those are documentation examples, not a run of mine. (score)

On a Noul, noul is the probability that the answer is yes, from 0 to 1. There is no confidence field. A value near 0.5 means the model splits yes and no. It does not mean “medium” on a skill you forgot to define. If you wanted a spectrum, you wanted a Score. (noul)

usage.input_tokens is what you pay for. usage.output_tokens is reported and not billed. There is no explanation string to parse. That is the point, and it is also why you cannot ask it why.

The tiny decision I would write on top of this call:

department = response.answers["department"]
urgent = response.answers["is_urgent"].noul

# 0.9 is the docs' illustration for a high-stakes approval.
# It is not a threshold I fitted, and I have not run the call.
if (
    department.choice == "technical"
    and department.confidence >= 0.9
    and urgent >= 0.9
):
    route = "page_oncall"
else:
    route = "queue"

On the quickstart’s printed sample, choice is technical and noul is 1.0, but confidence is 0.78. This gate would queue the ticket. Drop the confidence check and the same sample would page on-call. The confidence docs use 0.5 as a floor and 0.9 before an approval, and a different pattern page uses 0.6 and 0.85. Those are illustrations. Start conservative and move them on your own labels.

Two invariants people will assume, and the jaggedness page shows you should not. On one ticket, a Noul for “asking for a refund” returned 0.22 while a yes/no Choice put 0.01 on yes and 0.99 on no, at confidence 0.97. On another ticket, a Noul and its negation summed to 1.19 (0.72 and 0.47). A Choice is relative among the options you listed. A Noul is an absolute probability. Ask the decision once, in the type you will threshold, and enforce identities in code. (jaggedness, reviewed 2026-09-17)

How the score is produced

The picture is the contract you can check. A ticket goes in as state. Three questions go in beside it. One response comes back with a probability on each. TypeSafe has not published a layer diagram, a parameter count, or a paper. Almeida told TechCrunch he is tight-lipped about the architecture. TechCrunch describes Jev as transformer-based and says outside observers suspect an open-weight base. That suspicion is not a model card. The company says it built “a new model architecture, parallel sampler, and training method” called Reinforcement Learning for Calibrated Decisions (RLCD). Those three phrases are not a schematic.

What the API makes checkable is the contract. state is text: a string, a JSON object, or an array of text values. Images, audio, and video are not accepted. The context budget is two numbers. The models page sets 64k tokens for the whole request, meaning state plus every question, and 32k tokens for state plus the single longest question. Jev reads the state once and scores the questions against it in parallel. Adding a question still costs input tokens. The docs say it barely changes latency. (system one, models, fan-out)

You can mix Choice, Score, and Noul in one call. Each is scored independently, so one answer is not hidden context for the next. A Choice accepts at most 255 options. Above that, their wikiracing demo switches to a two-stage pattern: score independently, then choose.

Sampling is the other half of the speed claim. A normal language model emits one token conditioned on the last. TypeSafe says Jev generates all outputs in a single query, in parallel, and that this is why it can give up string generation and still return a full distribution. I can check the API shape. I cannot check the kernel. The speed numbers they publish were taken, they say, from laptops on the US West Coast, where the service was hosted at launch.

Calibration is a claim about groups. The primer says outcomes tagged 0.2 should happen about 20% of the time across many predictions, and that none of this is a promise about one answer. RLCD is the training objective aimed at that contract: decisions plus probabilities, higher probability lined up with a higher chance of being correct. The primer sets this next to RLHF, which optimizes for what people prefer to read. Almeida’s line to TechCrunch is the same point in plainer words: they had been “super good at human language for four years,” and that objective was the wrong language for automation. He also told TechCrunch that Jev is trained exclusively on synthetic data. The primer does not document the generator, the volume, or the filters. Treat “exclusively synthetic” as an interview claim until a paper exists. (primer)

The closest published class, and the open models in it, are the other post. Short version: a transformer that emits a distribution over a closed label set. Classifiers, zero-shot NLI models, cross-encoders, and reward models already do a narrower version of that job. Laya is the one that copies Choice, Score, and Noul.

Pin jev-1.13.0

There is no previous public Jev to diff. The launch calls this the first public System One model. The models page offers jev-1.13.0 and nothing older. Several cookbook pages still pin jev-1.12 and date their cached numbers to August 2026. Those pages do not say what 1.13 changed in the weights, the sampler, or the scores. A citation cookbook result on jev-1.12 is a snapshot of that ID on that day. If you have tuned a threshold against any earlier ID, pin jev-1.13.0 and fit the threshold again. The response model field is supposed to echo the versioned ID even if you sent an alias. GET /v1/models lists aliases. The versioned ID is accepted anyway.

Where jev-1.13 breaks

The jaggedness page is the most useful document in the launch, because it is a list of failures the vendor wrote down. It applies to jev-1.13, last reviewed 17 September 2026. Best on System One tasks, literal, weak at indirection, weak at numeric precision.

It reads the instruction you wrote, not the one you meant. Negations and implied scope are face value. If you find yourself explaining a wrong answer, that explanation is the missing half of the prompt.

It does not count reliably. Characters, term occurrences, long lists. The error grows with the size of the count. If a regex or a parser can see the unit, the count belongs in code. When the criterion is semantic, their pattern is one Noul per item and a sum in your process.

Dates are text, not ordered time. Extract month, day, and year as Choices over closed sets, including an explicit “not stated,” and do the calendar math in code.

Indirection costs accuracy. A property of a property, a double negative, a question that needs several hops. Point at the field by name.

Irrelevant bulk in state acts as a distractor. They call it context rot. Retrieve first.

Adversarial text can move the answer. Injected instructions, a passage arguing for its own label, a misleading frame. The model does not treat state as hostile by default. Precise criteria and a test set of attacks are part of the integration.

Contradictory instructions and criteria confuse it. A Noul whose true criterion means “no” is a self-inflicted wound.

Score levels are weak in numerical calibration. Use the expectation to cross a threshold you chose, or to rank. Do not use it as a calculator.

Generation is absolute. Chaining Choices to spell text “will not work well and will be very slow.” If the answer space is open, use a generative model, or use one to propose candidates and let Jev pick among them. (jaggedness)

“Can’t hallucinate,” in the launch post, means the service cannot return a shape outside the schema. They say a type error is “mathematically impossible,” and that the 0% they plot for themselves “is not empirical.” A guaranteed shape can still be the wrong option. A wrong option at confidence 0.97 is a confident mistake. Armin Ronacher’s comment to TechCrunch is the right user-side reading: if the probability comes back near 50%, treat it as a coin toss and drop it; if it comes back at 95%, you still decide what your code is allowed to do.

Which launch numbers to keep

Four days before the launch, Almeida published “Lies, Damned Lies, and Benchmarks.” The policy in that post: no standard benchmark table in model releases, new evals as dated snapshots that get retired rather than hill-climbed, and internal evals published with the caveats and the results that look bad. (benchmarks post, 11 September 2026) The workflow site is that policy in practice. Read it as a dated vendor chart.

The method, in their words: assume the harness is correct, give every model the same workflow, and score agreement with the average of GPT-6 Astra and Claude Fable 5.1, both at high thinking. Other models run at the provider’s default reasoning setting. The aggregate gives the four workflows equal weight. Accuracy here means agreement with those two reference models. It does not mean agreement with a human label, and it does not mean agreement with your policy. They also note that using Astra and Fable as the reference biases the chart toward OpenAI and Anthropic, and that they likely underestimate Jev and DeepSeek. The workflows were written by people on their model capabilities team. LLMs in this eval are wrapped in TypeSafe’s System One adapter so they emit the same decision schema. The speed and cost gaps are partly a gap against that wrapper. (evals)

The equal-weight workflow points, as the eval page displayed them on 26 September 2026:

Model, workflowAgreementCost per caseTime per case
sol74.1%$0.083623.3 s
opus 573.1%$0.176137.8 s
terra67.9%$0.030410.1 s
Jev67.8%$0.00040.4 s
sonnet 567.8%$0.117478.1 s
luna66.8%$0.003312.9 s
DS v4 pro65.5%$0.041386.5 s
DS v4 flash64.4%$0.005951.9 s
haiku 4.553.6%$0.019512.5 s

Source: evals.typesafe.ai, read 26 September 2026. The names are the labels on that page. I am not expanding sol, luna, or DS v4 past what the page prints. The launch post’s side-by-side calls one comparator GPT-5.6 Terra.

On that aggregate, Jev ties sonnet 5 at 67.8% and sits 0.1 points under terra. sol is 6.3 points higher. Dividing the rounded aggregates, terra takes about 25 times as long and costs 76 times as much. That is not 193 times, and it is not 444 times. Keep the per-point table. Do not keep a single multiplier.

The equal weight also hides the workflow where Jev falls off. Per-workflow Jev figures from the same pages, same day:

WorkflowJev agreementJev costJev timeA stronger workflow point on the same chart
Security incidents61.7%$0.00010.3 sopus 5 at 66.2%, $0.0574, 15.1 s. terra is lower, at 51.2%.
Agent-trace observability71.6%$0.00030.5 ssol at 76.6%, $0.0575, 40.3 s. luna at 76.1%.
Invoice processing61.8%$0.00110.5 ssol at 79.1%, $0.2152, 34.3 s. terra at 74.7%. opus 5 at 78.4%.
Customer service76.0%$0.00010.4 ssol at 78.3%, $0.0323, 10.1 s. DS v4 flash at 76.8%.

Invoice processing is the result that should stop a blanket rollout. Jev is cheap and fast there, and it is the clear accuracy laggard against the same reference. Security is a closer call: Jev beats terra’s workflow agreement and trails opus 5. Customer service is the flattering one: 76.0% against sol’s 78.3%, at a cost the table prints as $0.0001. If your job looks like invoices, the aggregate Pareto story is the wrong slide.

The page also says that, averaged across the four tasks, every model is more accurate, cheaper, and faster inside the workflow than with the same policy stuffed into one prompt. Jev itself is plotted only as a workflow point. A lot of the gain on this chart is decomposition plus code, and that gain shows up for the LLMs too. Jev is the cheap way to occupy the judgment slots inside a harness you already believe. It is not evidence that a single call replaces the harness.

A few other launch figures, with the scope attached:

  • Side-by-side demo against GPT-5.6 Terra at default reasoning. TypeSafe says the only disagreement in the recorded run was an ambiguous “churn likelihood” label, and that the short state “paints our model in an advantageous light.”
  • Doom demo: structured text state, not pixels. They mention about $7 an hour at 10 queries a second, and they say a non-AI bot could play better.
  • Wikiracing: speedups “a lot less” than the other demos, because the LLMs were in non-reasoning modes, except Astra at its lowest reasoning setting. Jev’s 255-option cap forces the two-stage path on huge link lists.
  • TechCrunch reported two outside tests. Pranit Sharma at Vercel said a safety classifier moved off ChatGPT Luna 5.6 onto Jev and came back 5 to 18 times faster, with greater accuracy, in that test. Nikhil Mudholkar at Bryo said Gemini was slightly more accurate on business-email classification and 10 to 20 times more expensive. I have not seen the sets, the prompts, or the labels. They are anecdotes in a news story.

What I would trust enough to design around: the request schema, the three primitives, the 64k and 32k budgets, the published input price, the fact that output is unbilled, the 255-option cap, the absence of a confidence field on Noul, and the jaggedness list. What I would not trust as a property of my traffic: 193.6x, 444.6x, “can’t hallucinate” as a synonym for “correct,” agreement with Astra and Fable as accuracy, a threshold copied from a docs snippet, and any August cookbook number still labeled jev-1.12.

Where this sits next to agent evals

I spend my time on agent evals. The workflow on their chart that matches that work is agent-trace observability: a support agent has finished, the trace includes every tool call, and the decision is whether a person should look, and how soon. Jev’s published point there is 71.6% agreement, $0.0003, 0.5 seconds. The sol point is 76.6% at $0.0575 and 40.3 seconds. I would not collapse that into one Noul that says “was this run good?” A broad question hides several judgments, and this model is weakest when you add hops. I would split the trace into separate questions a reviewer can audit: did a tool argument contradict the user, did the final message promise something the tools did not do, is there a policy term that should force a human. Code would combine those probabilities with the rules that are not judgments at all.

That is the same instinct as a data load. At Informatica I spent years on the part that happens after a mapping looks finished: refuse a bad extract, own the threshold, do not let a convenient field become a fact because it parsed. Jev can propose the feature. The schema guarantee means the feature will have the type I asked for. It does not mean the feature is fit to load. A probability is an input to a gate. The gate stays in code, with a labeled slice I control, and with the model ID pinned so a quiet alias change cannot move the gate under me.

I would also keep the reference-model trap in view. TypeSafe scored Jev by how well it matches two frontier models. If my eval product then scores agents by how well they match Jev, I have built a chain of agreement. Useful for a cheap first pass. Not a ground truth. The way I would use the probabilities is the way the primer defines calibration: bin a few hundred labeled traces, check whether the 0.8 bin is right about 80% of the time on my policy, and only then let a high bin auto-pass. Until that plot exists for my data, confidence is a routing hint.

Serving follows the same rule. Measure latency from the region you actually call, not from their West Coast laptop band of 70 ms to 500 ms, and not from the “about 100 ms” sentence on the how-to-build page. Batch every question the request will need. Log usage.input_tokens, the resolved model, and the full distribution, not only the argmax. Alert on 429, because the published rate limit is explicitly unstable while they absorb demand. And if the service blips the way TechCrunch described in the first week, the fallback is the deterministic branch you should have written anyway.

Jev is a new kind of call, not a new kind of oracle. The release is real, the schema is the product, and the numbers worth acting on are the ones you recompute on labels you trust.

Three pages. Each one is the next decision, not a tour of the site.

If you want the weights. Open models for a closed decision is the downloadable side of this same job. Laya speaks Choice, Score, and Noul. BART and DeBERTa already return a label probability, and they do not pretend to be Jev. Read that before you treat a Hugging Face card as a second benchmark.

If you already have an agent trace. Your Text-to-SQL Agent Passed the Demo. Did It Survive Production? is the eval this post keeps pointing at. Jev can propose “should a person look.” That post is how you tell whether the proposal is fit to load.

If the fast judge still ships a bad action. Your Gemini Agent Just Crashed. Here’s Why is the failure that a 100 ms probability does not remove. The gate is only as good as the question you asked.

Sources

0 0 votes
Article Rating
Subscribe
Notify of
guest
0 Comments
Oldest
Newest Most Voted