Skip to content

LAYA: Open models for a closed decision

Laya blog nucleusbox

Jev, from TypeSafe, takes unstructured text and returns a probability over labels you wrote. The weights are not public, and TypeSafe has not published a layer diagram. This note is the open-source side of that job: classify, score, or route, with repo ids I opened on Hugging Face. The Jev call itself is in Jev is a typed decision model.

Laya is a real model. I searched Hugging Face for Laya, Laia, and the close spellings. convaiinnovations/laya is a public text-classification repo with a model.safetensors file, created 18 September 2026, three days after the Jev launch. A search for laia returns unrelated repos, including image classifiers and empty model pages. There is no Laia decision model to download. I am not going to invent one.

I have not installed the laya package, and I have not run a forward pass. Everything below is the model card, the config files, and the Hub metadata.

Laya copies the question types

The card describes Laya as a non-autoregressive decision model. You pass a state and typed questions. It returns choice, score, or noul, the same three names TypeSafe uses. choice is one option plus a distribution. score is an expected level on an ordinal rubric. noul is a probability that the answer is yes. The card says it does not generate text.

The card’s own quickstart, which I did not run:

pip install laya
import laya

agent = laya.load("convaiinnovations/laya")
result = agent.predict(state, questions)

state and questions are the caller’s. The card’s worked example uses a billing email and the same three question types. The printed comments on that card (billing, a confidence, a noul) are their sample, not a result I observed. Python 3.10 or newer, Apache-2.0, author line Convai Innovations. The code repo is NandhaKishorM/laya, created 18 September 2026. The Hub page says downloads are not tracked for this model. I am not treating a star counter as evidence that anyone has put it on a ticket queue.

Their demo space is convaiinnovations/laya-demo. I have not clicked through a prediction. The card also documents laya-serve, which it says exposes POST /v1/systemone in Jev’s request shape. I have not started that server.

Three checkpoints, and the files in them

The card lists three checkpoints. I opened each repo’s rl_agent_config.json, the encoder config.json where it sits in the repo, and the Hub safetensors summary.

RepoWhat the config namesHub parameter totalRuntime cap in rl_agent_config.json
[convaiinnovations/laya](https://huggingface.co/convaiinnovations/laya)encoder answerdotai/ModernBERT-large, head_layers 2421,293,830max_len 512, head_max_len 192
[convaiinnovations/laya-multilingual](https://huggingface.co/convaiinnovations/laya-multilingual)encoder jhu-clsp/mmBERT-base, head_layers 2321,908,998max_len 1024, head_max_len 256
[convaiinnovations/laya-typed-decisions](https://huggingface.co/convaiinnovations/laya-typed-decisions)encoder answerdotai/ModernBERT-large, head_layers 2, fine_tuned true421,293,830max_len 1024, head_max_len 256

[answerdotai/ModernBERT-large](https://huggingface.co/answerdotai/ModernBERT-large) is a fill-mask encoder. The Hub reports 395,881,664 parameters for that repo, Apache-2.0. The English Laya encoder config I opened is model_type: modernbert, architecture ModernBertForMaskedLM, 28 hidden layers, hidden size 1024, vocab 50,368, max_position_embeddings 8192. The decision runtime still caps that checkpoint at 512 tokens. The position field and the cap are different numbers. Read the cap.

[jhu-clsp/mmBERT-base](https://huggingface.co/jhu-clsp/mmBERT-base) is the multilingual encoder the second config names. Its own card says 307M parameters, 22 layers, hidden size 768, vocab 256,000, max sequence 8,192, and 1,800+ languages. The encoder config inside convaiinnovations/laya-multilingual matches the shape: 22 layers, hidden 768, vocab 256,000, max_position_embeddings 8192. The Laya runtime default in that repo’s rl_agent_config.json is max_len 1024. The card says you can pass max_len=8192 for a long document, and that accuracy past a few thousand tokens is something you should measure yourself. I did not.

temperature in the English config is a fitted triple, plus a temperature_by_options map (choice:2, choice:3-5, noul:2, and so on). The multilingual config is temperature: [1.0, 1.0, 1.0] and an empty temperature_by_options. That matches the card’s warning that the multilingual checkpoint ships uncalibrated.

The typed-decisions checkpoint is marked fine_tuned: true and fine_tuned_from_checkpoint: true. The card says the Router will not select it unless you construct the router with auto_task_detection=True, because it was specialised to four synthetic workflows. Those four names are the same four on TypeSafe’s workflow site: agent-trace observability, customer service, invoice processing, security incidents.

The architecture the card publishes

TypeSafe has not published a layer diagram for Jev. Laya has. The picture above is Laya’s card, not a guess at Jev’s sampler.

Take the billing email on the card. “We were billed twice for March. Please refund the duplicate today.” Three questions sit beside it: department as a Choice, urgency as a Score, churn as a Noul.

Inside one forward pass, as the card describes it:

  1. The state and the question text are written into one sequence. This is not a chat thread and not a next-token loop.
  2. Every option is placed at its own [MASK] token. billing, technical, and other each get a mask. The score levels each get a mask. The yes and no of a Noul each get a mask.
  3. A bidirectional ModernBERT encoder reads that whole sequence. It can see the email and every option at once. It does not emit the answer one word at a time.
  4. A decision head, two transformer layers in all three configs, reads one logit per mask.
  5. A softmax is taken inside each question, not across the whole call. The department distribution sums to 1. The urgency distribution sums to 1. They do not share a denominator.
  6. Choice returns the label with the largest probability. Score returns the probability-weighted level. Noul returns the probability on yes.

One email, three pictures

The bars below are a teaching sketch so the shapes are easy to see. They are not a run I measured, and they are not the printed sample on the Laya card.

The email is just context. The thing the head scores is the mask in front of each option. Department has three masks. Urgency has three. Churn has two. A new label is a new mask in the request. You do not train a new classification head to add “refunds” to the department list.

Each question has its own softmax. Department sums to 1. Urgency sums to 1. Churn sums to 1. A high urgency does not steal probability from billing. That is why you can ask all three in one call and still threshold them separately.

Choice keeps the name on the tallest bar: billing at 0.70. Score does not keep a name. It keeps the weighted level, 0*0.10 + 1*0.60 + 2*0.30 = 1.20, which sits between “soon” and “blocking.” Noul keeps only the yes bar, 0.80. Same machinery, three ways to read it. If you need a spectrum, ask a Score. If you need yes or no, ask a Noul. Do not threshold a Choice probability as if it were a Noul.

Questions in one call are batched into that same pass. The card’s T4 table says about 39.5 ms for one English question and 158.6 ms for ten. I did not time it.

The card, and the config fields above, also say this:

  • The backbone is a bidirectional ModernBERT-style encoder, fully fine-tuned on the English checkpoint.
  • A decision head trained from scratch adds a type embedding, two transformer layers, an option-marker scorer, and an act/escalate head. head_layers is 2 in all three configs.
  • Every option is written into the sequence at its own [MASK] token. The scorer reads one logit per marker. A softmax over that question’s markers is the distribution. The label set is part of the request, so a new rubric does not require a new classification head.
  • Questions in one call are batched into a single forward pass. The card’s T4 table says about 39.5 ms for one English question and 158.6 ms for ten. I did not time it.

The card is blunt about the act head. action.act_probability “reads 1.0 for almost every input,” and the raw logits run against correctness (AUROC 0.30 on 396 labelled decisions, on their card). They tell you to gate on confidence, which they say reaches an AUROC of 0.77 on the same items. I did not recompute either figure.

noul can follow its own option labels. The card says the English checkpoint renders the two options as false: / true:, and that pair can dominate the state, returning a confident no for clearly positive input. Their workaround is to ask the same question as a two-option choice with neutral keys. Check noul on your data before you threshold it.

High-cardinality choice is the other published failure. Options share head_max_len (192 tokens on the English checkpoint, 256 on the others). The card says a 77-way question then gets about 3 to 4 tokens per label, and reports 0.425 accuracy on Banking77 at that default. Jev’s docs cap a single Choice at 255 options and tell you to switch to two stages past that. Different mechanisms, same practical lesson: a huge label list in one question gets worse.

Their benchmark is not TypeSafe’s chart

The Laya card prints a comparison to “TypeSafe Jev 1.13.0.” The same card says those Jev figures were not measured there, because the authors had no TypeSafe API access. Do not paste them next to the table in the Jev post.

TypeSafe’s public chart, read 26 September 2026, scores agreement with GPT-6 Astra and Claude Fable 5.1. On that chart Jev’s equal-weight workflow point is 67.8%, and invoice processing is the weak one at 61.8% agreement. Laya’s card scores a different set: 400 cases, 2,000 decisions, four workflows, and it calls the metric accuracy. The numbers below are the card’s measurements of Laya, not a second reading of evals.typesafe.ai.

Checkpoint, on their typed-decisions setAccuracySoft accuracyBrierECEScore MAE
laya-typed-decisions0.7660.4710.0620.2130.242
laya0.3620.3320.3160.1750.694
laya-multilingual0.3420.3260.4390.2850.687

The card also prints a per-question majority-class baseline of 0.461 and a teacher self-agreement ceiling of 0.735. The two base checkpoints sit under that majority baseline. The 0.766 belongs to the checkpoint fine-tuned on that benchmark’s own training split (their card says 1,200 cases, 6,000 decisions). Per workflow, the card gives that fine-tune 0.804 on invoice processing, 0.766 on security incidents, 0.764 on customer service, and 0.730 on agent-trace observability. That invoice number is not a zero-shot win over Jev’s 61.8% agreement. It is a specialist, trained on the set, scored on a metric TypeSafe did not publish.

The card’s other self-report I would keep in view: the shipped probabilities are over-confident. Refitting one temperature per question type and option count moves mean ECE from 0.466 to 0.081 on the English checkpoint, and from 0.314 to 0.106 on the multilingual one, on their data. Do that on a labeled slice you trust before a threshold goes anywhere near a refund or a page.

The older models that already do the job

“Similar architecture,” for Jev, can only mean the closest published class. TypeSafe has not published a layer diagram. The class is a transformer that emits a distribution over a closed label set: a classifier, a zero-shot NLI model, a cross-encoder, a reward model. They do not speak Choice, Score, and Noul. They do put a probability or a score on a label you can branch on, and you can download them today.

Zero-shot labels, one pass per label. [facebook/bart-large-mnli](https://huggingface.co/facebook/bart-large-mnli) is BartForSequenceClassification, trained on MultiNLI. The Hub reports 407,344,133 parameters, MIT license, pipeline tag zero-shot-classification. The card describes the Yin et al. method: the text is the NLI premise, each candidate label is turned into a hypothesis such as “This text is about politics,” and the entailment and contradiction probabilities become the label distribution. The Hugging Face zero-shot pipeline runs that once per label. Labels that can all be true are a separate call with multi_label=True, so the scores are not forced to sum to 1. That is the open version of “the label set arrives at request time.” It is not one forward pass for a whole rubric, and it has no ordinal score type.

[MoritzLaurer/deberta-v3-large-zeroshot-v2.0](https://huggingface.co/MoritzLaurer/deberta-v3-large-zeroshot-v2.0) is the same job on a different backbone. The Hub lists DebertaV2ForSequenceClassification, 435,063,810 parameters in float16, MIT, pipeline tag zero-shot-classification, base_model microsoft/deberta-v3-large. The card positions the series against facebook/bart-large-mnli and publishes its own 28-task table. I am not copying that table. Read the metric on the card, and read the -c variant if you care that the training mix stay commercially licensed. The repo without -c is the one I opened.

Pair scoring. [cross-encoder/ms-marco-MiniLM-L6-v2](https://huggingface.co/cross-encoder/ms-marco-MiniLM-L6-v2) is BertForSequenceClassification, Apache-2.0, pipeline tag text-ranking. The Hub safetensors summary reports 22,713,601 float32 parameters. The card’s table gives this id NDCG@10 of 74.30 on TREC DL 19 and MRR@10 of 39.01 on MS MARCO dev, at 1,800 docs per second on their hardware. You pass a query and a passage. The model scores the pair. Sort the scores and you have a rerank. That is routing among documents, not among departments, unless you phrase each department as a passage and accept a score that is not a distribution over the set.

Reward models. [OpenAssistant/reward-model-deberta-v3-large-v2](https://huggingface.co/OpenAssistant/reward-model-deberta-v3-large-v2) is DebertaV2ForSequenceClassification, MIT, tagged as a reward model. The card’s usage is tokenizer(question, answer) and then logits[0]. A higher logit is the answer the model ranks as better. The card’s validation-split accuracy for this id is 61.57 on WebGPT, 71.47 on Summary, 99.88 on SyntheticGPT, and 69.25 on Anthropic RLHF. The card itself says SyntheticGPT likely has a surface pattern that makes the pair trivial. I did not run the snippet, and the Hub response I opened did not include a parameter total, so I am not guessing one.

A reward model answers “which of these two completions did the preference data like?” It does not answer “which queue, and how sure?” If I used one inside an agent eval, it would rank two final messages. The questions I actually want on a trace are narrower: did a tool argument contradict the user, did the reply promise a tool result that never happened. Those are closed judgments. A reward logit is one feature. The gate stays in code.

What I would download

The label set is fixed and you have labels: fine-tune a classifier, or start from one of the sequence-classification repos above. The label set changes per request and only one label should win: facebook/bart-large-mnli or the DeBERTa zero-shot repo, and pay one forward pass per label. You are reranking passages: the MiniLM cross-encoder. You are ranking two completions: the OpenAssistant reward model, after you read the SyntheticGPT caveat.

You want Choice, Score, and Noul in one local call, with the rubric written in the request: convaiinnovations/laya for English, convaiinnovations/laya-multilingual when the script is not Latin, and convaiinnovations/laya-typed-decisions only if your task is one of those four workflows and you accept that it was fine-tuned on that benchmark. Then fit a temperature on your own labels. The English context cap in the config is 512 tokens, against Jev’s published 64k for the whole request. A long trace does not fit the English checkpoint without the multilingual max_len override, and the card tells you to check that override yourself.

Jev remains the hosted contract: one versioned ID, a 64k request budget, output tokens unbilled, no weights. I have not called it. Whichever probability you take, the lesson from the loads I used to own is the same. Refuse a bad extract. The field parsed. That does not make it fit to load.

Suppose you have an API key and want the hosted call. Jev is a typed decision model is the request, the response fields, and the gate I would put on top of them.

If you are about to score a trace with one of these models. Your Text-to-SQL Agent Passed the Demo. Did It Survive Production? is the labeled check. A local probability is still not ground truth.

If you run Laya on a rented GPU. The English card times a T4. How to Choose a GPU for LLMs is the cost decision before you leave a card on for a 512-token classifier.

Sources

0 0 votes
Article Rating
Subscribe
Notify of
guest
0 Comments
Oldest
Newest Most Voted