What is behind System One and the open Jev clones?
With Jev, TypeSafe promises a new model class for agent decisions, and within ten days there were a dozen open clones. Measured on an RTX 4090, zero-shot, against German and English test cases: the strongest clones are a Qwen model with its letter logits read out. And a home-built version in 150 lines does just as well.
October 1, 2026system-one · benchmark · calibration · qwen
On 2026-09-15 TypeSafe presented Jev as the first "System One model". The line that came with it was "unstructured state in, typed probabilistic decisions out". Ten days later Hugging Face held more than a dozen open clones, all with the same interface, all with comparison numbers against Jev. We put the most important ones side by side on a workstation with an RTX 4090, together with a home-built version without any training and a large language model as a cross-check.
The result first, because it carries the rest of the text. Behind the hype sits a solid but old technique: a language model picks an option via its logits instead of writing text. The best open clones are a Qwen model with a small LoRA on top. And they are barely better than an untrained 4B model with 150 lines of code.
What does System One actually promise?
A System One model writes no text. None at all. It receives a state, as text or JSON, plus a set of typed questions, and returns a probability distribution for each question.
| Question type | Purpose | Answer |
|---|---|---|
| choice | pick one option from a list | chosen option, distribution, confidence |
| score | place on an ordered scale | expected value, distribution, confidence |
| noul | does the statement hold? | a probability between 0 and 1 |
All questions go out in one call to POST /v1/systemone. The promises are 70 to 500 ms response time, $0.042 per million input tokens with free output, calibrated confidence and "frontier-level" decision quality. Weights? Not published. Neither are architecture, parameter count or training data, and there is no paper on the training method "RLCD".
What of this can you check without knowing the weights? The only certain thing is the format guarantee. The answer is always one of the allowed options. Whether it is the right one is another matter, and the same guarantee has been available for a while from constrained decoding with Outlines or XGrammar, just as from simply reading out the option logits of any language model.
The clones copy the interface one to one, and the official typesafe-sdk runs against each of them via base_url. Is there more in them than that interface? That is what we wanted to know.
How was it measured?
We sent every system exactly the same requests in Jev format and scored them on the same test split with the same metrics. Everything ran zero-shot. No fine-tuning, no examples in the prompt.
The computing was done by a Windows 11 workstation with an NVIDIA RTX 4090 and 24 GB of video memory, and because two of these models never fit on the card at once, only one runs at a time. A runner starts a system's server, measures all datasets and shuts it down again. Each record sends exactly one request with all its questions, and latency is the client-side time for that one request, after three warm-up requests that don't count. CLM needs vLLM, which doesn't exist for Windows, so its Qwen3-8B encoder ran in Docker (vllm/vllm-openai, vLLM 0.27.1, pooling mode).
The test cases
How do you measure something like this if you want to know whether it holds up in your own agent? We took two groups. Four use cases of our own in German, the way they actually come up in agents, and four public benchmarks from Hugging Face in English.
| Group | Dataset | Questions | Records | Labels |
|---|---|---|---|---|
| own | mail_triage | category (choice, 7), urgency (score, 3), reply needed (noul) | 120 | synthetic, German, written by language models |
| own | tool_auswahl | first tool (choice, 9), writing (noul) | 120 | as above |
| own | guardrail | injection, sensitive data, off-topic (noul each) | 120 | as above |
| own | lead_quali | need (score, 4), offer (choice, 5), next step (noul) | 120 | as above |
| HF | jev-bench | 17 classic NLP tasks | 306 | labelled by humans |
| HF | evalsafe-customer-service | customer service decisions | 330 | consensus of two large models, by TypeSafe |
| HF | tasksource-jev-typed-decisions | test split from 203 sources | 280 | derived from the original datasets |
| HF | system-one-decisions | ticket type and priority, German and English | 150 | from a generator |
One third of each dataset is the dev split, two thirds the test split. The assignment depends only on the hash of the ID and is therefore the same for every system. On the dev split each system gets its own fitted temperature. It never changes the chosen option. Only the confidence with which it is claimed. Only the test split counts: 883 decisions on our own use cases, 755 on the benchmarks.
What counts in the end? Three metrics. Accuracy. ECE, the gap between claimed confidence and actual accuracy, over ten bins. And the number that says most for agents: automatable at 5 % error, the share of decisions you can accept unchecked when sorted by confidence, if at most 5 % of them may be wrong.
And Jev itself?
We had no API key from TypeSafe. Jev is therefore only represented through answers others have published: the author of jev-bench, with Jev 1.13.0 and measured latency, and TypeSafe itself for evalsafe. An exact key per record and question matched 450 of the 755 benchmark decisions. Can't you just go through OpenRouter? Unfortunately not. All it offers is typesafe/jev-router, which forwards chat requests to large language models without returning distributions.
What is inside the clones?
Technically there are exactly three designs, and none of them is a new model class. Two are an ordinary language model that returns a distribution over answer options instead of text. The third measures similarity between embeddings. The details come from the installed code, the configurations and the model cards.
Option logits
home-built, JevK5
- 1state, question and options as A), B), C) in one prompt
- 2one forward pass, no generated token
- 3logits of the letters at the last position
- 4softmax yields the distribution
Pointer head
Kev-4B
- 1base model reads state and options
- 2one vector at the decide token, one at the end of each option
- 3dot product per option, learned head
- 4softmax with a fixed temperature
Embedding similarity
CLM, embed-router (Laya similar)
- 1embed state and each option separately
- 2text and question are never read together
- 3cosine between state and option
- 4softmax over the similarities
| System | Base | what was trained | how the probability comes about | Training data |
|---|---|---|---|---|
| home-built | Qwen3.5-4B, unchanged | nothing | softmax over the logits of the answer letters after one forward pass | none |
| JevK5 v0.3 | Qwen3.5-4B, LoRA r=16 on attention, merged | one epoch of cross-entropy on the letter logits | like the home-built one, fixed temperature 1.22, at most 16 options per pass | 17,408 synthetic questions labelled by Qwen3.6-27B and GPT-6 Luna, plus 30,052 items from 26 public train splits |
| Kev-4B | Qwen3.5-4B-Base, LoRA r=16 on all projections, plus a pointer head | cross-entropy in several stages | dot product between decision token and option vector, fixed temperature 2.41 | 10,000 items from ten public datasets plus generated cases |
| Laya multilingual | mmBERT-base, 322 M, encoder | whole encoder plus head, four epochs, "RLCD" | classifier on a [MASK] token in front of each option | not fully published |
| CLM-v0.1-8B | Qwen3-8B, frozen, as embedding model | two small projection heads | cosine between state and option vector | around 60 M question-answer pairs, no classification training |
| embed-router | qwen3-embedding 0.6B | nothing | cosine between state and option description | none |
Three of the four clones are built on Qwen, the fourth, Laya, on mmBERT. That is no accusation. Qwen is the obvious base for this kind of thing right now, and open weights exist for exactly that. It does help to place the names correctly, though: "open-source Jev" in practice almost always means Qwen with a different read-out. What is inside Jev itself, nobody outside TypeSafe knows.
Option logits: the home-built model and JevK5
This is the core of almost every clone, and it is surprisingly simple. State, question and options go into a chat prompt as A), B), C), the model computes a single forward pass, and at the last position you read out the logits of the letter tokens. Softmax over them. Done. Because no token is generated, nothing outside the options can come out.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
MODEL = "Qwen/Qwen3.5-4B" # instruct variant, no training
tok = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForCausalLM.from_pretrained(
MODEL, torch_dtype=torch.bfloat16, device_map="cuda"
)
def decide(state, question, options, temperature=1.0):
letters = [chr(65 + i) for i in range(len(options))]
listing = "\n".join(f"{l}) {o}" for l, o in zip(letters, options))
prompt = tok.apply_chat_template(
[{"role": "user", "content": f"{state}\n\nQuestion: {question}\n{listing}\n\nAnswer with the letter only."}],
tokenize=False, add_generation_prompt=True, enable_thinking=False,
)
ids = tok(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
logits = model(**ids).logits[0, -1] # one forward pass
candidates = [tok.encode(l, add_special_tokens=False)[0] for l in letters]
p = torch.softmax(logits[candidates].float() / temperature, dim=-1)
return dict(zip(options, p.tolist()))
That is the whole principle. What the home-built model in the test does beyond it is craft: score questions, where the distribution over the scale steps becomes an expected value, noul questions as a yes/no pair, several questions in one request and the Jev-compatible interface around it, around 150 lines in total.
JevK5 is exactly this technique with the same base model and a small LoRA, trained on the answers of large language models. One of the two teachers is Qwen3.6-27B, the same model that runs as our large cross-check. At its core JevK5 is a distilled Qwen3.6-27B in a 4B model. What does the training buy? In English 3.5 points, 67.9 against 64.4 %. In German nothing. 82.2 against 83.4 %. Out of the box JevK5 is better calibrated, with an ECE of 0.030 against 0.136. The home-built model catches up with one fitted temperature, though (0.027).
Pointer head: Kev
Kev takes the base model without instruct tuning and learns a head of its own that compares each option with a decision token. Each question is computed in isolation, and up to 255 options are possible. That is a cleaner build than letter logits. It does need training, however, because a base model can't do this task on its own, and in the test Kev lands below the untrained home-built model: 75.4 against 83.4 % in German, 63.0 against 64.4 % on the benchmarks. The built-in temperature of 2.41 makes Kev too unsure on this data, the fit would like 0.59. To be fair, Kev is the best-documented clone, with pre-registered rounds and locked tests, and states itself that its gains on its own suites are "in distribution".
Embedding similarity: CLM, Laya, embed-router
CLM embeds the state and each option separately and compares them by cosine. For choice questions it embeds only the description, for noul it compares "Yes. This is true: …" against "No. This is false: …". That is an embedding router with an 8B encoder. And that's how it behaves. 42.8 % in German, against 42.7 % for the 0.6B embedding router. Both sit below what "always the most frequent answer" achieves, which is 45.1 %. The advertised "Jev parity" comes from ranking tasks in agents, meaning tool calls and games, and partly only holds with specially trained heads. CLM was never trained for classification. Could it be our setup? We checked, and no. The manufacturer's real vLLM setup gives the same result as a rebuild of the encoder without the end token.
Laya sits in between. A small encoder reads the options together with the text and scores each one with a classifier, and at 17 ms that is the fastest solution in the field. Its own README, however, calls the base model "a fast base to specialise, not a zero-shot decision engine" and reports 0.362 zero-shot itself, below the majority baseline of 0.461. The headline "beats Jev" (0.766 against 0.727) applies to a checkpoint retrained on exactly this benchmark, and according to the README no result file for it has been published. Measured: 41.4 % in German, 48.7 % on the benchmarks. With one exception. On ticket types Laya reaches 84 %, the best value of all systems. The domain probably lies close to its training, but that can't be proven.
Have the clones seen the test already?
Partly. Kev trained on the train splits of banking77, boolq, mnli, sst5 and yelp, and these very tasks are in jev-bench and tasksource, there taken from the test splits. Individual items are probably unknown, the tasks and label spaces are not. For JevK5 the list of its 26 train splits is only in the online model card, and an overlap is plausible. For evalsafe there is no indication for any clone. The advantages of Kev and JevK5 on jev-bench are therefore upper rather than lower bounds.
Who comes out ahead?
Among the local systems, three sit close together: JevK5, the home-built model and Kev. The home-built model without training is as good as the trained clones. Laya and CLM land zero-shot at or below the level of "always the most frequent answer".
Own use cases, German (883 decisions)
- Qwen3.6-27B (JSON)94.8 %
- Home-built (Qwen3.5-4B)83.4 %
- JevK5 v0.382.2 %
- Kev-4B75.4 %
- CLM-v0.1-8B42.8 %
- embed-router 0,6B42.7 %
- Laya multilingual41.4 %
┆ always the most frequent answer: 45.1 %
HF benchmarks, English (755 decisions)
- Qwen3.6-27B (JSON)69.7 %
- JevK5 v0.367.9 %
- Home-built (Qwen3.5-4B)64.4 %
- Kev-4B63.0 %
- Laya multilingual48.7 %
- CLM-v0.1-8B33.6 %
- embed-router 0,6B30.5 %
- Reference (Jev, large LLM)
- Clone built on a language model
- Clone built on embedding similarity
On the English benchmarks everyone scores lower. The tasks are harder, and the labels are partly fuzzy. The order of the three 4B models barely changes.
The large language model wins clearly in German, with 94.8 %. But it costs around three seconds per request instead of a hundred milliseconds. Is that worth it? It depends on how often your agent has to decide.
- Reference (Jev, large LLM)
- Clone built on a language model
- Clone built on embedding similarity
One request carries every question of a record. Measured locally, one request after another.
One forward pass instead of generation makes the best 4B models 18 to 34 times faster than the 27B model, 43 to 101 ms against 1.4 to 3.1 seconds. In return they trail by 11 to 13 points in German. For an agent that makes twenty small routing decisions in a loop, that is the difference between two seconds and a minute. Noticeable.
And against Jev?
Jev is only present on 450 of the 755 benchmark decisions. On exactly these 450, the comparison looks like this:
Accuracy
- Jev 1.13.084.9 %
- Qwen3.6-27B (JSON)83.1 %
- JevK5 v0.378.7 %
- Kev-4B76.2 %
- Home-built (Qwen3.5-4B)76.0 %
- Laya multilingual44.0 %
- CLM-v0.1-8B34.9 %
Same answer as Jev
- Jev 1.13.0-
- Qwen3.6-27B (JSON)86 %
- JevK5 v0.382 %
- Kev-4B78 %
- Home-built (Qwen3.5-4B)80 %
- Laya multilingual44 %
- CLM-v0.1-8B39 %
- Reference (Jev, large LLM)
- Clone built on a language model
- Clone built on embedding similarity
Jev only via answers published by third parties (jev-bench, evalsafe), no API access of our own.
On the available cases Jev is the most accurate system, at 84.9 %. The 27B model is close behind, JevK5 six points back. Where do those six points come from?
| Benchmark · type | Jev 1.13.0 | Qwen3.6-27B | JevK5 v0.3 | Kev-4B | Home-built |
|---|---|---|---|---|---|
| evalsafe · choice | 97 | 99 | 93 | 82 | 93 |
| evalsafe · score | 95 | 84 | 65 | 66 | 67 |
| evalsafe · noul | 99 | 99 | 99 | 97 | 97 |
| jev-bench · choice | 80 | 80 | 80 | 74 | 70 |
| jev-bench · score | 49 | 52 | 52 | 51 | 44 |
| jev-bench · noul | 88 | 86 | 86 | 87 | 84 |
Accuracy in percent. The stronger the fill, the higher. Marked is the only cell with a clear lead.
Almost the whole lead comes from a single cell: score questions on evalsafe, TypeSafe's own eval set. There Jev reaches 95 %, the 4B models 65 to 67. On the human-labelled jev-bench, Jev at 72.5 % is level with JevK5 and the 27B model. And on evalsafe-noul, 91 % of the labels are "no". Any system that simply says "no" gets close to 100 % there.
For the evalsafe numbers, keep in mind where the labels come from. They are a consensus of two large models, and the set comes from TypeSafe itself. A model trained on similarly produced labels has an advantage there.
So what is it all for?
The confidence. A distribution instead of a text is the better form for agent decisions, because it gives you a dial: confident decisions go through, uncertain ones go to a human or to a larger model. But that can only work if the claimed confidence is right. Is it?
- as shipped (ECE 0.136)
- after temperature scaling, T = 1.86 (ECE 0.027)
Below the diagonal the model is overconfident. A single number, fitted on the dev split, puts the curve on it.
As shipped, the home-built model is clearly overconfident in English. At a claimed 86 % it hits 64 %. A single number, the temperature, fitted on a few hundred cases of the dev split, puts the curve on the diagonal. That is no more than an afternoon's work. And it decides whether you can trust the number in production, because a threshold on an uncalibrated confidence reliably lets the wrong cases through.
German, own use cases
- Jev 1.13.0-
- Home-built (Qwen3.5-4B)64 %
- JevK5 v0.351 %
- Kev-4B9 %
- Laya multilingual0 %
- CLM-v0.1-8B0 %
English, HF benchmarks
- Jev 1.13.072 %
- Home-built (Qwen3.5-4B)24 %
- JevK5 v0.333 %
- Kev-4B7 %
- Laya multilingual0 %
- CLM-v0.1-8B0 %
- Reference (Jev, large LLM)
- Clone built on a language model
- Clone built on embedding similarity
Share of decisions you can accept when sorted by confidence. Jev on two benchmarks only. The 27B LLM is missing because it returns no distribution.
On the German use cases, the home-built model lets you accept 64 % of decisions without review at no more than 5 % errors, JevK5 51 %. In English it is 24 to 33 %, Jev 72 %, though only on the two benchmarks with published answers. The 27B model returns no distribution and therefore has no such dial. All it has is its accuracy.
How much of it is hype?
The idea behind System One is right. It isn't new. The value lies in interface, price and latency, not in a new kind of model. And most of the clones' headlines don't survive a neutral measurement.
"Cannot hallucinate" only means the answer is always an allowed option. It can still be wrong, even with high confidence. Every language model whose letter logits you read out offers the same guarantee.
Jev's lead is narrow and depends on its own test set. On the independent jev-bench, Jev is level with JevK5 and the 27B model, and the gap comes almost entirely from the score cell on evalsafe, a test set TypeSafe built itself and labelled with the consensus of two large models.
"Open-source Jev" is mostly a known technique under a new name. The strongest clone, JevK5, is the base model plus LoRA, distilled from Qwen3.6-27B and GPT-6 Luna. Compared with 150 lines without training, that buys 3.5 points in English. In German, none.
Comparisons with Jev almost always come from their own suites. Kev measures on its own dev sets. Laya only beats Jev with a checkpoint retrained on the benchmark. By "on par", CLM means ranking in agent games with fine-tuned heads. JevK5 cites a rank the tested version v0.3 never received, and according to its changelog v0.2 even contained test items from MMLU-Pro.
More parameters only help if the design is right. CLM at 8B is exactly as good zero-shot as a 0.6B embedding router, because it does the same thing. It measures similarity instead of weighing the text against the question.
And Jev? Jev itself is good. On the available cases it is the most accurate and best-calibrated system, at 186 ms over the internet. The distance to a local 4B model is just smaller than the announcement suggests.
What does this mean for your own projects?
So do you have to wait for Jev? No. For German decisions that should run locally, an open 4B instruct model with option logits and a temperature fitted on a few hundred labelled cases of your own is enough today. That is the home-built model. JevK5 is an equivalent ready-made alternative, as long as the content is English.
What the clones don't deliver can only come from training on your own domain. The sensible next step is therefore fine-tuning with 200 to 500 checked examples of your own, with Kev, with Laya or as a LoRA on the home-built model. Another zero-shot clone usually brings nothing new.
Jev is worth it as a yardstick and for English, non-critical data. Whether it leads in German, we can only measure with a TypeSafe key. Until then the question stays open.
Where are the limits of this test?
What can you read from these numbers and what not? They are enough for a ranking and for rough distances. Differences of two to three percentage points are not reliable with 62 to 89 test cases per question.
- Jev is only measured indirectly. There are published answers on two benchmarks, none in German. Whether Jev beats the 83 % of the best 4B models on our own use cases is open.
- Our own labels are unchecked. Language models wrote the texts and target answers, nobody proofread them. That probably favours the 27B model, which "thinks" much like the label authors.
- Ticket priorities can hardly be derived from the text. No system gets above 52 %, so this column measures noise more than anything.
- Zero-shot only. According to their makers, Laya and Kev are meant as a base for fine-tuning. How good they are after 200 to 500 examples of your own is not something this test measures.
- The prompts are not optimised. The questions are worded identically for all systems, in German for our own use cases. Models with an English training prompt may lose a little as a result.
- The latencies are lab values. One request after another, no batching, no load. Jev's 186 ms include the trip over the internet and come from the jev-bench author's measurement.
- 4B sizes only. kev-9b and kev-0.8b are registered but were not run.