66 comparisons · checked 20 September 2026
Compare decision models
Head to head on the things that can actually be checked: licence, architecture, size, and whether anyone has measured it. Not on accuracy, because for most of these nobody has published a number worth repeating.
How to use these
Start with licence. It rules models out of a product faster than any benchmark and it is the one dimension that is never ambiguous. Then architecture, because it decides whether the confidence number is something you can route on. Size last.
Against the commercial model
Jev is the only model here we have measured end to end, so every one of these has a known quantity on one side and an unknown on the other.
Jev vs laya
421M · apache-2.0
Jev vs Qwen-2.5-1B-RLCD
decoder · apache-2.0
Jev vs cua-s1-forms
mit
Jev vs laya-multilingual
322M · apache-2.0
Jev vs decider-2b
1.9B · decoder · apache-2.0
Jev vs laya-typed-decisions
421M · apache-2.0
Jev vs open-jev-deberta-v3-large
434M · encoder · apache-2.0
Jev vs LFM2.5-350M-RLCD
354M · decoder · other
Jev vs modernbert-ja-310m-jev
315M · encoder · cc-by-sa-4.0
Jev vs JEV-CPU
decoder · mit
Jev vs laya-vision-smolvlm-256m
237M · decoder · cc-by-nc-sa-4.0
Open model against open model
Neither side of these has an independent benchmark. What they do have is a licence, an architecture and a size, and those are enough to narrow a shortlist.
The models in these comparisons
The 12 most-watched, excluding quantised builds and adapters. Comparing a model to its own GGUF conversion is not a comparison.
Jev
TypeSafe's hosted decision model. Returns a typed choice with a full probability distribution instead of text. The only model in this directory we have benchmarked end to end.
laya
The most-liked open-weight decision model of the post-Jev wave, a 421M ModernBERT-large encoder. The base release of a three-model family.
Qwen-2.5-1B-RLCD
Qwen2.5-1.5B-Instruct fine-tuned to return typed decisions.
cua-s1-forms
A System One model narrowed to one job, filling forms. MIT licensed, and the clearest example in the directory of the interface spreading past the benchmark it was born on.
laya-multilingual
The only multilingual decision model we have found, a 322M encoder on mmBERT. Also the one our English-only test set cannot say anything useful about.
decider-2b
The most-downloaded open-weight decision model in the post-Jev wave. A Qwen3.5-2B decoder fine-tuned to emit typed decisions, and the largest member of a three-model family.
laya-typed-decisions
The Laya variant tuned on the typed-decisions benchmark, and the only open model whose card publishes head-to-head numbers against Jev. One of those numbers is the reason we started measuring.
open-jev-deberta-v3-large
A DeBERTa-v3-large encoder positioned as an open reimplementation of Jev. The highest ratio of likes to downloads in the directory, which is usually what a credible claim looks like early.
LFM2.5-350M-RLCD
A typed-decision model from notnotsamuel, 354M parameters. Licensed other, check before commercial use.
modernbert-ja-310m-jev
modernbert-ja-310m fine-tuned to return typed decisions, 315M parameters. Licensed cc-by-sa-4.0, check before commercial use.
JEV-CPU
Qwen3-0.6B fine-tuned to return typed decisions.
laya-vision-smolvlm-256m
SmolVLM-256M-Instruct fine-tuned to return typed decisions, 237M parameters. Licensed cc-by-nc-sa-4.0, check before commercial use.
Find out what your data actually supports
We turn your historical decisions into a frozen test set, then tell you which model, which threshold, and how much of it you can safely automate.