Built with Evals and benchmarks
jev-directory
A browsable directory of what Jev can do — 50 runnable judge-model evals with the exact experimentalevaluate prompt, plus 1,300+ real community builds from the TypeSafe AI #show-an
13 ★ on GitHub
What it is
A browsable directory of what Jev can do — 50 runnable judge-model evals with the exact experimentalevaluate prompt, plus 1,300+ real community builds from the TypeSafe AI #show-and-tell channel, each linked to its project and source post.
Where the decision sits
jev-directory runs its own weights or its own runtime rather than calling a hosted API, so its decisions do not depend on a vendor being up or on a key being funded. The trade is that the model is now yours to keep current.
How active it is
Created 2026-09-18, last pushed 2026-09-19, 13 stars and 1 forks when we read it on 2026-09-20. Everything in this category is weeks old at most, so a stale push date means something different here than it would on a mature project. Check the repository before you depend on it.
What we have not checked
Whether it works. We have not run this project, measured the model behind it, or verified any performance claim in its README. This page records what the repository says about itself, read on 2026-09-20. The one model we have measured end to end is Jev.
Find out what your data actually supports
We turn your historical decisions into a frozen test set, then tell you which model, which threshold, and how much of it you can safely automate.