Built with Evals and benchmarks
jev-capability-atlas
Independent, evidence-based map of when TypeSafe's Jev actually holds up vs. breaks down — real API-call receipts, not a leaderboard. 中文為主的雙語 repo。
21 ★ on GitHub
What it is
獨立、非官方、非 TypeSafe 贊助的社群專案。 這個 repo 存在的目的,是幫你、我、還有任何想用 Jev(TypeSafe 的 System One 模型)的人,正確識別手上的任務適不適合交給它——用真實 API 呼叫的收據,畫出它「校準決策」這個宣稱在哪裡站得住、在哪裡站不住的邊界地圖。給人看怎麼用,給 agent 帶進專案看哪裡能試著換上去,也給任何做過真實驗的人一個把結果貢獻進來的地方。
Where the decision sits
jev-capability-atlas calls the hosted decision API rather than running its own weights. That keeps the project small and means its behaviour moves whenever the vendor ships a new model version. It also means an API key and a network round trip sit in the path of every decision.
Models it references: Jev.
How active it is
Created 2026-09-18, last pushed 2026-09-20, 21 stars and 3 forks when we read it on 2026-09-20. Everything in this category is weeks old at most, so a stale push date means something different here than it would on a mature project. Check the repository before you depend on it.
What we have not checked
Whether it works. We have not run this project, measured the model behind it, or verified any performance claim in its README. This page records what the repository says about itself, read on 2026-09-20. The one model we have measured end to end is Jev.
Find out what your data actually supports
We turn your historical decisions into a frozen test set, then tell you which model, which threshold, and how much of it you can safely automate.