8 projects · checked 2026-09-20

Evals and benchmarks for decision models

Head-to-head races, task suites and benchmark harnesses. Read the method before the headline: most of these compare on speed and cost, and very few report calibration.

7 of these 8 call the hosted API rather than running open weights.

Find out what your data actually supports

We turn your historical decisions into a frozen test set, then tell you which model, which threshold, and how much of it you can safely automate.