Evaluation lab
Benchmarks, in the open.
A small, repeatable view of the latest OurToken model canaries. Scores are the percentage of fixed test cases answered correctly.
Current canaryMMLU-Pro + IFEval0 models with a completed run
No published results
There are no completed model runs yet.
Run a canary from the admin Evaluation Lab to publish the first result.