Evaluation lab

Benchmarks, in the open.

A small, repeatable view of the latest OurToken model canaries. Scores are the percentage of fixed test cases answered correctly.

Current canaryMMLU-Pro + IFEval0 models with a completed run

No published results

There are no completed model runs yet.

Run a canary from the admin Evaluation Lab to publish the first result.