Evaluation lab
Benchmarks, in the open.
A small, repeatable view of the latest OurToken model canaries. Scores are the percentage of fixed test cases answered correctly.
Current canaryMMLU-Pro + IFEval2 models with a completed run
Latest results
Model scores
DeepSeek V4 FlashDeepseek ·
deepseek/deepseek-v4-flash80/ 1008/10cases passedGPT 5.6 LunaOpenai · openai/gpt-5.6-luna100/ 10010/10cases passedResults combine the latest completed MMLU-Pro and IFEval starter canaries available for each model. This is directional evidence, not an official full-suite ranking. Last updated Sep 21, 2026, 11:00 AM.