$ big_bench run --all-models --one-shot
big_bench
a benchmark you can touch: one prompt for every model, live artifacts, real prices. no MMLU.
✓ 12 models · 1 test · 0 auto-retries · $5.96 api
Season standings
Points are awarded by place in each test using the 10, 7, 5, 3, 2, 1 scheme; lower places score nothing. The standings are recomputed after every new test.
| # | model | pts |
|---|---|---|
| 01 | claude-opus-5🏆 | 10 |
| 02 | gpt-5.6-sol | 7 |
| 03 | qwen3.8-max | 5 |
| 04 | gpt-5.5 | 3 |
| 05 | claude-fable-5 | 2 |
| 06 | deepseek-v4-pro | 1 |
| 07 | kimi-k3 | 0 |
| 08 | gemini-3.1-pro | 0 |
| 09 | claude-opus-4-8 | 0 |
| 10 | grok-4.5 | 0 |
| 11 | claude-opus-4-6 | 0 |
| 12 | qwen3.7-max | 0 |
Tests
Jul 11, 20263d The Big Mac test: twelve models, one Big Mac winner: claude-opus-5
coming soon in the works? next test logic? a clock? tetris? — in the works.