← max_tokens

$ big_bench run --all-models --one-shot

big_bench

a benchmark you can touch: one prompt for every model, live artifacts, real prices. no MMLU.

12 models · 1 test · 0 auto-retries · $5.96 api

Season standings

Points are awarded by place in each test using the 10, 7, 5, 3, 2, 1 scheme; lower places score nothing. The standings are recomputed after every new test.

#modelpts
01 claude-opus-5🏆 10
02 gpt-5.6-sol 7
03 qwen3.8-max 5
04 gpt-5.5 3
05 claude-fable-5 2
06 deepseek-v4-pro 1
07 kimi-k3 0
08 gemini-3.1-pro 0
09 claude-opus-4-8 0
10 grok-4.5 0
11 claude-opus-4-6 0
12 qwen3.7-max 0

Tests

Jul 11, 20263d The Big Mac test: twelve models, one Big Mac winner: claude-opus-5
coming soon in the works? next test logic? a clock? tetris? — in the works.