← max_tokens

$ big_bench run --all-models --one-shot

big_bench

a benchmark you can touch: one prompt for every model, live artifacts, real prices. no MMLU.

✓ 13 models · 1 test · 0 auto-retries · $8.64 api

Season standings

Points are awarded by place in each test using the 10, 7, 5, 3, 2, 1 scheme; lower places score nothing. The standings are recomputed after every new test.

#modelpts
01 claude-fable-5-1🏆 10
02 claude-opus-5 7
03 gpt-5.6-sol 5
04 qwen3.8-max 3
05 gpt-5.5 2
06 claude-fable-5 1
07 deepseek-v4-pro 0
08 kimi-k3 0
09 gemini-3.1-pro 0
10 claude-opus-4-8 0
11 grok-4.5 0
12 claude-opus-4-6 0
13 qwen3.7-max 0

Tests

Jul 11, 20263d The Big Mac test: twelve models, one Big Mac winner: claude-fable-5-1
coming soon in the works? next test logic? a clock? tetris? — in the works.