← big_bench

The Big Mac test: twelve models, one Big Mac

One prompt, no edits: whatever came back is published — live sandboxed artifacts (drag to rotate), metrics and prices.

The prompt

promptCreate a 3D visualization of a Big Mac in three.js. The burger slowly rotates on its own; the camera can orbit and zoom. Make it as photorealistic as possible.

Technical requirements: a single self-contained HTML file; load three.js from a CDN; the scene fills the entire browser window and adapts to its size; no localStorage, cookies, or network requests beyond CDN libraries. Return only the finished HTML file as a single code block, with no text before or after.

Results

Score is a weighted blend: 0.4 big-mac-ness, 0.3 render quality, 0.2 interactivity, 0.1 first-run success. Points are awarded by place in this test using the 10, 7, 5, 3, 2, 1 scheme; lower places score nothing. A tie does not split a place: the cheaper run goes higher, then the faster one, and the last key is the snapshot id alphabetically.

#modelscoreptscostvaluetime
01 claude-opus-5🏆 0.875 +10 $1.31 0.7 617s
02 gpt-5.6-sol 0.840 +7 $0.36 2.3 141s
03 qwen3.8-max 0.830 +5 $0.22 3.8 751s
04 gpt-5.5 0.800 +3 $0.63 1.3 238s
05 claude-fable-5 0.745 +2 $0.98 0.8 234s
06 deepseek-v4-pro 0.740 +1 $0.03 25.9 403s
07 kimi-k3 0.705 $0.53 1.3 1162s
08 gemini-3.1-pro 0.640 $0.16 3.9 99s
09 claude-opus-4-8 0.635 $0.15 4.3 57s
10 grok-4.5 0.615 $0.07 8.7 68s
11 claude-opus-4-6 0.530 $1.49 0.4 864s
12 qwen3.7-maxdidn't render 0.000 $0.03 0.0 149s
total $5.96 79:43

value = score ÷ cost (single run) — a rough “score per dollar.” score is half-subjective and cost varies between runs, so treat it as a hint, not a hard metric. The gap is real, though: deepseek lands 0.14 behind the winner at ~a fortieth of the cost.

Compare two

Defaults to the winner vs the best value; pick any pair from the menus.

run conditions and versions

12/12 responded · 299.8k tok · $5.96 · 84:42 wall

one-shot · single-delivery · temperature default · reasoning high · max_tokens uncapped · 2026-07-11

how score is computed. The owner sets big-mac-ness and render quality by eye, on a 0–1 rubric. A headless-browser probe produces first-run success and interactivity (loaded without errors + responds to drag). The subjective part is disclosed honestly: it is an author’s judgment, not an “objective” number — so the leaderboard order and the value column inherit that subjectivity.

re-run by the owner: the first attempt produced no working artifact (no usable response / truncated / broken), this result is a re-attempt, single-delivery held per attempt: kimi-k3, claude-opus-5

snapshots: claude-opus-4-8=anthropic/claude-opus-4.8 · claude-opus-4-6=anthropic/claude-opus-4.6 · claude-fable-5=anthropic/claude-fable-5 · gpt-5.5=openai/gpt-5.5 · gpt-5.6-sol=openai/gpt-5.6-sol · gemini-3.1-pro=google/gemini-3.1-pro-preview · grok-4.5=x-ai/grok-4.5 · qwen3.7-max=qwen/qwen3.7-max · deepseek-v4-pro=deepseek/deepseek-v4-pro · kimi-k3=moonshotai/kimi-k3 · claude-opus-5=anthropic/claude-opus-5 · qwen3.8-max=qwen/qwen3.8-max