← big_bench

The Big Mac test: twelve models, one Big Mac

One prompt, no edits: whatever came back is published — live sandboxed artifacts (drag to rotate), metrics and prices.

The prompt

promptCreate a 3D visualization of a Big Mac in three.js. The burger slowly rotates on its own; the camera can orbit and zoom. Make it as photorealistic as possible.

Technical requirements: a single self-contained HTML file; load three.js from a CDN; the scene fills the entire browser window and adapts to its size; no localStorage, cookies, or network requests beyond CDN libraries. Return only the finished HTML file as a single code block, with no text before or after.

Results

Score is a weighted blend: 0.4 big-mac-ness, 0.3 render quality, 0.2 interactivity, 0.1 first-run success. Points are awarded by place in this test using the 10, 7, 5, 3, 2, 1 scheme; lower places score nothing. A tie does not split a place: the cheaper run goes higher, then the faster one, and the last key is the snapshot id alphabetically.

#modelscoreptscostvaluetime
01 claude-fable-5-1🏆 0.880 +10 $2.68 0.3 653s
02 claude-opus-5 0.875 +7 $1.31 0.7 617s
03 gpt-5.6-sol 0.840 +5 $0.36 2.3 141s
04 qwen3.8-max 0.830 +3 $0.22 3.8 751s
05 gpt-5.5 0.800 +2 $0.63 1.3 238s
06 claude-fable-5 0.745 +1 $0.98 0.8 234s
07 deepseek-v4-pro 0.740 — $0.03 25.9 403s
08 kimi-k3 0.705 — $0.53 1.3 1162s
09 gemini-3.1-pro 0.640 — $0.16 3.9 99s
10 claude-opus-4-8 0.635 — $0.15 4.3 57s
11 grok-4.5 0.615 — $0.07 8.7 68s
12 claude-opus-4-6 0.530 — $1.49 0.4 864s
13 qwen3.7-maxdidn't render 0.000 — $0.03 0.0 149s
total $8.64 90:36

value = score ÷ cost (single run) — a rough “score per dollar.” score is half-subjective and cost varies between runs, so treat it as a hint, not a hard metric. The gap is real, though: deepseek lands 0.14 behind the winner at ~a fortieth of the cost.

Compare two

Defaults to the winner vs the best value; pick any pair from the menus.

run conditions and versions

13/13 responded · 353.6k tok · $8.64 · 96:49 wall

one-shot · single-delivery · temperature default · reasoning high · max_tokens uncapped · 2026-07-11

how score is computed. The owner sets big-mac-ness and render quality by eye, on a 0–1 rubric. A headless-browser probe produces first-run success and interactivity (loaded without errors + responds to drag). The subjective part is disclosed honestly: it is an author’s judgment, not an “objective” number — so the leaderboard order and the value column inherit that subjectivity.

re-run by the owner: the first attempt produced no working artifact (no usable response / truncated / broken), this result is a re-attempt, single-delivery held per attempt: kimi-k3, claude-opus-5

snapshots: claude-opus-4-8=anthropic/claude-opus-4.8 · claude-opus-4-6=anthropic/claude-opus-4.6 · claude-fable-5=anthropic/claude-fable-5 · gpt-5.5=openai/gpt-5.5 · gpt-5.6-sol=openai/gpt-5.6-sol · gemini-3.1-pro=google/gemini-3.1-pro-preview · grok-4.5=x-ai/grok-4.5 · qwen3.7-max=qwen/qwen3.7-max · deepseek-v4-pro=deepseek/deepseek-v4-pro · kimi-k3=moonshotai/kimi-k3 · claude-opus-5=anthropic/claude-opus-5 · qwen3.8-max=qwen/qwen3.8-max · claude-fable-5-1=anthropic/claude-fable-5.1