01 claude-fable-5-1 winner 0.880
| tokens | cost | time | code | file |
|---|---|---|---|---|
| 162 → 53.6k | $2.68 | 653s | 423 loc | 24 KB |
One prompt, no edits: whatever came back is published — live sandboxed artifacts (drag to rotate), metrics and prices.
promptCreate a 3D visualization of a Big Mac in three.js. The burger slowly rotates on its own; the camera can orbit and zoom. Make it as photorealistic as possible.
Technical requirements: a single self-contained HTML file; load three.js from a CDN; the scene fills the entire browser window and adapts to its size; no localStorage, cookies, or network requests beyond CDN libraries. Return only the finished HTML file as a single code block, with no text before or after. Score is a weighted blend: 0.4 big-mac-ness, 0.3 render quality, 0.2 interactivity, 0.1 first-run success. Points are awarded by place in this test using the 10, 7, 5, 3, 2, 1 scheme; lower places score nothing. A tie does not split a place: the cheaper run goes higher, then the faster one, and the last key is the snapshot id alphabetically.
| # | model | score | pts | cost | value | time |
|---|---|---|---|---|---|---|
| 01 | claude-fable-5-1🏆 | 0.880 | +10 | $2.68 | 0.3 | 653s |
| 02 | claude-opus-5 | 0.875 | +7 | $1.31 | 0.7 | 617s |
| 03 | gpt-5.6-sol | 0.840 | +5 | $0.36 | 2.3 | 141s |
| 04 | qwen3.8-max | 0.830 | +3 | $0.22 | 3.8 | 751s |
| 05 | gpt-5.5 | 0.800 | +2 | $0.63 | 1.3 | 238s |
| 06 | claude-fable-5 | 0.745 | +1 | $0.98 | 0.8 | 234s |
| 07 | deepseek-v4-pro | 0.740 | — | $0.03 | 25.9 | 403s |
| 08 | kimi-k3 | 0.705 | — | $0.53 | 1.3 | 1162s |
| 09 | gemini-3.1-pro | 0.640 | — | $0.16 | 3.9 | 99s |
| 10 | claude-opus-4-8 | 0.635 | — | $0.15 | 4.3 | 57s |
| 11 | grok-4.5 | 0.615 | — | $0.07 | 8.7 | 68s |
| 12 | claude-opus-4-6 | 0.530 | — | $1.49 | 0.4 | 864s |
| 13 | qwen3.7-maxdidn't render | 0.000 | — | $0.03 | 0.0 | 149s |
| total | $8.64 | 90:36 |
value = score ÷ cost (single run) — a rough “score per dollar.” score is half-subjective and cost varies between runs, so treat it as a hint, not a hard metric. The gap is real, though: deepseek lands 0.14 behind the winner at ~a fortieth of the cost.
Defaults to the winner vs the best value; pick any pair from the menus.
01 claude-fable-5-1 winner 0.880
| tokens | cost | time | code | file |
|---|---|---|---|---|
| 162 → 53.6k | $2.68 | 653s | 423 loc | 24 KB |
02 claude-opus-5 0.875
| tokens | cost | time | code | file |
|---|---|---|---|---|
| 160 → 52.4k | $1.31 | 617s | 806 loc | 32 KB |
The full spec, finally: a three-part bun with real crust mottling and properly seated sesame, two grained patties, cheese with actual thickness, shredded lettuce, onion, pickles. Best materials of the twelve — and the only entry whose garnish explodes: lettuce shards and cheese wings stick out a third past the bun. 10 minutes and $1.31, second-priciest of the run. Two disclosures: it was delivered on the second attempt (the first stream was cut at ~100s by the transport, not by the model — one of the run's two reruns), and its two subjective marks were set on the owner's behalf rather than by his own eye.
03 gpt-5.6-sol 0.840
| tokens | cost | time | code | file |
|---|---|---|---|---|
| 108 → 12.0k | $0.36 | 141s | 739 loc | 25 KB |
The clearest Big Mac of the first ten: two patties, a proper middle bun, white-sauce specks — and the most convincing light of them. Won the first ten on structure, not just polish.
04 qwen3.8-max 0.830
| tokens | cost | time | code | file |
|---|---|---|---|---|
| 149 → 36.8k | $0.22 | 751s | 509 loc | 22 KB |
One version on from the sibling that failed: qwen3.7-max pulled dead CDN paths, qwen3.8-max ships an import map for three@0.165.0 and ESM addons — and it runs. Third place. Big-mac-ness level with claude-opus-5 (0.80): three-part bun, two patties, lettuce on both levels, honest big sesame, onion crumb and sauce dots — the cheese is even modelled with a melted droop, but it is buried under lettuce sheets grown too wide, and pickles are the one thing missing. The render (0.70) is plasticine rather than meat: near-black patties, mint-green lettuce, sesame in blobs. $0.22 — cheaper than everything above it. 12.5 minutes, third-slowest of the twelve.
05 gpt-5.5 0.800
| tokens | cost | time | code | file |
|---|---|---|---|---|
| 108 → 21.1k | $0.63 | 238s | 707 loc | 25 KB |
Same double-decker-plus-sauce reading as gpt-5.6-sol, a step behind on lighting and materials. Solid, unspectacular.
06 claude-fable-5 0.745
| tokens | cost | time | code | file |
|---|---|---|---|---|
| 160 → 19.5k | $0.98 | 234s | 364 loc | 15 KB |
Gets the two-patty, middle-bun structure right — the lettuce is flat, stylized sheets rather than leaves. $0.98 for fifth, pricier than both GPTs above it.
07 deepseek-v4-pro 0.740
| tokens | cost | time | code | file |
|---|---|---|---|---|
| 106 → 22.3k | $0.03 | 403s | 793 loc | 38 KB |
Mid-table (0.740) for $0.03 — about a tenth of gpt-5.6-sol's $0.36. A double-decker on a plate, fillings a touch sparse. The cheapest strong result of the run.
08 kimi-k3 0.705
| tokens | cost | time | code | file |
|---|---|---|---|---|
| 186 → 35.1k | $0.53 | 1162s | 475 loc | 20 KB |
The slowest by far: 19 minutes of reasoning before a single line of HTML — it timed out on the first run at the standard 15-minute budget and only finished on a doubled one (one of the run's two reruns). Nails the Big Mac structure — two patties, a middle bun, sesame top — but dark wood-like patties and sparse lettuce keep it mid-table. $0.53. The one entry with a doubled time budget; its two subjective marks were also set on the owner's behalf rather than by his own eye.
09 gemini-3.1-pro 0.640
| tokens | cost | time | code | file |
|---|---|---|---|---|
| 102 → 13.5k | $0.16 | 99s | 478 loc | 23 KB |
Clean single-patty burger, but rendered small in the frame — asked to fill the window, it didn't quite. Fastest of the mid-pack at 99s.
10 claude-opus-4-8 0.635
| tokens | cost | time | code | file |
|---|---|---|---|---|
| 160 → 5.9k | $0.15 | 57s | 368 loc | 12 KB |
The best bun of the first ten — real sesame texture, soft shading — around a single-patty cheeseburger, not a Big Mac. Craft high, big-mac-ness low. Also the fastest: 57s.
11 grok-4.5 0.615
| tokens | cost | time | code | file |
|---|---|---|---|---|
| 308 → 11.7k | $0.07 | 68s | 328 loc | 9 KB |
A double-decker, but geometric and bare — no lettuce, no sauce. Second-cheapest of the renderers at $0.07, behind deepseek-v4-pro's $0.03.
12 claude-opus-4-6 0.530
| tokens | cost | time | code | file |
|---|---|---|---|---|
| 114 → 59.6k | $1.49 | 864s | 334 loc | 14 KB |
The cautionary tale: 14 minutes and 59.6k tokens — the most expensive run ($1.49) — for a plain domed bun with no visible patties. Effort ≠ result.
13 qwen3.7-max 0.000
| tokens | cost | time | code | file |
|---|---|---|---|---|
| 111 → 8.2k | $0.03 | 149s | 660 loc | 26 KB |
Didn't render — THREE is not defined. It pulled three.js from dead CDN paths (the old global three.min.js and examples/js/OrbitControls.js, both 404 since three.js dropped them around r150). A real signal: stale knowledge of how three.js ships today.
13/13 responded · 353.6k tok · $8.64 · 96:49 wall
one-shot · single-delivery · temperature default · reasoning high · max_tokens uncapped · 2026-07-11
how score is computed. The owner sets big-mac-ness and render quality by eye, on a 0–1 rubric. A headless-browser probe produces first-run success and interactivity (loaded without errors + responds to drag). The subjective part is disclosed honestly: it is an author’s judgment, not an “objective” number — so the leaderboard order and the value column inherit that subjectivity.
re-run by the owner: the first attempt produced no working artifact (no usable response / truncated / broken), this result is a re-attempt, single-delivery held per attempt: kimi-k3, claude-opus-5
snapshots: claude-opus-4-8=anthropic/claude-opus-4.8 · claude-opus-4-6=anthropic/claude-opus-4.6 · claude-fable-5=anthropic/claude-fable-5 · gpt-5.5=openai/gpt-5.5 · gpt-5.6-sol=openai/gpt-5.6-sol · gemini-3.1-pro=google/gemini-3.1-pro-preview · grok-4.5=x-ai/grok-4.5 · qwen3.7-max=qwen/qwen3.7-max · deepseek-v4-pro=deepseek/deepseek-v4-pro · kimi-k3=moonshotai/kimi-k3 · claude-opus-5=anthropic/claude-opus-5 · qwen3.8-max=qwen/qwen3.8-max · claude-fable-5-1=anthropic/claude-fable-5.1