The answer
- Best drawing: GPT-6 Astra at
max- $2.29 and 13.8 min for one laptop with an Apple-style menu bar, a dock of recognisable apps, a notch and side ports - everything the prompt asked for, with no visible assembly error, at half the tokens and price of the Claude Fable 5.1 cell it dethroned on 2026-09-05. Previous holder: Claude Fable 5.1 atmax, for $4.95 - still the best Claude drawing here, with one corner defect Astra does not have. - Best value: Muse Spark 1.3 at
default- $0.035, a 65x smaller bill for a laptop with a real menu bar, a browser window and a joined hinge. Its tells are a glow spill past the lid and up to a minute of silence before the first token - Meta's cheap tier, reachable only through OpenRouter. - The trap: usually the budget, not the model. 9 MacBook runs billed $12.86 and returned no drawing - and $10.26 of that was our own output budget cutting a model off mid-thought. Every capped cell we re-ran at the model's real ceiling then finished. Claude Opus 5.5 at max is the exception that ends that rule: given the whole 128,000-token ceiling the streaming API allows, it spent every one of them thinking and returned an empty file. It is not a runaway - handed 300,000 on the Batch API it finished, and needed 255,957. What effort itself buys is ornament, sometimes nothing, never an understanding of the object.
26 models170 runs$54.50 measuredupdated 2026-09-23
One API call per cell, first answer kept, providers' own billing. The MacBook is the test - everyone knows what one looks like, so a 5% error shows. The pelican is the control every model passes. Order is our quality rank, judged by eye; sort by newest or price instead.
GPT-6 Astra #
The new best MacBook in the benchmark, at max for $2.29 - a real menu bar, a dock of recognisable apps, a notch, ports, and not one visible assembly error - at half the price and time of the drawing it dethrones.
More
The first model to take the top seat off the Fables, and it does it on the criterion this page actually grades: assembly. All four of its cells - even default at $0.33 - draw a laptop whose lid and base share a hinge and a vanishing point, with nothing floating, torn, or overhanging; no other ladder in the catalog manages that. At max it renders what looks like a product shot: an Apple-style menu bar with status icons and the 9:41 clock, a dock of individually recognisable apps, a notch, side ports, speaker grilles, a full keyboard with function row - 45.8K tokens, 13.8 minutes, $2.29, roughly half the tokens, time and money of Fable 5.1's dethroned best. The pelican at max is a children's-book illustration: saddle, bell, brake lever, fenders, a chain that actually wraps the sprocket. The effort dial is cleanly ordered on both tasks - the first frontier model here where paying more reliably buys more. OpenAI launched it behind a limited-access program on 2026-09-04; our access opened the next morning and it went straight through smoke, tokenizer verification (o200k, ratio 1.0000) and this ladder in the same hour.
- default - Perfect assembly out of the gate: joined hinge, one vanishing point, keyboard, trackpad, ports - no menu bar or dock yet, but nothing wrong.
- xhigh - A dock with a badge on the Mail icon, a dated menu bar, ports and grilles - $0.98 for a drawing that would have led this page a week ago.
- max - The new best drawing in the benchmark: an Apple-style menu bar with status icons and the 9:41 clock, a dock of recognisable apps, notch, side ports, full keyboard - and not one visible assembly error. Half the tokens, time and price of the Fable 5.1 cell it dethrones.
In Playcode since 2026-09-05 · runs 2026-09-05
The best drawing per dollar on this page by a distance: a full Finder menu bar, the 9:41 clock, a dock of real app icons, a notch and a keyboard whose keys are individually labelled - for 39 cents, a sixth of what the drawing above it cost.
More
The GPT-6 price cut lands on the benchmark the same day it lands on the API, and it rearranges the top of this page on value rather than on quality. Its max cell is a genuine product shot - an Apple menu bar with named menus, wifi and battery glyphs and the 9:41 clock, a dock of individually recognisable apps, a notch, perforated grilles, a trackpad with its lip, and a full keyboard where the letter, modifier and function keys all carry legible legends - and it costs $0.386 against the $2.29 GPT-6 Astra spent on the drawing that still leads. Astra keeps rank 1 by a hair, on side ports Sol does not draw. What makes this more than one good cell is the ladder: all four assemble correctly, including the cheapest at 3.6 cents, which still has a notch, a real key grid, grilles, a trackpad and a lid and base that share one vanishing point. Nothing floats, tears or overhangs anywhere in the set, and the dial is cleanly ordered on both tasks - 3.6K tokens at default rising to 38.5K at max. Its pelican is a tidy, correct illustration at every effort. Half the price of GPT-5.6 Sol, which it replaces on the Playcode Pro tier, and eleven places above it here.
- default - Three and a half cents, 48 seconds, and nothing wrong with it: notch and camera, a real key grid, perforated grilles both sides, a trackpad, and a lid and base that share one vanishing point. No menu bar or dock - the screen is bare wallpaper - which is the only thing separating this from cells costing fifty times more.
- high - The furniture arrives: an Apple menu bar with Finder, File, Edit and View, a wifi glyph and the 9:41 clock, a dock of coloured app tiles over a mountain wallpaper. Body, keyboard, grilles and trackpad all still correct.
- xhigh - A fuller menu bar (through Window and Help), rounded dock icons with distinguishable glyphs, and keys that begin to carry legends. The only blemish is the Help menu label running into the wallpaper behind it.
- max - The product shot, for 39 cents. A full Finder menu bar with wifi and battery glyphs and the Mon 9:41 clock; a dock of individually recognisable apps; a notch; perforated grilles; a trackpad with its lip; and a keyboard whose letter, modifier and function keys are all legibly labelled. Only side ports are missing against the drawing that still leads this page at six times the price.
In Playcode since 2026-09-22 · runs 2026-09-23
Its xhigh draws a product shot - real Finder menu bar, 9:41 clock, recognisable dock, notch, ports, wordmark - for $1.72. Its max cannot finish inside the 128,000 tokens the streaming API allows, and returns an empty file; handed 300,000 on the Batch API it produces the best drawing here and needs 255,957 of them.
Spent all 128,000 output tokens - the whole ceiling the streaming API allows this model - on thinking, and returned an empty file. Nothing arrived until the stream closed 20 minutes in. $2.56 for no drawing. Re-running it on the same API cannot help: 128,001 is rejected by name. The row below is the same request given 300,000 tokens on the Batch API, and it needed 255,957 of them.
More
Released 2026-09-21 and benchmarked the next day, and it rearranges the top of this page without taking it. Its xhigh cell is a product shot: an Apple menu bar with named menus, status icons and the 9:41 clock, a dock of individually recognisable apps, a notch, side ports, speaker grilles, the MacBook Pro wordmark on the chin, a keyboard that recedes with the deck, and one vanishing point shared by lid and base. That is detail parity with the rank-1 Astra drawing - it even has the chin wordmark Astra omits - and it costs $1.72 against Astra's $2.29 and a third of the $4.95 Fable 5.1 cell it matches. Astra keeps the top seat on two things: legible legends on every key, and a ladder where all four cells land. Cheap is where Opus 5.5 is most convincing: high is $0.40 for a correctly assembled laptop with grilles, a real key grid and a trackpad - the cell Opus 5 gets faulted for, drawn right - at 21% less cost and 32% less time than Opus 5 spent on its near-identical version, which is Anthropic's claimed 30% speed-up showing up in an independent measurement. The ladder then breaks at the top, and expensively. At max it spent all 128,000 output tokens - its own ceiling, not our sweep budget - on thinking, emitted nothing until the stream closed, and billed $2.56 for an empty file. Fable 5.1 failed this way against a 64K budget and finished when given its real ceiling; Opus 5.5 had its real ceiling from the start, so there is no larger number to re-run it with. The pelican dial is ordered and cheap: $0.05 buys a complete bird and bike with no saddle and a stub of a chain, $0.13 buys a finished illustration - chain wrapped over both sprockets, cranks, legs on the pedals, motion lines, sun. A postscript on that empty max cell: it was never a runaway. Re-run on the Batch API, the one surface that allows this model more than 128,000 output tokens, the identical request finished naturally at 255,957 - and drew the most complete laptop in this catalog, menu bar and clock and dock and side ports and labelled keys and glow on the desk. The model is not failing the task at max effort; the streaming API simply cannot hold the answer. That drawing costs $2.56 at the Batch API's half price, arrives asynchronously hours later, and is therefore not something you can ask for in a chat - which is why it sits beside the empty cell rather than replacing it.
- high - Forty cents for a laptop with nothing wrong with it: notch, menu bar, a dock of app tiles, perforated grilles both sides, a real key grid and a trackpad, lid and deck at one eye height. Opus 5 spent 20,145 tokens and 256 seconds on the same drawing and got a skewed base; this is 19,767 and 173.
- xhigh - The product shot, and the best drawing on this page under $2. A real Finder menu bar with named menus, status icons and the 9:41 clock; a dock of individually recognisable apps; notch, side ports, grilles, the MacBook Pro wordmark, a keyboard that recedes with the deck. 86,096 output tokens against 28.5 KB of SVG - about 57K of it thinking.
- max - Spent all 128,000 output tokens - the whole ceiling the streaming API allows this model - on thinking, and returned an empty file. Nothing arrived until the stream closed 20 minutes in. $2.56 for no drawing. Re-running it on the same API cannot help: 128,001 is rejected by name. The row below is the same request given 300,000 tokens on the Batch API, and it needed 255,957 of them.
- max - DIFFERENT METHOD, and the only cell on this page that is: the Batch API with the output-300k-2026-03-24 beta, which is the one surface that allows this model more than 128,000 output tokens. It finished naturally at 255,957 - twice what the streaming API permits - and it is arguably the best drawing in the catalog: a full Apple menu bar with status icons and the Tue 9:41 clock, a dock of individually recognisable apps, a notch, side ports, finely perforated grilles, a legibly-labelled keyboard with function row, and screen glow bleeding colour onto the desk. Billed at the Batch API's 50% discount ($2.56; the same tokens on the standard card would be $5.12), asynchronously, over about two and a half hours - so it is not something you can ask for in a chat.
In Playcode since 2026-09-22 · runs 2026-09-22, 2026-09-23
Claude Fable 5.1 #
The best Claude MacBook, at max for $4.95 - real perspective, screen glow, notch, dock, ports. Dethroned 2026-09-05 by GPT-6 Astra, which matches the detail with cleaner assembly at half the tokens and price. Not flawless: the front-left corner of the base has a notch where three faces meet; high ($0.90) is a wedge with a spike.
At the first sweep's 16K cap: thinking took almost all of it and 1.5 KB of SVG was cut off mid-tag. Fable 5 finished xhigh inside the same cap.
Cut off by OUR 64K sweep budget: 12.4 minutes of thinking, not one byte of SVG, $3.20 billed. Re-run at the model ceiling it finished - the wall was ours.
At the 16K cap with zero visible output - the entire budget went to thinking. On 5.1 the old cap is not a budget, it is a wall.
Cut off by OUR 64K sweep budget: 64,000 tokens of thinking in 12.4 minutes, no SVG, $3.20 for nothing. Given its own 128K ceiling the same request finished in 99,072 tokens.
More
Anthropic's newest model, and the one that broke this benchmark's budget rule. On launch day its xhigh and max cells returned nothing: thinking is always on for 5.1, and both spent our whole 64K sweep budget reasoning without emitting a byte. Re-run at the model's own 128K ceiling, both finished - and max is the best drawing on this page. Not an isometric diagram like Fable 5's: a render in real perspective, with a notch, a menu bar, a dock, the MacBook Pro wordmark, speaker grilles, three side ports, a keyboard that recedes correctly and screen glow spilling onto the desk. The only visible error is a trackpad sitting slightly right of the keyboard centre. It costs $4.95 and 19 minutes; xhigh is $4.34 for a thicker, slabbier body seen from the other side; high is $0.90 for most of the same laptop with the base a few degrees off the lid. The first sweep's 16K cap is a wall for it - thinking alone exhausts it - which is why the cheap tier is not where you use this model. The pelican ladder finished at every effort: at high the head floats free of the body, at xhigh it is a proper bird leaning over the bars, and at max it is the cleanest pelican in the benchmark for $1.92.
- high - Uncapped (64K): silver body, notch, menu bar, dock, grilles, a perspective keyboard - and a body the checklist takes apart. The base is a wedge rather than a slab, it ends in a spike at the front-right corner, and its footprint does not match the lid. Read by eye this passed for "a few degrees of skew"; three separate checklist runs all named the wedge and the spike.
- xhigh - At the model ceiling (128K): finished in 17.2 minutes and 86,716 tokens. The same silver three-quarter body seen from the other side, with a menu bar, a dock and side ports - but a thick, slab-like base and a wide empty palm rest. The 5.1 cell that changed the benchmark: at our old 64K budget this run drew nothing.
- max - At the model ceiling (128K): 99,072 tokens, 19 minutes, $4.95 - and the best drawing in the benchmark. Perspective, not isometry: notch, menu bar, dock, wordmark, grilles, three side ports, a keyboard in perspective and screen glow spilling onto the desk. The trackpad sits slightly right of the keyboard centre; nothing else is visibly wrong.
- high @16K - At the first sweep's 16K cap: thinking took almost all of it and 1.5 KB of SVG was cut off mid-tag. Fable 5 finished xhigh inside the same cap.
- xhigh @64K - Cut off by OUR 64K sweep budget: 12.4 minutes of thinking, not one byte of SVG, $3.20 billed. Re-run at the model ceiling it finished - the wall was ours.
- xhigh @16K - At the 16K cap with zero visible output - the entire budget went to thinking. On 5.1 the old cap is not a budget, it is a wall.
- max @64K - Cut off by OUR 64K sweep budget: 64,000 tokens of thinking in 12.4 minutes, no SVG, $3.20 for nothing. Given its own 128K ceiling the same request finished in 99,072 tokens.
In Playcode since 2026-09-01 · runs 2026-09-01, 2026-09-02
Claude Fable 5 #
The steadiest geometry here - no run found a broken join at any effort - but isometric, and the screen stays a gradient: no menu bar at any level, no glow below max. $2.13 and ten minutes at max.
More
The most careful draughtsman here. Its xhigh MacBook won the launch sweep, and at max it draws the only laptop on this page with no visible error at all - notch, dock, grilles, side ports, every edge meeting where it should. What keeps it second is the projection: it draws a flat isometric diagram, where the prompt asks for a perspective render with screen glow, and 5.1 delivers exactly that. $2.13 and nearly ten minutes for one laptop; half the price of its successor's best cell, and half the ambition.
- xhigh - Winner of the launch sweep - the most detailed MacBook anyone drew inside the 16K cap.
- max - Uncapped (64K, streamed): the best drawing in the benchmark - isometric, dock, side ports - for $2.13 and 9.8 minutes.
In Playcode since 2026-07-10 · runs 2026-07-13, 2026-07-21
GPT-5.6 Sol #
At xhigh a genuinely good laptop, more dark Lenovo than MacBook; at max on this task it never returns at all.
Timed out. It did not return in 36 minutes, even with a 64K budget - exactly as in the first run.
More
Superseded by GPT-6 Sol, which costs half as much and ranks fifteen places higher - see rank 2. The most ambitious - more gradients, more perspective, more parts than anyone. At high that ambition bends the geometry; at xhigh it pulls together into a genuinely good laptop, admittedly more dark Lenovo than MacBook. Its party trick is the failure: it finishes the pelican at max, but on the MacBook it never returns at all.
- max - Timed out. It did not return in 36 minutes, even with a 64K budget - exactly as in the first run.
In Playcode since 2026-07-09 · runs 2026-07-13, 2026-07-21
Claude Opus 5 #
Its cheapest MacBook is its soundest: high, at $0.50, is the one Opus 5 cell no checklist run faulted twice. Spend more and the construction breaks - max, at its 128K ceiling, draws a wedge base ending in a spike.
Cut off by OUR 64K sweep budget: thinking is on by default, so it spent almost all of it reasoning and emitted only the background glow, $1.60 billed for nothing. Given its own 128K ceiling it finished (above).
More
The most detailed, ambitious Anthropic laptop here - notch, dock, traffic lights, a code window, convincing materials - and the first Anthropic model whose effort dial actually changes the drawing instead of redrawing the same thing. But it fails the one thing the MacBook test is for: the geometry. At high the base is skewed; at xhigh the isometry is broken outright - the screen and the keyboard deck sit in different perspectives, so it reads slick at a glance and falls apart on a second look. And max was the worst of the three until we found the cause: thinking is on by default, so it spent almost our entire 64K sweep budget reasoning and emitted only the background glow, $1.60 for nothing. Given its own 128K ceiling on 2026-09-02 the same request finished in 74,061 tokens and $1.85 - its best laptop by a distance, and the first Opus 5 cell whose lid and deck belong to one object: silver body, notch, menu bar, a dock of real app icons, grilles, trackpad. The base is still a wedge that thins to nothing at the front, and the view is nearer flat-on than three-quarter. A clear step up from Opus 4.8's flat mid laptop in detail and effort-response - but what separates Fable is that it makes no visual errors, where Opus 5 only looks the part.
- high - All the detail is here - notch, dock with app icons, a code window, "MacBook Pro" under the screen. Read closely the construction does not hold up: the base is wider than the lid on both sides, sits at a far more top-down angle, and ends in a sharp point at the front-right. $0.50 and 20K output tokens, because thinking is on by default.
- xhigh - More detail and better materials than high, but the isometry is broken: the screen and the keyboard deck sit in different perspectives, so it looks slick at a glance and falls apart on a second look. 47K output tokens and $1.18 for it.
- max - At the model ceiling (128K): finished in 74,061 tokens and 14.9 minutes for $1.85. The richest Opus 5 drawing in content - silver body, notch, menu bar, a dock of real app icons, grilles, trackpad - and the most broken in construction. Every checklist run named the same five faults: sharp corners, a spike at a corner, a base that is a wedge rather than a slab, wrong proportions, and keys that never resolve into rows. Its high cell, at $0.50, is the one Opus 5 drawing no run faulted twice.
- max @64K - Cut off by OUR 64K sweep budget: thinking is on by default, so it spent almost all of it reasoning and emitted only the background glow, $1.60 billed for nothing. Given its own 128K ceiling it finished (above).
In Playcode since 2026-07-24 · runs 2026-07-24, 2026-09-02
Kimi K3 #
A well-finished dark MacBook - grilles, dock, code window - drawn nearly flat-on with a tapered lid over a dead-on base; 17.5 minutes at max.
More
The new challenger, and a serious one on finish: a proper dark MacBook with a dock, a notch and speaker grilles at both levels, for a fifth of Fable's price. On closer inspection it is not error-free, though. It draws the laptop nearly flat-on rather than in the three-quarter view asked for, the lid tapers while the base stays dead-on so the two perspectives disagree, and the base is a depthless slab whose ends splay past the lid. Tidy, cheap, and wrong in the way this test is built to expose - and the clock is brutal at max (17.5 minutes). As ever, the pelican it already aces at high gains nothing from max.
- high - A proper dark MacBook - dock, notch, keyboard, trackpad - already here for $0.15. That was a hair over Gemini 3.6 Flash at max until Google halved the Flash rate card; it is now roughly three times the dearest Gemini cell.
- max - The best-finished cheap laptop here - speaker grilles, a code window, a dock - but not error-free: it is drawn nearly flat-on instead of the three-quarter view asked for, and the tapered lid disagrees with the dead-on base. 30,579 tokens and 17.5 minutes for it.
In Playcode since 2026-07-23 · runs 2026-07-24
GLM 5.3 #
The best-assembled cheap laptop yet - a real three-quarter MacBook with a clean hinge - but its default cell takes 73K tokens, 16 minutes and $0.32 to draw it.
More
The first cheap model to draw a three-quarter MacBook that actually holds together: twilight wallpaper with a moon, a dock, a sculpted key grid, speaker grilles, and a hinge whose halves meet - the assembly GLM 5.2 and every Gemini never managed. The catch is the meter. That drawing took 73K output tokens, 16.5 minutes and $0.32 - forty-five times 5.2's price for the same task - because 5.3's server-side default effort sits between high and max, and reasoning bills as output. The dial, at least, is real: the first ORDERED effort ladder in this catalog, low 3.9K tokens to max 93K, the second-longest run on this page at 20 minutes. Its pelicans are the cheap tier's best, and max even gives the bird a saddle - the accessory nearly every other model forgets. Run it at low and it is a one-cent model that draws like a toy; leave the default and it spends like a mid-tier reasoner. Same $1.40/$4.40 card as 5.2.
- default - The best-assembled cheap laptop yet: a real three-quarter MacBook with a twilight wallpaper, dock, sculpted key grid and a hinge whose halves actually meet. It took 73K output tokens and 16.5 minutes - the server default effort sits between high and max.
- low - Crude at a cent and a half: the screen slab sits skewed on a flattened base.
- high - Front-on with a dated menu bar and full keyboard - and a ghost second lid edge behind the screen.
- max - 93K output tokens over 20 minutes - the second-longest run on this page. Tidier than high, same ghost slab.
In Playcode since 2026-09-04 · runs 2026-09-04
Muse Spark 1.3 #
Meta's cheap tier, via OpenRouter: a MacBook with a real menu bar, a browser window and a joined hinge for $0.035 - behind up to a minute of silence before the first token.
More
The surprise of the batch. Meta's Muse Spark 1.3, reachable only through OpenRouter, draws the second-best cheap laptop on the page: a menu bar with real Finder text, a notch, a browser window with traffic lights, a dock, a sculpted keyboard, a joined hinge - for three and a half cents. Its tell is a cyan glow spilling past the lid's edge and a slightly splayed deck. Its pain is latency, not price: up to 70 of a cell's seconds pass before the first token arrives through the aggregator, so a ninety-second drawing is mostly waiting. The pelicans are plainer flat cartoons - a helmet-and-cape rider at medium, no saddle anywhere - and the dial barely moves it: default, medium and high all land within 8.2-8.7K tokens on the laptop. Not selectable in Playcode; benchmarked because a frontier lab's cheap tier deserves a row.
- default - The surprise of the batch: menu bar with real Finder text, a notch, a browser window with traffic lights, a dock, a joined hinge - for three and a half cents. Its tells: a cyan glow spilling past the lid, and 59 of the 97 seconds passed before the first token.
- high - The dial barely moves it: default, medium and high all land within 8.2-8.7K tokens.
In Playcode since 2026-09-04 · runs 2026-09-04
GLM 5.2 #
Correctly assembled - notch, grilles, key grid, trackpad - for $0.075, but flat-on instead of the three-quarter view asked for. No effort dial.
More
Superseded by GLM 5.3, which finally adds the three-quarter view - at forty-five times the default-effort price. The most complete cheap laptop here, though not the best one. Everything is assembled correctly - notch, speaker grilles on both sides, a real key grid, a trackpad, a lid whose bezel meets the hinge - and nothing is detached or floating, which is more than Opus 5 at twenty times the price manages. Google has since halved the Flash rate card, so GLM's old price peer is gone: Gemini 3.6 Flash at low now costs $0.036 against GLM's $0.075, and 3.8 Flash packs in more detail at $0.049 - but every Gemini laptop is crooked somewhere, and GLM's is not. GLM still renders individual keys and grilles where Gemini smears the deck into a dark band, but it does not do the one thing both Geminis do - a proper three-quarter view. GLM ignores that instruction and draws the laptop nearly flat-on. Its base is skewed against the lid, sliding right as it comes forward, so the two halves do not share a vanishing point. The pelican inverts the failure: the bicycle is excellent - spoked wheels, a correct diamond frame, a crank - but the bird is draped over the top tube rather than sitting on it, with arrow-headed legs that stop short of the pedals. Same cost, same 17K tokens, same three minutes for both drawings: it has one budget and spends all of it whatever you ask for.
- default - Correctly assembled and cleanly shaded - notch, speaker grilles, key grid, trackpad - but drawn almost flat-on instead of the three-quarter view asked for, and the base is skewed against the lid. Only one cell: Z.ai has no effort dial, so the ladder does not apply.
In Playcode since 2026-06-19 · runs 2026-08-14
Gemini 3.8 Flash #
Tries harder than 3.7 and is the first Gemini heavyweight with valid markup in all 8 cells - but the geometry is still off everywhere: skewed decks, overhanging keys, a lid that floats at max.
More
Better behaved than 3.7, and still not a well-built laptop. Every cell is misaligned somewhere: default's key block overhangs the deck's left edge and the two halves do not quite share a vanishing point, max's lid floats a hair off the hinge while 14.9K tokens - the heaviest Gemini Flash cell ever - go partly on ambient glow circles, and low smears the screen's colours across the keys. The parts are all present (screen glow, notch, menu bar, dock, a real key grid for $0.049) but the assembly never closes, which is the thing this page ranks. What 3.8 genuinely fixes is reliability: 8 cells, 8 valid SVGs, where 3.7 shipped a file no browser would finish - and the pelican keeps the tier's charm, cap, scarf, fish in the pouch, webbed foot on the crank. Google says 3.8 may think harder at higher efforts and the dial is still unordered: medium is its cheapest MacBook cell, max its dearest at nearly 2x. Same halved rate card as 3.7 through 2026-12-31. A crooked laptop for a fortieth of Fable's price is still a crooked laptop.
- default - Its best try: a three-quarter MacBook with screen glow, a notch, a menu bar, a dock and a real key grid - still misaligned, with the key block overhanging the deck edge and the two halves not quite sharing a vanishing point.
- low - The screen's colours smear across the keyboard deck - the weakest of its four cells, and dearer than default.
- medium - Its cheapest cell and arguably its cleanest: a tidy three-quarter laptop with an aligned key grid and speaker grilles.
- max - The heaviest Gemini Flash cell ever measured here - part of the budget goes on ambient glow circles behind the laptop, and the lid hovers a hair off its hinge line.
In Playcode since 2026-09-04 · runs 2026-09-04
Gemini 3.7 Flash #
Most detail in the cheap tier - notch camera, menu bar, dock, a real key grid - but the lid and base do not share a hinge, and 1 of its 8 cells is invalid SVG.
More
Superseded by 3.8 Flash, which Google prices identically and which draws a straighter - still not straight - laptop. The most detailed draughtsman in the cheap tier by a distance, and the least reliable. Its MacBook has a notch camera, a menu bar, a dock of coloured icons and a real key grid for $0.052; its pelican genuinely rides the bicycle - helmet, scarf, a fish in the pouch, webbed feet on the pedals - which is more than most models here manage. Then it drops the assembly: the laptop's lid and base do not share a hinge, so the deck reads as a slab sliding out beneath the screen, and at max on the pelican it emitted invalid SVG that no browser will finish drawing. Read the price twice. It is a fortieth of Fable's best drawing, but ten to sixteen times Gemini 3 Flash for the same prompt, because it thinks eight times as hard and thinking bills as output. Effort buys nothing: its four cells scatter between 9.3K and 13.9K tokens in no order.
- max - The most detailed laptop in the benchmark - notch camera, menu bar, dock, a real key grid - but the lid and the base do not share a hinge, so the deck reads as a separate slab sliding out beneath it.
In Playcode since 2026-08-26 · runs 2026-08-26
DeepSeek V4.1 Flash #
Three real MacBooks in about a minute for two cents each - and each with one thing wrong: a stray slab, a sheared keyboard, a lid and base drawn from different heights.
More
The first DeepSeek that draws the machine. Where V4 Flash shipped orphan lines and gave up on perspective, every V4.1 Flash cell is a MacBook with a notch, screen content, speaker grilles and a trackpad, finished in 52 to 58 seconds - three times faster than its predecessor's laptops - for about two cents. None of the three holds together, and each fails differently: at low a huge grey slab hangs off the base like a second, empty laptop; at high the keyboard shears sideways off the deck and a black spike sticks out of the right edge; at max - the cleanest surface, with a menu bar, a dock and a macOS-style wallpaper - the lid is drawn from below and the base from above, with a translucent ghost of the deck floating off to the right. Assembly, not detail, is what keeps it out of the top ten. The dial barely moves the meter: 15.7K, 16.3K and 17.5K output tokens. Its pelicans are better than its laptops again, and low is the keeper. Released 2026-09-10 with native image input on a card cheaper than V4 Flash's ($0.30 in / $1.20 out).
- low - Two cents and 52 seconds for a recognisable MacBook - notch, a tiny wordmark on the chin, grilles both sides, a trackpad - drawn flat-on above a deck in steep perspective, with a huge grey slab hanging off the base like a second, empty laptop. The keyboard is horizontal bands, not keys.
- high - The most three-dimensional of the three, with individual keys, a dock and a screen glow on a dark ground - but the key block shears sideways and spills off the deck, the trackpad sits in a different projection from the keys, and a black spike sticks out of the right edge. Five seconds slower than low, for the same two cents.
- max - The cleanest finish DeepSeek has produced here: a menu bar with traffic lights, a dock, a macOS-style wallpaper, grilles, trackpad, a proper notch. Then the geometry: the lid is drawn from below and widens towards the top while the base is seen from above, the keyboard is a blank dark rectangle, and a translucent ghost of the deck floats off the right side. Its best laptop, and still a rank-12 laptop.
In Playcode since 2026-09-10 · runs 2026-09-10
A real MacBook for under two cents - notch, perforated grilles, a full key grid, a trackpad - let down by a keyboard drawn flatter than the deck it sits on, and a stray wedge off the front-left corner. Pay more and it gets worse before it gets better: the high cell has no keyboard at all.
More
The cheapest model in this catalog by a wide margin, and the first at this price to draw a machine you would recognise without being told. Its max cell is $0.019 - under two cents for 38,003 tokens and nine minutes - and it puts a notch and camera on the bezel, perforated speaker grilles on both sides, an individually-keyed keyboard and a properly proportioned trackpad on an aluminium body. What it does not do is furnish the screen: no menu bar, no dock, just a wallpaper, where models a few places above manage both. The assembly fault is consistent - the key block is drawn flatter than the deck it sits on, so the keyboard reads as pasted onto the laptop rather than set into it, with a thin wedge protruding past the front-left corner. Its dial is the sharpest illustration on this page of why token counts are not quality: output rises cleanly from 6.3K to 38K, but the DRAWINGS do not follow. Default, the cheapest cell at a third of a cent, has a full key grid. High, which costs nearly twice as much, has no keyboard at all - just a blank silver plate where the keys should be. Xhigh brings them back and adds grilles and a side port, but hangs the lid at a different eye height from the base. Only max gets everything into one object. It ranks above Grok 4.6 because its lid and base do at least join, which Grok's do not, and below DeepSeek V4.1 Flash, which furnishes its screens. Half the price of GPT-5.6 Luna, which it replaces on the Playcode Fast tier, and six places above it.
- default - A third of a cent. A recognisable laptop with a notch, a full key grid and a trackpad - but drawn nearly flat-on where the prompt asks for three-quarters, with no grilles and a wide empty stretch of deck to the right of the keyboard.
- high - Nearly twice the price of default and it has NO KEYBOARD - a blank silver plate where the keys should be, with only a faint rectangle marking the area and a single grille patch on the right. The clearest case on this page of more tokens buying a worse drawing.
- xhigh - Keys return, and with them dotted grilles on both sides and a port on the right edge. The lid is drawn flatter than the base, so the two halves sit at slightly different eye heights and the hinge does not quite close.
- max - Its best, and still under two cents: notch, grilles both sides, a full key grid, a trackpad and a bright wallpaper on an aluminium body. The key block is drawn flatter than the deck it sits in, and a thin white wedge juts past the front-left corner.
In Playcode since 2026-09-22 · runs 2026-09-23
Grok 4.6 #
Finally tries - real laptops with screen glow and a dock where 4.5 shipped 668 tokens of nothing - but the deck floats free of the screen and the key lattice slides off the edge.
More
The redemption of the xAI seat, partially. Grok 4.5 was this benchmark's floor - 668 output tokens for a laptop it barely attempted - and 4.6 actually engages: ten times the tokens, a screen with glow and a dock, a dense keyboard, a trackpad. What it cannot yet do is join the parts. At default the deck floats detached below the screen with a stray line poking from the left edge and the key lattice sliding off the deck's right side; at high the deck's corner tears away like a paper fold while the keyboard fades to near-white. The pelicans are tidier than the laptops: high is a calm, coherent flat cartoon done in 38 seconds for a cent - the fastest finished drawing on this page - and default puts two feet on the pedals of a solid dark-blue bike, though neither bird gets a saddle. Same $2/$6 card as 4.5, reasoning still always on, and the effort dial moves output in no consistent direction: high spent more tokens than default on the laptop and fewer on the bird.
- default - A real laptop at last from xAI - screen glow, dock, dense keyboard, trackpad - but the deck floats detached below the screen, the key lattice slides off its right edge, and a stray line pokes from the left.
- high - Calmer and paler: browser chrome and a camera notch on the screen, but the keyboard fades to near-white and the deck's right corner tears away like a paper fold.
In Playcode since 2026-09-04 · runs 2026-09-04
Gemini 3.6 Flash #
Plainer than 3.7 and correct wherever it commits; it smears the keyboard into dark bands, and one MacBook cell is invalid SVG.
Emitted invalid SVG - the error was: Specification mandates value for attribute points. A browser draws only as far as the error, so there is no finished picture.
More
Superseded by 3.7 Flash, which Google prices identically. Two things changed since the launch sweep. Google halved the Flash rate card, so the run this article once billed at $0.16 is $0.041; and a fuller four-effort sweep killed the claim that this was the first Gemini with a working effort dial - its cells sit between 9.4K and 10.9K tokens in no order at all. What survives is the drawing: flat-shaded, plainer than 3.7, and correct everywhere it commits, though it smears the MacBook keyboard into dark bands. It broke once, emitting invalid SVG on the MacBook at default. It still takes the pelican off 3.7, which broke on that one.
- default - Emitted invalid SVG - the error was: Specification mandates value for attribute points. A browser draws only as far as the error, so there is no finished picture.
In Playcode since 2026-07-21 · runs 2026-08-26
DeepSeek V4 Flash #
About a cent per drawing and among the worst laptops here - two perspectives in one, stray lines; max is clean but flat-on.
More
Superseded by V4.1 Flash, which DeepSeek prices lower and which finally draws the machine - see rank 12. The single clearest piece of evidence that the pelican test is contaminated. Its pelican at max is genuinely good - a properly built bicycle with a chain and cranks, and a pelican actually sitting on it - and its MacBooks are among the worst drawings in this entire benchmark. Not merely weak: broken. Stray lines and orphan shapes drift across the canvas, the lid and the keyboard deck are drawn in different perspectives, key rows shear off to one side. Even its best laptop gives up on the three-quarter view the prompt asks for and retreats to a flat front-on cartoon with a keyboard too narrow for its own deck. That gap is not a fluke of effort - a model this cheap and this fluent at the bird cannot draw the object nobody has trained it on, which is the whole reason we stopped grading on pelicans. Costs about a cent per drawing, and on the MacBook that is still not a bargain.
- low - Three quarters of a cent, and the least broken of its three - which is the only sense in which it is good. A generic silver wedge: the lid is a paper-thin sliver with no depth at all, the screen dwarfs the body, and nothing marks it as a MacBook rather than any laptop in a stock-vector pack.
- high - One of the worst drawings in this benchmark, at any price. The lid and the keyboard deck are in two different perspectives, the key rows shear away to the right and spill past the deck, the trackpad is a skewed parallelogram, and loose lines and orphan shapes hang off the front edge and the hinge. It is also the cheapest of its three runs - raising the dial one notch from low makes the picture visibly worse.
- max - Double the tokens and double the wait to stop drawing stray lines. It finally names the machine - notch, "MacBook Pro" wordmark, symmetric body - but only by abandoning the brief: the prompt asks for a 3D model and this is a flat, front-on cartoon with no perspective at all, its keyboard a narrow trapezoid floating on a deck it does not fit. The best MacBook it can draw, and still near the bottom of the field.
In Playcode since 2026-05-27 · runs 2026-08-02
Claude Opus 4.8 #
Three tidy, identical, mid drawings for $0.05, $0.10 and $0.18 - the effort dial buys nothing.
More
Cheap, fast, consistent - and consistently mid. Its three effort levels cost $0.05, $0.10 and $0.18 and produce what is recognisably the same drawing. The effort dial buys nothing here.
In Playcode since 2026-05-31 · runs 2026-07-13
Gemini 3 Flash #
A simplified laptop with nothing broken for $0.005, valid markup in all 8 cells; the dial does nothing to it.
More
The cheap-and-fast control, and still the bargain of the benchmark: a simplified laptop with flat shading where nothing is broken, for $0.005. It knows what it can execute and stays inside that envelope - eight cells across both tasks and not one line of invalid markup, which neither Gemini heavyweight manages. The effort dial does nothing to it: max is its cheapest cell of the four, not its dearest.
- max - The dial is inert on this model: max is its CHEAPEST cell of the four, not its dearest.
In Playcode since 2026-01-20 · runs 2026-08-26
GPT-5.6 Luna #
OpenAI's nano tier after the 5x price cut: a recognisable laptop for half a cent, a cleanly ordered effort dial - and a high cell that ships blocks of static over the deck.
More
Superseded by GPT-6 Luna at half the price again ($0.20/$1.20 -> $0.10/$0.50), which draws a better laptop six places above - see rank 15. The cheapest OpenAI entry ever on this page until then. After the 2026-09-04 reprice ($1/$6 cut to $0.20/$1.20) it competes with Gemini 3 Flash on price - its default MacBook costs half a cent - and what that buys is uneven. xhigh draws its best laptop, a rounded silver body with a sculpted keyboard, coherent and assembled if a little toy-ish; high smears rectangles of noise texture across the deck; max mistakes the brief for a tablet propped on a serving tray. Like GLM 5.3 the same day, its dial is genuinely ordered - 4.2K tokens at default rising monotonically to 23.4K at max, on both tasks - which no OpenAI sibling here manages. The claim that it matches Gemini 3 Flash is close but not true: Flash never breaks, Luna sometimes does. It ranks one below.
- default - A recognisable laptop for half a cent on the freshly cut nano pricing - the base is a thin angular wedge and the keyboard floats on it.
- high - Ships rectangles of noise-texture static across the deck. Valid SVG, broken picture.
- xhigh - Its best cell: a rounded silver body with a sculpted keyboard - coherent, assembled, a little toy-ish.
- max - Mistakes the brief for a tablet propped on a serving tray.
In Playcode since 2026-09-04 · runs 2026-09-04
Claude Opus 4.6 #
Recognisable but loose, with parts in the wrong places; rejects xhigh outright.
More
Rejects xhigh outright - the bug this benchmark found in its first ninety seconds. What it does draw is recognisable but loose, with parts in the wrong places under inspection.
In Playcode since 2026-06-29 · runs 2026-07-13
Claude Sonnet 4.6 #
A laptop-shaped object with misplaced parts; rejects xhigh; $0.10-0.11 whatever you ask for.
More
Also rejects xhigh. Draws a laptop-shaped object with misplaced parts, and effort barely moves it: $0.11, $0.11, $0.10 across the ladder.
In Playcode since 2026-03-02 · runs 2026-07-13
GPT-5.6 Terra #
A wide, thin generic ultrabook at every budget - even with 29,834 tokens at max.
More
Fast and reliable, and it does not know what a MacBook looks like. Even uncapped at max, with 29,834 tokens to spend, it draws a wide, thin generic ultrabook. The clearest evidence in this benchmark that budget does not buy understanding.
- max - Uncapped it finishes - and the form is still wrong. Budget does not buy understanding.
In Playcode since 2026-07-09 · runs 2026-07-13, 2026-07-21
Claude Sonnet 5 #
Broken geometry, and the longest run on the page: 121,713 tokens to close one SVG, drawn tipped over.
Cut off by OUR ~64K sweep budget (the provider reported 65,538 output tokens against it): it never closed the SVG. At its own 128K ceiling it needed 121,713 tokens - nearly double our budget.
More
Broken geometry on the laptop, and the most expensive way to get there in tokens: at max it needs 121,713 output tokens and 21 minutes to close one SVG - nearly double our old 64K sweep budget, which is why this cell read as a runaway until 2026-09-02. What it finally draws is a complete laptop, dock and key grid included, tipped over: the lid and the base are two skewed slabs sharing no hinge and no vanishing point. Perfectly fine on a pelican, lost on an object you know well.
- max - At the model ceiling (128K): it closes the SVG at last - 121,713 tokens, 21 minutes, $1.22, the longest run on this page. A complete laptop with a dock, a window and a full key grid, drawn tipped over: the lid and the base are two skewed slabs that share no hinge and no vanishing point.
- max @64K - Cut off by OUR ~64K sweep budget (the provider reported 65,538 output tokens against it): it never closed the SVG. At its own 128K ceiling it needed 121,713 tokens - nearly double our budget.
In Playcode since 2026-06-30 · runs 2026-07-13, 2026-09-02, 2026-07-21
Grok 4.5 #
668 tokens, more logo than laptop - the cheapest MacBook run here at $0.004; high is its ceiling.
More
Superseded by Grok 4.6, which xAI prices identically and which actually tries. Barely tries. 668 output tokens on the MacBook - more logo than laptop - though that makes it the cheapest run in the benchmark at $0.004. Its pelican is fine. high is its ceiling.
- high - high is its ceiling. 668 output tokens - it barely tried.
In Playcode since 2026-07-10 · runs 2026-07-13
What one drawing costs #
Every MacBook run, cheapest first, including the ones that billed and returned nothing (greyed). * = requested xhigh, clamped to high by the model; @16K and @64K = a run cut off by our own output budget, kept here because it was still billed.
| Model · effort | Cost | vs cheapest |
|---|---|---|
| GPT-6 Luna · default | $0.003 | 1.0x |
| Gemini 3 Flash · max | $0.004 | 1.3x |
| Grok 4.5 · high | $0.004 | 1.3x |
| Gemini 3 Flash · medium | $0.005 | 1.5x |
| GPT-5.6 Luna · default | $0.005 | 1.6x |
| Gemini 3 Flash · default | $0.005 | 1.6x |
| Gemini 3 Flash · low | $0.005 | 1.6x |
| GPT-6 Luna · high | $0.006 | 1.8x |
| DeepSeek V4 Flash · high | $0.006 | 2.0x |
| DeepSeek V4 Flash · low | $0.008 | 2.4x |
| GPT-6 Luna · xhigh | $0.008 | 2.5x |
| GPT-5.6 Luna · high | $0.009 | 2.9x |
| DeepSeek V4 Flash · max | $0.013 | 3.9x |
| GLM 5.3 · low | $0.017 | 5.4x |
| GPT-5.6 Luna · xhigh | $0.019 | 5.8x |
| DeepSeek V4.1 Flash · low | $0.019 | 5.9x |
| GPT-6 Luna · max | $0.019 | 5.9x |
| DeepSeek V4.1 Flash · high | $0.020 | 6.1x |
| DeepSeek V4.1 Flash · max | $0.021 | 6.6x |
| Muse Spark 1.3 · low | $0.021 | 6.6x |
| Claude Sonnet 5 · high | $0.026 | 8.3x |
| GPT-5.6 Luna · max | $0.028 | 8.8x |
| Grok 4.6 · default | $0.029 | 9.1x |
| Muse Spark 1.3 · default | $0.035 | 11x |
| Gemini 3.6 Flash · medium | $0.035 | 11x |
| Gemini 3.6 Flash · low | $0.036 | 11x |
| Gemini 3.7 Flash · low | $0.036 | 11x |
| GPT-6 Sol · default | $0.036 | 11x |
| Muse Spark 1.3 · high | $0.036 | 11x |
| Muse Spark 1.3 · medium | $0.037 | 12x |
| Grok 4.6 · high | $0.040 | 12x |
| Gemini 3.6 Flash · max | $0.041 | 13x |
| Gemini 3.6 Flash · default no drawing | $0.041 | 13x |
| Gemini 3.8 Flash · medium | $0.043 | 13x |
| Gemini 3.7 Flash · medium | $0.045 | 14x |
| Gemini 3.7 Flash · default | $0.047 | 15x |
| Gemini 3.8 Flash · default | $0.049 | 15x |
| Gemini 3.7 Flash · max | $0.052 | 16x |
| Claude Opus 4.8 · high | $0.052 | 16x |
| Gemini 3.8 Flash · low | $0.056 | 17x |
| Gemini 3.8 Flash · max | $0.075 | 23x |
| GLM 5.2 · default | $0.075 | 23x |
| GPT-5.6 Terra · high | $0.085 | 27x |
| GPT-6 Sol · high | $0.086 | 27x |
| Claude Opus 4.8 · xhigh | $0.098 | 31x |
| Claude Sonnet 4.6 · max | $0.10 | 32x |
| Claude Sonnet 4.6 · high | $0.11 | 34x |
| Claude Sonnet 4.6 · xhigh* | $0.11 | 34x |
| Claude Sonnet 5 · xhigh | $0.12 | 38x |
| Kimi K3 · high | $0.15 | 46x |
| GPT-6 Sol · xhigh | $0.16 | 51x |
| GPT-5.6 Terra · xhigh | $0.17 | 55x |
| Claude Opus 4.8 · max | $0.18 | 56x |
| Claude Opus 4.6 · max | $0.19 | 61x |
| Claude Opus 4.6 · high | $0.20 | 64x |
| Claude Opus 4.6 · xhigh* | $0.21 | 66x |
| GPT-5.6 Sol · high | $0.26 | 80x |
| GLM 5.3 · high | $0.27 | 84x |
| Claude Fable 5 · high | $0.27 | 84x |
| GLM 5.3 · default | $0.32 | 101x |
| GPT-6 Astra · default | $0.33 | 104x |
| GPT-6 Sol · max | $0.39 | 121x |
| Claude Opus 5.5 · high | $0.40 | 124x |
| GLM 5.3 · max | $0.41 | 128x |
| GPT-5.6 Terra · max | $0.45 | 140x |
| GPT-5.6 Sol · xhigh | $0.45 | 141x |
| Kimi K3 · max | $0.46 | 143x |
| Claude Opus 5 · high | $0.50 | 158x |
| Claude Sonnet 5 · max @64K no drawing | $0.66 | 205x |
| Claude Fable 5 · xhigh | $0.73 | 228x |
| Claude Fable 5.1 · high @16K no drawing | $0.80 | 250x |
| Claude Fable 5.1 · xhigh @16K no drawing | $0.80 | 250x |
| GPT-6 Astra · high | $0.88 | 274x |
| Claude Fable 5.1 · high | $0.90 | 281x |
| GPT-6 Astra · xhigh | $0.98 | 307x |
| Claude Opus 5 · xhigh | $1.18 | 369x |
| Claude Sonnet 5 · max | $1.22 | 380x |
| Claude Opus 5 · max @64K no drawing | $1.60 | 500x |
| Claude Opus 5.5 · xhigh | $1.72 | 538x |
| Claude Opus 5 · max | $1.85 | 579x |
| Claude Fable 5 · max | $2.13 | 665x |
| GPT-6 Astra · max | $2.29 | 716x |
| Claude Opus 5.5 · max | $2.56 | 800x |
| Claude Opus 5.5 · max no drawing | $2.56 | 800x |
| Claude Fable 5.1 · xhigh @64K no drawing | $3.20 | 1000x |
| Claude Fable 5.1 · max @64K no drawing | $3.20 | 1000x |
| Claude Fable 5.1 · xhigh | $4.34 | 1355x |
| Claude Fable 5.1 · max | $4.95 | 1548x |
| GPT-5.6 Sol · max no drawing | not billed | - |
The best drawing costs 65x the best-value one. Whether that detail is worth it depends on what you are building - but you cannot ask the question from a $/Mtok price list. Why the sticker price misleads across vendors: The Same TypeScript Costs 73% More on Claude Than on GPT.
How much room each model has #
One answer can only be so long, and the limit is the provider's, not ours. Measured 2026-09-02 by asking every API in this benchmark for 900,000 output tokens in one answer and reading the rejection. Against it: the most tokens that model actually spent on one MacBook here. Sorted by appetite.
| Model | Output ceiling | Biggest drawing it wrote |
|---|---|---|
| Claude Opus 5.5 | 300,000Batch API with the output-300k-2026-03-24 beta - the ONLY surface above 128,000; the streaming Messages API caps this model at 128,000 and rejects 128,001 by name | 255,957 tok · max |
| Claude Sonnet 5 | 128,000Anthropic Models API | 121,713 tok · max |
| Claude Fable 5.1 | 128,000Anthropic Models API | 99,072 tok · max |
| GLM 5.3 | 131,072Z.ai rejection (probed 2026-09-04: legal range [1, 131072]) | 93,272 tok · max |
| Claude Opus 5 | 128,000Anthropic Models API | 74,061 tok · max |
| GPT-6 Astra | no limit enforcedaccepted 900,000 without error (model page documents 128,000; unlike Sol/Terra, no rejection) | 45,836 tok · max |
| DeepSeek V4 Flash | 393,216API rejection | 44,895 tok · max |
| Claude Fable 5 | 128,000Anthropic Models API | 42,560 tok · max |
| GPT-6 Sol | no limit enforcedaccepted 900,000 without error (model page documents 128,000; GPT-5.6 Sol rejected the same request) | 38,546 tok · max |
| GPT-6 Luna | no limit enforcedaccepted 900,000 without error (model page documents 128,000) | 38,003 tok · max |
| Kimi K3 | no limit enforcedaccepted 900,000 without error | 30,579 tok · max |
| GPT-5.6 Terra | 128,000API rejection | 29,834 tok · max |
| GPT-5.6 Luna | 128,000model page; the same cap the Sol/Terra probes measured | 23,418 tok · max |
| DeepSeek V4.1 Flash | 393,216API rejection (probed 2026-09-10: legal range [1, 393216], unchanged from V4 Flash) | 17,461 tok · max |
| GLM 5.2 | 32,000vendor docs (our account had no balance to probe) | 17,007 tok · default |
| GPT-5.6 Sol | 128,000API rejection | 15,042 tok · xhigh |
| Gemini 3.8 Flash | 65,536Vertex model page (same 65,536 cap the 3.7 probe measured) | 14,877 tok · max |
| Gemini 3.7 Flash | 65,536Vertex rejection | 13,863 tok · max |
| Gemini 3.6 Flash | 65,536Vertex rejection | 10,824 tok · max |
| Muse Spark 1.3 | no limit enforcedaccepted 900,000 without error (OpenRouter card lists 943,718) | 8,694 tok · medium |
| Claude Opus 4.6 | 128,000Anthropic Models API | 8,328 tok · xhigh* |
| Claude Sonnet 4.6 | 128,000Anthropic Models API | 7,340 tok · xhigh* |
| Claude Opus 4.8 | 128,000Anthropic Models API | 7,189 tok · max |
| Grok 4.6 | no limit enforcedaccepted 900,000 without error (probed 2026-09-04) | 6,477 tok · high |
| Gemini 3 Flash | 65,536Vertex rejection | 1,673 tok · default |
| Grok 4.5 | no limit enforcedaccepted 900,000 without error | 668 tok · high |
Two things follow. The ceiling is low relative to the task. Claude and GPT-5.6 stop at 128,000 output tokens whatever you pay; Gemini Flash at 65,536; only DeepSeek (V4 Flash, and V4.1 Flash after it) goes past a quarter of a million, at 393,216. Claude Fable 5.1 spent 99,072 of its 128,000 on one laptop - it has less than a third of its budget left, so this task is close to the largest single answer that model can give. Newer models eat more. Fable 5.1 writes 2.3x the tokens of Fable 5 for the same prompt (99,072 against 42,560), and Sonnet 5 writes 121,713 where Gemini 3.7 Flash writes 13,863 - thinking tokens are billed as output and they dominate. A per-million price is not comparable across models that spend this differently.
What the effort dial buys #
- Up to
xhigh, effort buys real quality on the strongest models. Fable 5's and Sol's best drawings inside the 16K cap are both atxhigh. On mid models it buys nothing: Opus 4.8's three drawings cost $0.05, $0.10 and $0.18 and look the same. - Every truncation on this page turned out to be our budget, not the model. The 16K cap of the first sweep hid four "cliff" failures; the 64K budget that replaced it hid three more. Given its own 128K ceiling, each of those three finished - Fable 5.1 at
maxin 99,072 tokens, Opus 5 in 74,061, Sonnet 5 in 121,713. A thinking model does not warn you that it is running out of room: it bills you and stops. - Finishing is not getting it right. Terra completes a full drawing at
maxwith 29,834 tokens and the form is still a generic ultrabook. Budget does not buy understanding of the object. - On some models the dial is not even a dial. Gemini Flash cells scatter inside a band in no order - Gemini 3 Flash's
maxis its cheapest cell. DeepSeek V4 Flash draws its worst MacBook athighand a better one atlow. - The top of the dial is a different product. On Fable 5.1,
maxcosts 5.5xhigh($4.95 against $0.90) and takes 19 minutes instead of 4 - for one laptop. Sol atmaxstill does not return at all: 36 minutes, twice. Measure the top setting on your own workload before you let users pay for it.
The checklist - and why there is no score #
"Looks better" does not survive from one sweep to the next, so we wrote a checklist: 42 yes-or-no tests in seven blocks, each one a thing that has actually gone wrong in this benchmark. It is what we now look for, in this order:
- Projection and geometry (10). One projection for lid and base; a shared hinge line along the full back edge; no skew between the halves; upright; lid width equal to base width; three-quarter view; perspective rather than an isometric diagram; lid open 95-120 degrees.
- Body, edges, corners (8). Rounded corners of equal radius; no triangular spikes from bad polygon joins; a base that is a uniform slab, not a wedge; a lid with thickness; a plausible deck depth; correct proportions; and a count of stray or orphan shapes.
- Display (7). A uniform bezel with a taller chin; a centred notch with one camera; macOS-shaped content clipped to the screen; screen glow; no gibberish text; no Apple logo on the front bezel (a modern MacBook has none).
- Deck (7). Six key rows with a function row and a space bar, inside their well; a trackpad centred on the keyboard, 35-60% of its width; grilles on both sides; no foreign-brand cues.
- Ports (4), materials and light (4), canvas (2). Ports on a side edge, 3-4 on the left, 3 on the right, inside the thickness band; metallic gradients, one light direction, a ground shadow; no z-order errors; the laptop centred and unclipped.
Then we tried to automate it, and it did not work. A judge model (Claude Sonnet 5) answered all 42 tests per drawing with evidence, three passes each, scored 100 - 8 x major - 4 x minor - 1 x cosmetic in code. We ran it three times, twice after fixing a real flaw in it - an answer schema the judge kept inverting, then a 100px measuring grid drawn over each image so positions were read rather than guessed. The scores moved by a mean of 26 points between runs and by 56 on one cell, and the ranking reshuffled every time. One Opus 5 drawing whose base is visibly wider than its lid on both sides scored 96, then 84. Of 86 defect claims made across the three runs, 17 were repeated by all three.
So there is no score on this page. What we publish instead is the part that reproduced: the defects all three runs named appear under the drawings they belong to, marked "checked". Two cells drew no repeated complaint from any run - and both turned out to have one when we went back and looked by hand at full size. Fable 5.1 at max has a notch where three faces meet at the front-left corner of the base; Opus 5 at high, which one run scored 96 out of 100, has a base wider than its lid on both sides ending in a knife point. Those two are marked "by hand". Silence from the judge is not evidence of a clean drawing.
The ranking above stays what it has always been: our reading of the pictures, printed as an opinion rather than a number - and it turns on a choice. Grade the brief, a perspective render with screen glow on an aluminum body, and Fable 5.1 at max leads. Grade the craft - corners, joins, a slab that is a slab - and Fable 5 at max leads, because there is nothing to pick at. We grade the brief. The checklist still earned its place: it is what made us look at base thickness and corner joins at all, and it is what the notes under each drawing are now written against.
Method - prompts, caps, dates, model ids
MacBook prompt (sent verbatim, no system prompt):
Generate an SVG that looks like a realistic 3D render of a MacBook Pro 16" (opened, slight three-quarter view, screen glow, aluminum body). Use gradients and perspective to fake the 3D. Respond with ONLY the SVG markup, no explanation.Pelican prompt:
Generate an SVG of a pelican riding a bicycle. Respond with ONLY the SVG markup, no explanation.- One API call per cell, no retries, no examples, first answer kept. Failures are shown as failures, never re-rolled. Drawings are the model's unedited SVG rendered in its own browser page (the models reuse the same
#ids, so inlining them into one document mixes their gradients and clip-paths). - Effort set through each provider's native parameter: Anthropic
output_config.effort, OpenAIreasoning.effort, xAIreasoning_effort, Googlethinking_level, Moonshot and DeepSeekreasoning_effort. Z.ai exposes none. Each model runs the ladder its provider supports; Grok 4.5 tops out athigh. - Output budget: 16,000 tokens in the first sweep (2026-07-13/14), 64K from 2026-07-21, and since 2026-09-02 each model's own maximum - 128,000 output tokens for the Anthropic and GPT-5.6 models, 65,536 for Gemini Flash, 393,216 for DeepSeek V4 Flash and V4.1 Flash alike (measured, per model). A fixed number is our cap, not a model limit, and on a model that always thinks it can be spent before the first tag; cells cut off by one are labelled and kept.
- Costs are the providers' own usage-reported token counts at list price; total measured spend on this page is $54.50. Every cell records the sweep it came from (under "More"), so mixed dates stay visible rather than blended.
- Model ids:
gpt-6-astra,gpt-6-sol,claude-opus-5-5,claude-fable-5-1,claude-opus-5,kimi-k3,gemini-3.6-flash,claude-fable-5,gpt-6-luna,grok-4.6,grok-4.5,gpt-5.6-luna,gpt-5.6-sol,gpt-5.6-terra,claude-sonnet-5,claude-opus-4-6,claude-opus-4-8,deepseek-v4-flash,deepseek-flash,glm-5.3,muse-spark-1.3,glm-5.2,claude-sonnet-4-6,gemini-3.8-flash,gemini-3.7-flash,gemini-3-flash-preview. - The measurements are ours; the prose was drafted with AI assistance and edited by a human.
Benchmark changelog #
- The empty Claude Opus 5.5 max cell is resolved, and it was never a runaway. That request was re-run on the Batch API with the output-300k-2026-03-24 beta - the only surface that allows this model more than 128,000 output tokens - and it finished naturally at 255,957, twice the streaming ceiling, drawing what is arguably the most complete laptop in this catalog: a full Apple menu bar with status icons and the Tue 9:41 clock, a dock of individually recognisable apps, a notch, side ports, perforated grilles, a legibly-labelled keyboard with function row, and screen glow bleeding onto the desk. It sits BESIDE the empty cell rather than replacing it, and is marked as a different method: batch is asynchronous, took about two and a half hours, and bills at half rate ($2.56; the same tokens on the standard card would be $5.12). Every other cell on this page is still one synchronous API call at the provider's own billing. The model's CEILING row now reads 300,000 with that caveat, because the streaming API rejects 128,001 by name.
- GPT-6 Sol and GPT-6 Luna added, both tasks, the day OpenAI halved the family's prices - and Sol enters at rank 2, the best drawing per dollar this page has measured. Its max cell is a product shot (full Finder menu bar, wifi and battery glyphs, the 9:41 clock, a dock of recognisable apps, a notch, grilles, and a keyboard with legible legends on every key) for $0.386, against the $2.29 GPT-6 Astra spent on the cell that still leads; Astra keeps rank 1 on side ports Sol does not draw. All four Sol cells assemble correctly, including the cheapest at 3.6 cents. Luna enters at rank 15 as the cheapest model in the catalog by a wide margin - a recognisable MacBook with notch, grilles, key grid and trackpad for $0.019 - and as the clearest demonstration here that token counts are not quality: its output rises cleanly from 6.3K to 38K, but its HIGH cell has no keyboard at all, a blank plate where the keys should be, while its cheaper default has a full key grid. Both replace their GPT-5.6 namesakes on the Playcode tiers at half the price; 5.6 Sol falls to 6 and 5.6 Luna to 21. Output ceilings probed the same day: both accepted a 900,000-token request without error, where GPT-5.6 Sol rejected it. Ranks 2-24 shift down.
- Claude Opus 5.5 added, both tasks, the day after Anthropic released it - and it enters at rank 2, the best drawing on this page under $2. Its xhigh cell is a product shot (real Finder menu bar with named menus and the 9:41 clock, a dock of recognisable apps, notch, side ports, grilles, the MacBook Pro wordmark) for $1.72, against the $4.95 Fable 5.1 drawing it matches and GPT-6 Astra's $2.29 rank-1 cell, which still wins on legible key legends and on landing every cell in its ladder. Its own ladder does not: at max it spent all 128,000 output tokens - the model's OWN ceiling, not our sweep budget - on thinking and returned an empty file for $2.56, so unlike Fable 5.1's capped cells there is no larger budget to re-run it with. Cheap is where it is most convincing: high is $0.40 for a correctly assembled laptop, 21% cheaper and 32% faster than the near-identical drawing Opus 5 produced at the same effort - an independent check on Anthropic's claimed 30% speed-up. Card: $4 / $20 with cache reads at $0.20, 20% under Opus 5 on tokens and 60% under on reads. Ranks 2-23 shift down one, and the stale NEW badges left on six earlier models were cleared.
- DeepSeek V4.1 Flash added, both tasks, on release day - the model that replaces V4 Flash on the Playcode Fast tier, with native image input and a cheaper card ($0.30 / $1.20). It enters at rank 12, three places above the V4 Flash it retires: every cell is now a real MacBook (notch, dock, grilles, trackpad) finished in under a minute for about two cents, but each carries one assembly error - a stray slab at low, a sheared keyboard at high, a lid and base drawn from different eye heights at max. The effort dial is nearly flat (15.7K / 16.3K / 17.5K output tokens). Output ceiling probed the same day: rejects above 393,216, unchanged from V4 Flash. Ranks 12-22 shift down one.
- GPT-6 Astra added, both tasks, the morning our API access opened - and it takes rank 1, the first time a Fable has not led this page. All four of its MacBook cells assemble without a visible error (no other ladder manages that), its max cell is the new best drawing at half the tokens, time and price of the dethroned Fable 5.1 cell, and its effort dial is cleanly ordered on both tasks. Tokenizer verified o200k (ratio 1.0000) by real probe the same hour; unlike its Sol/Terra siblings it silently accepts a 900,000-token output request. The best-drawing pin and the answer block moved accordingly.
- Three models added in one evening sweep: GLM 5.3 at rank 6 (the best-assembled cheap laptop yet and the first ORDERED effort dial in this catalog - but its server default spends 73K tokens and 16 minutes on one drawing), Meta's Muse Spark 1.3 via OpenRouter at rank 7 (a menu bar with real Finder text for $0.035, behind up-to-70-second first-token silences; the best-value pick moves here), and GPT-5.6 Luna at rank 16 (half-a-cent laptops on OpenAI's freshly 5x-cut nano pricing, a second ordered dial, and one cell of shipped static). Also recorded: OpenAI silently repriced the whole GPT-5.6 family - Sol $5/$30 to $4/$20, Terra $2.50/$15 to $2/$12, Luna $1/$6 to $0.20/$1.20; each run's cost reflects the card on its run date.
- Grok 4.6 added, both tasks (registered a month late - xAI shipped it ~2026-08-05 at the same $2/$6 card as 4.5). It ends 4.5's famous 668-token non-attempt: real laptops with a screen glow and a dock, ranked 9 because the parts still do not join - the deck floats free of the screen at default and tears at the corner at high. Its high pelican is the fastest finished drawing on the page: 38 seconds, one cent. Output ceiling probed the same day: accepted a 900,000-token request without error, like 4.5.
- Gemini 3.8 Flash added, both tasks, on launch day - same halved rate card as 3.7 (through 2026-12-31). It is the first Gemini heavyweight to return valid markup in all 8 of its cells, and it tries harder than 3.7 - but on editorial review the same day, both Gemini laptops were re-graded on assembly: everything is misaligned somewhere, so 3.8 enters at rank 7 and 3.7 falls to 8, below GPT-5.6 Sol, Opus 5, Kimi K3 and GLM 5.2. The best-value pick moves to GLM 5.2, the most correctly assembled cheap laptop. The Gemini effort dial stays unordered: medium was 3.8's cheapest MacBook cell, max its dearest at nearly 2x and the heaviest Gemini Flash cell measured here (14.9K output tokens), consistent with Google's own note that 3.8 "might use more tokens to maximize performance".
- Budget rule changed: every cell now runs at the model's own output ceiling (128,000 tokens for the Anthropic models here), not a number we picked. The four MacBook cells that had "run away" - Fable 5.1 at xhigh and at max, Opus 5 at max, Sonnet 5 at max - all finished when re-run that way, and Fable 5.1 at max is now the best drawing on the page. The capped runs stay in the table, labelled, because they were billed: $8.66 at 64K, $1.60 more at the old 16K cap. Every provider ceiling in the benchmark was measured the same day and published under "How much room each model has". A 42-item defect checklist was written the same day and run three times by a judge model over the three Anthropic ladders; the scores disagreed with each other by up to 56 points, so the page publishes only the defects every run repeated, and no score.
- Claude Fable 5.1 added, both tasks. On launch day only its high MacBook drew ($0.90): at xhigh and max it spent our whole 64K budget thinking and returned nothing, and the first sweep's 16K cap was exhausted by thinking alone. The next day's re-run showed the 64K number was the problem, not the model. The page was rebuilt as this catalog.
- Gemini 3.7 Flash added, both tasks, and every Gemini cell re-measured from scratch after Google halved the 3.6 and 3.7 Flash rate card (through 2026-12-31). No Gemini Flash model responds to the effort dial in any consistent order, and 2 of the 24 Gemini cells emitted invalid SVG.
- GLM 5.2 (Z.ai) added. It exposes no effort dial, so it gets one cell per task.
- DeepSeek V4 Flash added on its low / high / max ladder.
- Claude Opus 5 and Kimi K3 added.
- The max column re-run with a 64K output budget, streamed: four "cliff" failures in the first sweep were our 16K cap, not the model. The same sweep found that our harness had never wired Gemini's thinking_level, so every launch-sweep Gemini run had been at the default.
- Published from the first sweep (16K output cap). Its first ninety seconds found a production bug: Claude Opus 4.6 and Sonnet 4.6 reject xhigh with an API error, which our quality dial had been mapping onto them. Fixed the same night.
Try it yourself #
Paste the MacBook prompt above into any model, or swap in an object your audience knows intimately - a Coke can, a Vespa, a Stratocaster. The only requirements are that the ground truth lives in your reader's head and that style cannot hide broken geometry. We re-run this catalog when a frontier model ships.
Run it on a real project in PlaycodeEvery model here, one click apart, with the same effort dial.