← back to the archiveCover illustration for “Launch buzz measures demoability”
POSTday 10·7 weeks ago·by Andy Padia

Launch buzz measures demoability

400 builder posts from Claude Fable 5's first 48 hours tell you what the model makes easy to show off. They tell you almost nothing about what survives production review — the sample is selected for shareability.

Claude Fable 5 is days old and the launch-reaction industry is already at work: one newsletter says it classified 400 builder posts from the model's first 48 hours into a map of what people are building. I could not verify the sampling method or reproduce the categories — the piece is paywalled — but the exercise itself is the interesting artifact, because every major model launch now produces one, and every one of them gets read the same wrong way.

A pile of launch posts is not a capability survey. It is a demoability survey. The sample is selected — ruthlessly, structurally — for what performs well in public: visual output, one-shot wins, surprising behavior, anything that fits in a clip. What never appears in the sample: the boring extraction pipeline that got 2% better, the confidential legal workload nobody may post about, the task that failed quietly after six attempts, and the job that worked but cost too much to repeat. The absent categories are precisely the ones an enterprise runs.

So launch telemetry answers one question with real authority: what does this model make easy to demonstrate? That is genuinely useful — early posts are how interaction patterns and surprising capabilities get discovered, and I read them for exactly that. The failure is using them to answer a different question: should we build on this?

Where the two questions diverge

Fable 5 lists at $10 per million input tokens and $50 per million output. A launch-week demo pays that price once, for a clip. Your workload pays it at volume, every day, against an alternative that may be five times cheaper and good enough. No launch post carries that arithmetic, because the person posting ran the task once and was selecting for wow, not for unit economics. Demoability is measured in screenshots; production fitness is measured in acceptance rate at a price, and the two numbers are not even correlated in the cases that matter — the strongest demo categories are often the ones with the weakest tolerance for a 4% error rate.

At work I watched a product team run this movie earlier this year with a different model: the eval that chose their foundation model was, functionally, a highlight reel — a dozen viral examples reproduced in-house, all of which worked. Impressive demos, signed contract. The first month of production surfaced what the reel never could: their real documents were longer than anything in the demos, their acceptance criteria were stricter than a screenshot's, and recovery from failure — the thing no viral post ever shows — was the majority of the engineering. They re-ran the selection with their own fixtures and a cost curve; a less glamorous model won.

My rule: launch buzz is a discovery feed, never a decision input. Read the 400 posts to learn what is newly possible. Then test what you need on your own twenty fixtures, at your prompt lengths, with your acceptance bar and your failure-recovery path, and let that table pick the model. The demo economy is optimized to show you the best 48 hours of a model's life. You are buying the other 8,712.

Steal this for the next launch week: keep two lists. List one — "patterns worth stealing" — fill it from the viral posts freely. List two — "reasons to switch models" — may only be fed by results from your own eval fixtures. The discipline is refusing to let anything cross from list one to list two without passing through a test you own.

A launch tells you what the model shows well; only your fixtures tell you what it does daily — never let the first list make the second list's decision.

#model-selection#launches#evaluation#claude#product
← older drop
Agent go-live starts the expensive phase
newer drop →
Domain agents need disagreement signals

related drops

explore all 78 drops →
← back to the archiveday 59