← back to the archiveCover illustration for “An AI writing score still needs a reader on the other side”
POSTday 95·11d ago·by Andy Padia

An AI writing score still needs a reader on the other side

Tomasz Tunguz's writing experiment reports a higher AI-rated quality floor. The next test is whether readers understand and use the argument better.

Tomasz Tunguz reports that an AI panel scored the bottom end of his writing much higher after years of developing an AI-assisted workflow. In his September 2 account, the tenth-percentile score rose from 2.59 to 3.81 on a five-point scale. The panel evaluated 15 randomly selected published posts per year from 2021 through 2026. His experiment.

That is evidence about how the panel rated the sampled writing. It is not yet evidence that readers understood it better, remembered more or made better decisions. I would keep the quality claim provisional until one of those reader-side outcomes joins the score.

The distinction is especially interesting for a workflow that learns an author's preferred style. A system can become better at satisfying an editorial rubric while becoming more predictable in ways the audience does not value. That is a risk to test here, not a finding about Tunguz's work.

Editing volume is a different result

The operational detail I would keep is his reported median of 136 line-level edits per piece. His account says the workflow changed what he worked on more than how much editing he did. That is a useful description of effort being redirected toward refinement. Workflow account.

But a count of line changes is not a clock. One difficult edit can take longer than twenty routine substitutions, and a better argument may require fewer changed lines but more thinking. The count can describe a process without establishing equal labour cost across years.

The quality panel also lacks enough published detail for an independent reconstruction. The post does not provide its complete rubric, model identities, sampled articles and individual scores. That limits the conclusion; it does not make the experiment worthless. A useful personal measurement can be an invitation to a better shared test.

Ask readers to do something with the argument

For an illustrative AndyMental experiment, I would give readers two versions of a short technical argument without identifying which used AI assistance. I would hold the source evidence and intended claim constant. Changing topic, difficulty and assistance at once would make the result hard to interpret.

Then I would ask readers to name the decision the piece helps them make and the strongest limitation on its advice. Preference alone is useful, but it can reward an easy read that quietly loses an important qualification. A comprehension task checks whether the useful distinction survived the polish.

I would keep an AI rating alongside that result, rather than throwing it away. The interesting cases are disagreements: high-scoring prose that readers misinterpret, and lower-scoring prose that helps them make the right call. Those cases tell an editor what the rubric is missing.

This is my proposed evaluation, not a test already run on Tunguz's archive or this site. It would still need enough readers and examples to support a general claim. For a first pass, even a small set can reveal specific ambiguities worth repairing without pretending to prove a universal productivity gain.

The stronger writing workflow would optimise against those discovered failures. If readers repeatedly mistake a vendor claim for a measured result, the next revision should fix that distinction, even if the model panel already likes the cadence.

I like the attempt to measure craft. The next improvement is to let the audience disagree with the instrument. Otherwise the writing system and its judge can become very pleased with each other while the reader is still trying to find the point.

An AI score can guide revision; reader understanding decides whether the revision helped.

#ai-writing#evaluation#productivity
← older drop
EY's human-skills bonus rewards AI adoption
newer drop →
Harvey’s legal engineers make the delivery cost visible

related drops

explore all 329 drops →
← back to the archiveday 106