AI productivity tools, and the measurement problem underneath them

The hours-saved numbers in every AI roundup trace back to people estimating their own time. When someone measured it properly, the result went the other way.

7 minute read

METR's randomised controlled trial: 16 experienced developers, 246 real issues, AI access randomised per issue. Belief and measurement point in opposite directions.

MeasureDirectionSize
Expected effect, asked before the studyFaster24%
Believed effect, asked after the studyFaster20%
Measured effect, from task timingsSlower19%

Start with the number nobody quotes

Search for how much time AI saves and you will be told 6.1 hours a week, or 8.3, or 11. Those figures come from surveys where people estimated their own time, and they do not agree with each other because estimating your own time is not a reliable instrument. There is one study that measured it instead. METR ran a randomised controlled trial with 16 experienced open-source developers across 246 real issues, randomly assigning each issue to allow or disallow AI tools. Developers took 19% longer on the issues where AI was allowed. Before the study they expected AI to speed them up by 24%. After it — after being slower — they still believed it had sped them up by 20%. Note that there is no chart in this section on purpose: two of those numbers point one way and one points the other, and a bar chart would draw all three in the same direction.

Why this matters for creators, not just developers

The obvious objection is that editing video is not fixing GitHub issues. Fair. But the mechanism METR identified is not specific to code: the time goes into prompting, waiting, reading the output, and repairing it, and none of those feel like work in the way that doing the task yourself does. That gap — between effort felt and time spent — is exactly what makes self-reported hours-saved unreliable, and it applies to anyone who has ever regenerated a caption six times. Treat every unmeasured productivity claim in this category, including the ones in tool marketing, as a report about how the work felt.

The tools that survive the test, and the ones that do not

The distinction that actually predicts whether a tool saves you time is whether its output is verifiable at a glance. A tool that produces something you can check in a second — a transcript you skim, a cut you scrub, a background that is either removed or not — pays for itself, because the review step is cheap. A tool that produces something you have to think hard about — strategy, a content plan, an analysis of your numbers — moves the work rather than removing it, because now you have to audit prose that is confident and might be wrong. When a tool saves you time, it is usually because it removed a mechanical step, not because it did your thinking.

Where the time actually is for a one-person operation

For most creators the expensive hours are not writing. They are the assembly line after the recording: finding the moments worth cutting, reframing to vertical, captioning, writing something for each platform, and posting. Those are all verifiable-output tasks, which is why automating them holds up better than automating ideas. If you record long-form and publish short, the single biggest saving available to you is not a writing assistant — it is not manually scrubbing a two-hour file looking for the good bits. That is the job FrameOS does, and it is worth being specific that it is a mechanical saving, not a creative one.

Categories worth a subscription, honestly ranked

Transcription and captioning first: the output is checkable, the manual alternative is genuinely slow, and the accuracy is measurable. Clipping and reframing second, for the same reason. Dictation third, if you draft a lot of text — the speed advantage there has been measured properly rather than surveyed. Image and thumbnail generation fourth, useful but with a quality ceiling you will hit. Scheduling assistants fifth: the AI in your scheduler is almost always a rewriter, and rewriting is cheap. Strategy and analytics assistants last, not because they are useless but because their output is the hardest to verify and the easiest to accept when it is wrong.

How to measure it on your own workflow

You do not need a research team to avoid the METR trap, you need a stopwatch and two weeks. Pick one recurring task. Do it your normal way for a week and write down the wall-clock minutes each time — not your impression afterwards, the actual minutes. Then do it with the tool for a week, recording minutes the same way, and count the repair time as part of it, because it is. If the tool wins, it will win obviously. If it is close, you are paying a subscription for a feeling. The reason to do this at all is that your intuition has been shown, under controlled conditions, to point the wrong way.

FAQ

Does AI actually save time?

For mechanical tasks with checkable output — transcription, captioning, cutting, background removal — yes, and the saving is easy to verify yourself. For open-ended work, the only randomised trial to date found experienced developers were 19% slower with AI while believing they were 20% faster. Measure your own workflow rather than trusting either the surveys or your impression.

Why do AI productivity statistics disagree so much?

Because almost all of them ask people to estimate their own time saved, and self-estimates of time are unreliable. That is why published figures range from about six to eleven hours a week without any of them being able to reconcile with the others.

What is the highest-value AI tool for a solo creator?

Whichever one removes the most mechanical minutes from your specific week. For creators publishing short clips from long recordings, that is usually clipping and captioning, because the manual version is slow and the output is instantly checkable.

Sources

Related pages