Delegate by Task Length, Not by Task Type, practitioner guidance from TheAICommand
← AI News
Capability

Delegate by Task Length, Not by Task Type

METR measures AI capability as the length of task a model finishes at a given success rate. Read properly, it is the most useful delegation tool available, and it says something uncomfortable about regulated work.

·TheAICommand

Quick answer

Ask three questions before delegating to AI: how long the task would take a capable person, how reliable the output must be, and how quickly a wrong answer would be caught. METR's data puts the 80 per cent reliability horizon four to six times shorter than the headline 50 per cent figure, so size handovers against the reliability you need.

Stop asking what AI can do. Ask how long.

Almost every conversation about handing work to AI is organised around task type. Can it do research. Can it do drafting. Can it do analysis. The category is the wrong unit, because the same category contains a ten minute job and a three day job, and models behave completely differently across that range.

The research group METR has spent two years building a measure that uses a better unit. Its 50 per cent time horizon is the length of task, measured by how long a skilled human takes, at which a model is predicted to succeed half the time. Not how long the AI runs. How hard the task is, expressed in human hours.

The headline finding from the paper is that this length has been doubling roughly every seven months. Precisely, 207 days, with a 95 per cent bootstrapped confidence interval of 166 to 240 days. The measurement covers 170 tasks with more than 800 human baselines totalling 2,529 hours, across 12 frontier and 4 near frontier models released between 2019 and 2025. The paper has been revised four times, most recently on 10 July 2026, and METR now runs an updated methodology with a larger task suite on a public tracker.

That doubling curve is the part that gets shared. It is not the part that should change how you work.

The number that actually governs delegation

METR also reports the 80 per cent time horizon: same idea, higher reliability bar. Those horizons are four to six times shorter than the 50 per cent ones. The doubling time is almost identical, 204 days against 207, so the gap is not closing. It is a stable property of how these systems fail.

Measure50 per cent horizon80 per cent horizon
What it meansThe model succeeds half the timeThe model is right four times in five
Relative lengthThe headline figure on the chartFour to six times shorter
Doubling time207 days204 days

Sit with that. A model at the frontier of a chart might carry a multi hour 50 per cent horizon and a horizon measured in tens of minutes once you require it to be right four times in five. If your working assumption came from the headline figure, you have been sizing handovers using a number calibrated to a coin flip.

METR is unusually direct about this. Its own limitations note from 22 January 2026 says a 50 per cent time horizon of X hours does not mean tasks under X hours can be delegated to AI. It goes further: for work that is reliability critical and hard to verify, the group suggests success probabilities above 98 per cent before automation is worth it. There is no line on the public chart for a 98 per cent horizon, which tells you something about how short it would be.

Split composition contrasting a long measuring bar at fifty per cent reliability with a much shorter one at eighty per cent
The reliability discount does not shrink as capability grows

What the measure is not

Three misreadings are common enough to be worth naming.

It is not a measure of how long an AI can work unattended. METR frames it as difficulty, not duration of operation. A model with a four hour horizon is not a system you leave running for four hours. It is a system that can handle problems a person would need about four hours to solve, at the stated success rate.

It is not general across domains. METR reports horizons are broadly similar for mathematics but dramatically lower for visual computer use tasks, by a factor it puts at 40 to 100. Anything that runs through a graphical interface, and that includes a large share of ordinary office work, sits far below the software engineering numbers on the chart.

It is not precise. METR states that error bars have historically been around a factor of two in each direction, and worse for recent models. The tracker itself carries a note that measurements above 16 hours are unreliable with the current task suite, which is a striking piece of self limitation on a chart whose top end is the bit everyone screenshots. Its March 2026 methodology work found that reasonable alternative modelling choices can cut recent 50 per cent estimates by up to 35 per cent and raise 80 per cent estimates by up to 100 per cent.

And the benchmark is software tasks. The authors say plainly that their tasks do not perfectly represent the average segment of intellectual labour, and raise external validity as an open question. They also note that most of their tasks are self contained, while the long tasks humans actually do require collaboration. Extending the metric to compliance work, case management or policy drafting is an inference. It is a reasonable one, and it is still ours, not theirs.

The three question sizing rule

The useful version of all this is not a forecast. It is a way to size a handover before you make it.

How long would this take a competent person with no special context? METR notes the horizon is closer to what a low context person could do, someone like a new starter or a contractor working from a written brief. That is the right mental model. Not your best analyst who knows the history. Someone capable, working from the material you would actually hand over.

How reliable does the answer have to be? Be honest, and be specific. A brainstorm that a person will heavily edit can run at 50 per cent and still save time, because the failures are cheap and obvious. A figure that goes into a customer letter cannot. This is the question that most often gets skipped, because nobody wants to write down that their process tolerates a one in five error rate.

How quickly would a wrong answer be caught? This is the multiplier. A wrong answer that fails loudly, that will not compile or does not reconcile, is survivable at low reliability. A wrong answer that is fluent, plausible and buried in paragraph nine of a twelve page document is not, because the reliability of the whole workflow collapses to the reliability of the person skimming it.

Put the three together and the rule falls out. Short task, low stakes, obvious failures: delegate the whole thing. Long task, high stakes, quiet failures: delegate a bounded slice and verify at the boundary. The instinct to split work into stages is not project management hygiene. It is a direct consequence of the shape of the curve, because a series of short verified steps sits far higher on the reliability axis than one long unverified one. If your team has a written delegation charter, this is the measurement that should sit behind its categories.

Why does this bite hardest in regulated work?

Australian regulated work is close to the worst case on two of the three questions at once.

Its outputs are reliability critical. A determination, an assessment, a disclosure, a suitability statement or a regulatory notification is not a draft that gets fixed later. Errors are consequential and often not reversible.

And its errors are poorly verifiable. Verifying that a claim decision correctly applied a statutory test is not like running a test suite. It requires a person with the expertise to re examine the reasoning. In practice, that verification is exactly the labour the delegation was meant to save, which is how organisations end up paying twice.

That combination is precisely the case where METR suggests a very high success bar before automation is worth it. Read against the published horizons, that does not point at abandoning AI in regulated processes. It points at where the value actually sits: in the bounded, verifiable, shorter tasks that surround the decision rather than the decision. De identifying a file. Building a chronology from source documents. Assembling an evidence map. Drafting a structure to a fixed template. Each is short, each has a checkable output, and none of them asks the model to be right about the thing that matters.

That is also the pattern that survives a model change. If your workflow depends on a frontier model clearing a long unverified task, every vendor update is a revalidation event, the same drift problem that passing the pilot does not solve. If it depends on a series of short checked steps, it degrades gracefully.

Build your own ruler

METR's chart is measured on software engineering tasks. Yours are not those tasks. The metric is still usable, but only if you calibrate it against your own work, and that is a smaller exercise than it sounds. It is the delegation-sizing cousin of evaluating an AI tool on your own work before you buy it.

Pick eight to ten tasks that genuinely represent what you would want to hand over. Include a spread, from something a capable person finishes in fifteen minutes to something that takes most of a day. Write each one as a brief that a competent newcomer could work from, because that is the standard the metric assumes. Then have someone actually do them and record the time honestly, including the part where they go looking for the missing input.

You now have a ruler in your own units. When you consider delegating something, you are no longer guessing whether it is a long task. You can place it against work you have timed.

The second half is the harder half. Run the same tasks through the model, several times each, and count how often the output is acceptable without rework. Three runs will not give you a defensible success rate, but the pattern is usually obvious well before you reach statistical comfort, and the tasks where the model is confidently wrong will announce themselves. Keep the set. It becomes the evaluation you re run when the model underneath you changes, which is the only way to know whether a vendor update helped you or quietly moved your boundary.

This is deliberately modest. It is not a benchmark, it is a ruler and a smoke test, and it costs a day. The alternative is sizing delegation against a public chart built on tasks that are not yours, at a reliability level you never chose.

The hype check

The doubling curve invites extrapolation, and extrapolation is where the honest reading breaks down. METR itself warns that speculating about months long or years long horizons is fraught, and that the benchmark cannot really address those scenarios because long human projects involve collaboration its tasks do not model.

Treat the trend as a planning input with wide error bars, not a schedule. The confidence interval on the doubling time alone spans 166 to 240 days, before you account for the factor of two on individual measurements and the modelling sensitivity METR published in March. A curve that reliable is useful for deciding that your delegation boundary should be revisited every six months. It is not useful for deciding what to do in 2029.

Do this Monday

Take one workflow you have been thinking about handing to AI. Write down the three answers. How many hours for a capable person working from the brief. What success rate the output actually requires. How fast a wrong answer surfaces.

Then find the natural break points and delegate to the first one, with a check that a person can perform quickly. You will usually find the boundary sits earlier than expected, and that the earlier boundary is the one that holds when the model underneath you changes.

Bottom line

The useful question was never whether AI can do the task. It is how long the task is, how right it has to be, and who finds out when it is not. METR's 80 per cent horizon, four to six times shorter than the headline number, is the figure to size against, and a small ruler built on your own timed tasks beats a public chart built on someone else's.

TheAICommand. Intelligence, At Your Command.

Frequently asked questions

What is METR's time horizon metric?
The length of task, measured by how long a skilled human takes, at which a model is predicted to succeed at a stated rate. The 50 per cent horizon has doubled roughly every seven months, precisely every 207 days with a 95 per cent confidence interval of 166 to 240 days, measured over 170 tasks, more than 800 human baselines and 16 models released between 2019 and 2025.
Why does the 80 per cent horizon matter more?
Because most real work cannot tolerate a coin flip. METR's 80 per cent horizons are four to six times shorter than its 50 per cent horizons, and the gap is stable: both double at almost identical rates, 204 against 207 days. If your delegation assumption came from the headline figure, you have been sizing handovers at a reliability level you never chose.
Can the metric be applied outside software tasks?
Only as an inference, and METR says so. The benchmark is software tasks, horizons are 40 to 100 times lower for visual computer-use tasks, error bars have historically been around a factor of two in each direction, and the tracker marks measurements above 16 hours as unreliable. Extending it to compliance, case management or policy drafting is reasonable, but it is your inference, not METR's finding.
What does this mean for regulated Australian work?
Regulated outputs are reliability critical and poorly verifiable, which is close to the worst case on the sizing questions. METR suggests some such tasks need success probabilities above 98 per cent before automation is worth it. The value therefore sits in bounded, verifiable short tasks around the decision, such as de-identification, chronology building and evidence maps, not in the decision itself.
How do I calibrate delegation for my own team?
Build your own ruler. Pick eight to ten representative tasks, write each as a brief a capable newcomer could work from, time a person doing them honestly, then run the same tasks through the model several times and count how often the output is acceptable without rework. Keep the set and re-run it when the underlying model changes.

Tags

evaluationdelegationreliabilityagentsbenchmarks
← Back to AI News