METR's finding that AI task-completion time horizons double roughly every seven months is the most-cited capability chart in AI right now — here's what the 50%-reliability metric behind it actually measures, and what it quietly leaves out.
If you've seen a chart claiming AI models can now complete tasks that would take a human "several hours," it almost certainly traces back to one source: METR, the nonprofit AI evaluation group, and its "time horizon" metric. In March 2025, METR published research showing that the length of task a frontier model can complete with 50% reliability has been doubling roughly every seven months — a trend line so clean it's become the single most-cited capability chart in AI policy decks, lab announcements, and "AGI timeline" arguments alike.
The finding itself is real and the methodology is careful. But "time horizon" is a much narrower, more specific measurement than the headline suggests, and the gap between what it measures and what people think it measures is exactly the kind of thing worth being precise about before you use it to make a deployment decision.
METR (Model Evaluation and Threat Research) built two task suites to run this evaluation: HCAST (Human-Calibrated Autonomy Software Tasks) and RE-Bench (Research Engineering Benchmark), later supplemented with additional software and general-agency tasks. Both suites share a specific property: every task was first completed by skilled human contractors, whose completion times were recorded as the baseline "human time" for that task.
Models are then given the same tasks — things like fixing a bug in an unfamiliar codebase, writing a small ML training pipeline, or automating a data-processing workflow — and scored pass/fail by an automated or human grader. For each model, METR plots success rate against the human-baseline time for each task, which traces out a logistic curve: success rate near 100% for tasks that take humans seconds, dropping toward 0% for tasks that take humans many hours. The "time horizon" for that model is the x-axis value where that curve crosses 50% — in plain terms, the task length at which the model succeeds about half the time.
Plot that single number for each model release over its release date, and you get the now-famous chart: early GPT-era models clustered around a horizon of seconds to a few minutes, GPT-4-class models reaching into the tens of minutes, and by METR's 2025 tracking, frontier models crossing into the one-hour-plus range. Fit a trend line through those points and the slope implies a doubling period of roughly 7 months (METR's paper cites a figure in that neighborhood, sometimes reported as ~212 days).
That's a genuinely interesting empirical result. It is not, however, a general statement about "AI can now do X hours of work." Here's where the benchmark's actual scope and the public's reading of it diverge.
Time horizon isn't the only benchmark trying to capture "how much can an agent actually do," and it's worth placing it next to the others making similar claims, because each one is measuring something importantly different:
| Benchmark | What it actually scores | Task domain | Reliability bar | Main methodological critique |
|---|---|---|---|---|
| SWE-bench | Resolves a real GitHub issue with a merged patch | Open-source Python repos, pre-existing bugs | Binary pass/fail per issue | Issues skew toward well-specified, well-tested repos; leakage/contamination risk from training data |
| GAIA | Answers a multi-step question requiring tool use (web, files, code) | General assistant tasks, deliberately broad | Exact-match final answer | Many tasks have a single "correct" answer that under-credits valid alternative approaches |
| τ-bench (tau-bench) | Completes a customer-service-style conversation per a policy document | Airline/retail support simulations | Pass/fail against policy compliance, often averaged over repeated trials | Narrow domain (two simulated businesses); policy-following isn't the same as general reasoning |
| METR HCAST / RE-Bench | Completes a software/ML-engineering task within a human-calibrated time budget | Software engineering, ML research engineering | 50%-success time horizon (a threshold, not pass/fail per task) | Task suite is itself narrow (eng-adjacent); 50% threshold is far below production reliability |
Each of these is defensible as a research instrument and misleading as a one-line marketing claim. METR's is distinctive because it doesn't report a score at all — it reports a threshold, which makes it feel more rigorous than a percentage, but that framing hides a few important choices.
It's a 50% bar, not a usable bar. This is the most important caveat, and METR's own paper is explicit about it: the headline "time horizon" is defined at 50% success probability. A model that succeeds on a four-hour task half the time is not a model you'd let run unsupervised — in most production contexts you need something closer to 90-99% reliability. METR also computes an 80% time horizon in the same paper, and it's dramatically shorter than the 50% figure for every model tested. If you re-ran the famous doubling chart using an 80% bar instead of 50%, the "AI can now do hours of work" story shrinks substantially, and the practical deployment threshold for anything high-stakes sits even further out than that.
The task suite is narrow by design. HCAST and RE-Bench are software-engineering and ML-research-engineering tasks, performed by contractors METR hired and timed. That's a deliberate, defensible choice for an org focused on AI R&D automation risk — it's exactly the kind of work you'd want to track if your concern is models automating their own development. It is not a representative sample of "general knowledge work," and extrapolating the doubling trend to, say, legal drafting, clinical triage, or enterprise sales operations assumes a transfer that the benchmark never tested.
Autogradable tasks aren't messy tasks. Every task in the suite needs a grader — automated or human-scored against a rubric — which pushes the suite toward work with clean success criteria. Real-world knowledge work is full of ambiguous specs, stakeholders who change their minds mid-task, and "correct" answers that depend on context nobody wrote down. A benchmark built around gradeable tasks will systematically underweight exactly the kind of friction that makes long-horizon work hard in practice.
Few data points, exponential extrapolation. The headline doubling-time figure is fit across a relatively small number of model releases spanning a few years. Fitting an exponential to roughly a dozen points and projecting it forward is a classic place where confidence intervals get quietly dropped from the public retelling. METR's researchers have been more careful about this than most of the people citing their work — the paper reports uncertainty ranges around the doubling estimate that rarely survive the trip into a conference keynote slide.
Task "length" collapses a lot of different things into one number. A two-hour task that's two hours of linear, well-specified work is not the same difficulty as a two-hour task that requires juggling five interdependent subsystems, asking a human a clarifying question, or holding a huge amount of context simultaneously. Human baseline time is a reasonable proxy for difficulty, but it flattens dimensions — parallelism, context load, need for clarification — that matter a lot for whether an agent can actually do the work unsupervised.
None of this makes the metric useless — it makes it a specific instrument rather than a universal one. As a way to track relative capability growth across model generations on agentic, code-adjacent work, it's one of the better-designed benchmarks available: human-calibrated baselines are harder to game than raw leaderboard scores, and the logistic-curve-plus-threshold approach is more informative than a single pass-rate percentage. It's also a reasonably good predictor of something people actually care about: whether a given model generation can sustain a longer autonomous coding or research session before it goes off the rails, which lines up with what practitioners have anecdotally observed shipping agent-based coding tools over the same period.
The honest summary is that METR's time horizon is a well-built, narrow-scope research metric that measures 50%-reliability task duration on software- and research-engineering work, and it has been asked to carry the weight of a general claim about AI capability growth that it was never designed to support.
If you're deciding whether an agent is ready to handle a longer-running task in your own product, don't import METR's time horizon number as a guarantee — build the equivalent evaluation for your actual task distribution. Collect a human-timed baseline for the real tasks your agent needs to do, run your candidate model against them, and plot success rate against task length the same way METR does. Then pick your threshold based on how much supervision the task will realistically get in production, not 50%. A benchmark built from your own failure modes will tell you far more than a borrowed trend line, however good the trend line's methodology is.