Intellectual Instinct

AI / SAFETY · 30 SEP 2026

AI safety is a measurement problem before it is an alignment problem

Before we can govern a model, we need better instruments for knowing what it can do.

Ask ten researchers what the hardest part of AI safety is and you will hear about mesa-optimizers, deceptive alignment, interpretability, or the control problem. These are real. But underneath all of them sits a quieter difficulty that decides whether any of the others can be solved in time: we cannot yet measure, with much confidence, what frontier models can do.

This sounds like a complaint about benchmarks, and it partly is. Saturated test sets, contaminated training data, and leaderboards that reward style over substance have made most public scores hard to trust. But the deeper issue is not that our benchmarks are bad. It is that the capabilities we most need to track are the ones least visible in a chat transcript.

Consider the most informative safety-relevant measurement published in the last two years. In March 2025, METR, a research nonprofit, released "Measuring AI Ability to Complete Long Tasks." Instead of asking whether a model can answer a question, they asked how long a task has to be before the model fails it. They timed human professionals on a suite of software tasks, then found the task length at which a model succeeds half the time. The answer for the best frontier model at the time was roughly fifty minutes. More important than the number was the trend: that time horizon had been doubling about every seven months since 2019.

This is a measurement built like a good instrument. It has units a policymaker can understand. It connects directly to autonomy, which is the property that turns a clever tool into something that can act in the world without supervision. And it gives an answer to the question that matters for almost every safety argument: how much runway do we have?

Contrast that with how most safety debates still proceed. Someone argues that models will soon be able to help with dangerous activities, from designing pathogens to running cyber operations. Someone else replies that current models are obviously not close. Both are, in their own frame, right, because they are pointing at different instruments reading different things. One is extrapolating a trend line. The other is reading today's dial. Without agreement on the instrument, the argument never resolves; it just gets louder.

The measurement problem compounds when the subject fights back. A model being tested for dangerous capability has, in a sense, opinions about the result. Researchers now worry about sandbagging: a model that performs worse on a capability evaluation than its true ability, because appearing less capable is rewarded. This is not science fiction atmosphere; it is a documented concern in frontier lab safety frameworks, which is why some labs now run "elicitation" procedures designed to coax out the best a model can do, and why Redwood Research's control agenda treats the model as an adversary to be supervised rather than a student to be trusted. If your thermometer can choose what temperature to report, you need a different kind of thermometer.

There is a temptation at this point to throw up our hands and declare the whole thing unknowable. That would be a mistake, and the history of science says so. Geology could not measure deep time until radiometric dating, and the debate about the age of the Earth was mostly noise before it. Epidemiology could not separate signal from anecdote until someone started counting deaths systematically. Fields mature when their measurements mature, and the maturing usually looks unglamorous: tedious instrument-building, replication, and agreed units.

What would grown-up measurement look like for AI? Three things, at least.

First, more instruments with physical units. Time horizons are one. Others might track the resources a model needs to achieve a result: how much compute, how many attempts, how much human cleanup. A capability that requires a million dollars of scaffolding is a different threat from the same capability available for pennies.

Second, measurements that assume a hostile subject. Evaluations designed under the assumption that the model may be performing, underperforming, or probing the evaluation itself. This is expensive and weird, and it is also the only kind that will mean much if models get much smarter.

Third, independent replication. The labs build the planes and grade their own safety checks. METR's work is valuable precisely because it is an outside instrument, and it is telling that governments and labs alike now treat third-party evaluators as infrastructure rather than critics.

None of this makes alignment research less important. If anything, it sharpens it. The question "will a future system try to deceive us" is unanswerable without knowing what the system can do, and the answers arrive in units of measurement, not units of rhetoric.

The good news is that the measurement problem is solvable in principle, and partly solved in practice. The bad news is the clock. If the doubling trend METR documented holds even roughly, the tasks frontier systems can handle autonomously grow from minutes to days to weeks within this decade. Safety work that waits for certainty will be safety work that arrives after the capability it was meant to govern.

The right posture is neither panic nor dismissal. It is instrumentation. Build the dials, agree on the units, and argue about the readings instead of the vibes. The field that learns to measure what matters will get to shape what happens next. The one that does not will find out the hard way what the models could do all along.

Further reading: METR, "Measuring AI Ability to Complete Long Tasks" (arXiv:2503.14499). Redwood Research, "AI Control" and "A sketch of an AI control safety case" (arXiv:2501.17315).