SWE-bench
A benchmark for evaluating whether language models can resolve real GitHub issues in existing Python repositories.
Source: SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.
The paper first appeared as a preprint in 2023 and was published at ICLR 2024.
Referenced in An Empirical Study of Harness Design for Coding Agents.