SWE-bench
2,294 real-world GitHub issues and their verified patches from 12 popular Python repositories. The gold standard benchmark for evaluating LLMs on practical software engineering tasks — writing actual code that resolves real bugs in production codebases.
from datasets import load_dataset
ds = load_dataset("princeton-nlp/SWE-bench")Preview a sample row
{
"instance_id": "astropy__astropy-12907",
"repo": "astropy/astropy",
"issue": "ModelBoundingBox fails when input_shape is passed as a tuple",
"patch": "diff --git a/astropy/modeling/bounding_box.py...",
"test_patch": "def test_bounding_box_tuple_input_shape..."
}Frequently asked questions
Can I use SWE-bench commercially?
Yes — SWE-bench is released under MIT, a permissive license that allows commercial use, including training models you ship in a product. Check the dataset card for attribution requirements before release.
How much data does SWE-bench contain, and do I need all of it?
SWE-bench contains 2,294 Tasks. It is an evaluation benchmark, so it is used in full to measure models — never mix it into training data, or your benchmark scores become meaningless.
What is SWE-bench best used for?
Benchmarking real-world software engineering - never train on it. It belongs to the Evaluation & Benchmarks section of our dataset hub, where you'll find alternatives and complementary sets.