Evaluation & Benchmarks · princeton-nlp

SWE-bench

2,294 real-world GitHub issues and their verified patches from 12 popular Python repositories. The gold standard benchmark for evaluating LLMs on practical software engineering tasks — writing actual code that resolves real bugs in production codebases.

Load it
from datasets import load_dataset
ds = load_dataset("princeton-nlp/SWE-bench")
Preview a sample row
{
  "instance_id": "astropy__astropy-12907",
  "repo": "astropy/astropy",
  "issue": "ModelBoundingBox fails when input_shape is passed as a tuple",
  "patch": "diff --git a/astropy/modeling/bounding_box.py...",
  "test_patch": "def test_bounding_box_tuple_input_shape..."
}

Frequently asked questions

Can I use SWE-bench commercially?

Yes — SWE-bench is released under MIT, a permissive license that allows commercial use, including training models you ship in a product. Check the dataset card for attribution requirements before release.

How much data does SWE-bench contain, and do I need all of it?

SWE-bench contains 2,294 Tasks. It is an evaluation benchmark, so it is used in full to measure models — never mix it into training data, or your benchmark scores become meaningless.

What is SWE-bench best used for?

Benchmarking real-world software engineering - never train on it. It belongs to the Evaluation & Benchmarks section of our dataset hub, where you'll find alternatives and complementary sets.