Evaluation & Benchmarks · hendrycks

MATH (Hendrycks)

12,500 competition mathematics problems across 7 subjects (algebra, calculus, number theory, etc.) with step-by-step solutions. The key benchmark for evaluating mathematical reasoning in LLMs — requires multi-step problem-solving and symbolic manipulation.

Load it
from datasets import load_dataset
ds = load_dataset("hendrycks/competition_math")
Preview a sample row
{
  "problem": "Find the integer $n$, $0 \\le n \\le 11$, such that $n \\equiv 10389 \\pmod{12}$.",
  "solution": "Since $10389 = 865 \\cdot 12 + 9$, the remainder when 10389 is divided by 12 is $\\boxed{9}$.",
  "level": "Level 1",
  "type": "Number Theory"
}

Frequently asked questions

Can I use MATH (Hendrycks) commercially?

Yes — MATH (Hendrycks) is released under MIT, a permissive license that allows commercial use, including training models you ship in a product. Check the dataset card for attribution requirements before release.

How much data does MATH (Hendrycks) contain, and do I need all of it?

MATH (Hendrycks) contains 12,500 Problems. It is an evaluation benchmark, so it is used in full to measure models — never mix it into training data, or your benchmark scores become meaningless.

What is MATH (Hendrycks) best used for?

Benchmarking competition math ability. It belongs to the Evaluation & Benchmarks section of our dataset hub, where you'll find alternatives and complementary sets.