Reasoning · meta-math

MetaMathQA

395K high-quality mathematical question-answer pairs created by augmenting GSM8K and MATH through answer rewriting, backward reasoning, and FOBAR methods. Powers MetaMath models that significantly outperform base models on math benchmarks.

Load it
from datasets import load_dataset
ds = load_dataset("meta-math/MetaMathQA")
Preview a sample row
{
  "query": "James buys 3 shirts for $60. There is a 40% off sale. How much did he pay per shirt after the discount?",
  "response": "The original price per shirt is $60/3 = $20. The discount per shirt is $20 * 0.40 = $8. So James paid $20 - $8 = $12 per shirt. The answer is 12.",
  "type": "GSM_Rephrased"
}

Fine-tune with this dataset

Estimated VRAM to fine-tune with QLoRA (4-bit base model + LoRA adapters), using conservative defaults:

7B QLoRA · ~6GB VRAM13B QLoRA · ~10GB VRAM
Check if your GPU can fine-tune this →

New to fine-tuning? Follow the step-by-step walkthrough: Fine-Tune Your First LLM in 1 Hour

Frequently asked questions

Can I use MetaMathQA commercially?

Yes — MetaMathQA is released under MIT, a permissive license that allows commercial use, including training models you ship in a product. Check the dataset card for attribution requirements before release.

How much data does MetaMathQA contain, and do I need all of it?

MetaMathQA contains 395K Pairs. You rarely need all of it: for style and format fine-tuning, a few hundred to a few thousand examples are enough — load a slice (e.g. split="train[:1000]") and scale up only if quality plateaus.

What is MetaMathQA best used for?

Boosting GSM8K/MATH-style math skills in 7B models. It belongs to the Reasoning section of our dataset hub, where you'll find alternatives and complementary sets.