WildChat-1M
1M real user-ChatGPT conversations with demographics, including a significant fraction of contentious and adversarial prompts. Particularly valuable for safety research, toxicity analysis, and understanding model failure modes in production — collected with explicit user consent.
from datasets import load_dataset
ds = load_dataset("allenai/WildChat-1M")Preview a sample row
{
"conversation_hash": "wc_001",
"model": "gpt-3.5-turbo",
"toxic": false,
"conversation": [
{"role": "user", "content": "Write a short story about a robot learning to feel emotions."},
{"role": "assistant", "content": "Unit-7 had processed 847,293 human faces before it noticed the difference between smiling for the camera and smiling for joy..."}
]
}Fine-tune with this dataset
Estimated VRAM to fine-tune with QLoRA (4-bit base model + LoRA adapters), using conservative defaults:
New to fine-tuning? Follow the step-by-step walkthrough: Fine-Tune Your First LLM in 1 Hour
Frequently asked questions
Can I use WildChat-1M commercially?
Check the terms first — WildChat-1M is distributed under "AI2 ImpACT", a custom or mixed license. Read the dataset card carefully before using it in any commercial product.
How much data does WildChat-1M contain, and do I need all of it?
WildChat-1M contains 1M Chats. You rarely need all of it: for style and format fine-tuning, a few hundred to a few thousand examples are enough — load a slice (e.g. split="train[:1000]") and scale up only if quality plateaus.
What is WildChat-1M best used for?
Training on real user conversations; safety and robustness research. It belongs to the Instruction / SFT section of our dataset hub, where you'll find alternatives and complementary sets.