Anthropic HH-RLHF — LLM 偏好(RLHF / DPO) Dataset

170k human preference comparisons between two AI assistant responses, with one labeled 'chosen' (helpful, harmless) and one 'rejected'. The dataset that defined RLHF alignment for chat assistants. Used to train early Claude models and became the standard reference for alignment research and DPO fine-tuning.

Dataset Details

ProviderAnthropic
Category偏好(RLHF / DPO)
Size170k Pairs
LicenseMIT
Downloads2.8M
TagsRLHF, Alignment, Safety, Helpfulness, Human-Feedback
from datasets import load_dataset
ds = load_dataset("Anthropic/hh-rlhf")

用这个数据集微调

使用 QLoRA(4-bit 基础模型 + LoRA 适配器)微调的预计显存需求(保守默认参数):

7B QLoRA~6GB VRAM
13B QLoRA~10GB VRAM

检测你的 GPU 能否微调 →

微调新手?跟着分步教程走: 一小时微调你的第一个 LLM

相关数据集

常见问题

Anthropic HH-RLHF 可以商用吗?
可以——Anthropic HH-RLHF 采用 MIT 宽松许可证,允许商业使用,包括训练用于产品的模型。发布前请查看数据集卡片中的署名要求。
Anthropic HH-RLHF 有多少数据?需要全部使用吗?
Anthropic HH-RLHF 包含 170k Pairs。通常不需要全部:风格和格式微调只需几百到几千条样本——先加载切片(如 split="train[:1000]"),质量到达瓶颈时再扩大规模。
Anthropic HH-RLHF 最适合做什么?
Safety-focused preference training (helpfulness and harmlessness)。它属于数据集中心的「偏好(RLHF / DPO)」板块,那里有替代和互补的数据集。

← 全部数据集 | Fine-Tuning Guide