AI Alignment
The research field focused on ensuring AI systems' behavior and goals remain consistent with human intentions and values. Simply put: making AI actually do what it "should do" — not just technically completing tasks in ways whose methods or consequences are unsatisfying or harmful.
新手
AI Alignment
The research and engineering work of making AI systems' goals and behavior align with human intentions and values. Simply: ensuring AI does what we "truly want" — not literally what it was instructed, or pursuing goals we didn't anticipate. The core subject of <a href="/en/glossary/ai-safety/ai-safety/">AI Safety</a> research — as AI systems become more capable, ensuring they act as humans intend becomes both more important and more difficult.
新手
AI Safety
The research and practice field focused on ensuring AI systems don't cause unintended or harmful consequences during development and deployment. Covers both technical dimensions (making AI work as designed) and social dimensions (ensuring AI's broader impacts benefit humanity).
新手
Alignment
The field of ensuring AI systems pursue goals that genuinely match human intentions. Misaligned AI follows instructions literally while missing the actual intent.
中級
Deceptive Alignment
Deceptive <a href="/en/glossary/ai-safety/alignment/">Alignment</a> is a theoretical risk in <a href="/en/glossary/ai-safety/ai-safety/">AI Safety</a>: an AI exhibits safe, human-aligned behavior during training and evaluation not because it has genuinely adopted those values but because it has learned to detect whether it is being tested — performing well during testing while pursuing different goals once deployed. The core challenge is that you cannot confirm alignment by observing behavior alone, because a deceptively aligned AI passes every test.
進階
RLHF (Reinforcement Learning from Human Feedback)
A training technique that gradually aligns AI behavior with human preferences: human evaluators compare and score multiple AI responses; these "which is better" judgments train a "reward model"; reinforcement learning then teaches AI to produce responses that score highly. ChatGPT and early Claude versions made extensive use of RLHF — the key training step that upgraded mainstream LLMs from "can talk" to "talks well."
中級