GLOSSARY
AI Safety & Alignment
The practice of making AI systems helpful, honest and harmless — following human intent and values rather than producing capable but harmful outputs.
Alignment is the effort to make AI systems pursue what humans actually intend. Raw models predict text; alignment training (human feedback, constitutional methods, red-teaming) shapes them into assistants that refuse harmful requests, admit uncertainty and follow instructions faithfully.
It is also a moving target: jailbreaks evolve, capabilities expand into agents that act in the world, and policies like Anthropic's 'harmlessness' research or OpenAI's safety tiers formalize the practice. For users, alignment shows up as the guardrails you notice — refusals, warnings, content filters — whose quality marks the difference between careful labs and careless ones.