What is AI alignment?
AI alignment is the effort to make an AI system’s behavior consistent with human goals, values, and intentions. It involves deciding which goals and values should guide the system and testing whether its behavior follows them. Google DeepMind’s value-alignment analysis separates the technical question of implementing values from the normative question of choosing them.
An AI system can follow an instruction and still produce a result its user never intended. AI alignment concerns that gap between a system’s behavior and human goals and values. The question becomes more pressing as systems take on more consequential tasks.

Why is AI alignment important?
Alignment matters because literal instructions can produce unintended, harmful outcomes. The goal is to make the system’s actions serve the intended purpose and respect the people affected.
Trust should follow evidence of behavior. A statement that a system acts in a user’s interests is a claim to examine, not a substitute for testing.
If artificial general intelligence is the goal, AGI alignment must be a companion effort.
Why is AI alignment difficult?
One difficulty is deciding which human values a system should follow. People, cultures, and societies can disagree. Turning those values into instructions that work across unfamiliar situations is a further challenge.
A system can produce unintended behavior when instructions or goals leave important details unspecified. Human feedback can help shape behavior, but whose feedback is used and how it is applied remain part of the alignment problem.
Researchers explore ways to specify appropriate goals and to learn from human examples or feedback. Each approach still requires decisions about whose values and examples should guide the system.
How should alignment be evaluated?
Alignment requires attention to how a system behaves, including when its instructions are incomplete or conflict. A statement that a model is helpful or safe needs to be tested against the decisions it actually makes.
For a user, the practical question is whether the system’s actions serve the intended purpose and respect the people affected by them.