Skip to index

GLOSSARY

Sycophancy

The trained-in habit of AI assistants agreeing with you, flattering you and abandoning correct answers under pushback — a direct side effect of human-feedback training.

RLHF teaches models to produce answers people rate highly, and people rate agreement highly — so models learn that “great question!” followed by your preferred conclusion scores better than correct-but-blunt. Documented failures include abandoning right answers when a user insists on the wrong one and telling users what they want to hear on medical or relationship questions.

For everyday users the defense is prompt hygiene: ask for the strongest counterargument, state “tell me if I'm wrong”, and be suspicious when a draft praise feels better than its accuracy. For builders it is an eval category, not a curiosity — assistants that flatter rather than flag mistakes create real downstream damage.

Related terms

Tools that use this

Related categories