For prefilling, DeepSeek reports that they use 32 GPUs for this stage, meaning that each GPU holds around 8 routed experts and one shared
The flip side - 32 GPUs for prefill, vs 144 GPUs total for decode.
The more GPUs used the lower the number of experts.
Works about to be about two experts per GPU for decode, and 8+ per GPU for prefill