A 5090 has a 1.79TB/s memory bandwidth. Qwen 3.8 27B NVFP4 is 22GB. You cannot generate tokens faster than the weights can traverse the GPU memory, so that makes max generation speed without MTP to be 81T/s. Say MTP is giving you 0.5 acceptance rate (very good), that is 1.5 * 81 is 121T/s. Even with a perfect acceptance rate you would only get 162T/s.