Interesting they chose to split prefilling into its own independent service, I hadn't heard of that technique before. I found this paper that researches why that could be beneficial: https://arxiv.org/abs/2401.09670v1
There are a lot of shared prefixes from user prompts. You can save by first looking into the cache to find the longest prefix for a prompt. Their MLA makes the KV cache particularly efficient.