They weren't the first to do MTP like this, and arguably did it wrong: the MTP heads are kept in a separate file and have to be welded in by the inference engine.
Qwen 3.6 shipped with working MTP first, and had working MTP in llama.cpp first.
Qwen 3.6 shipped with working MTP first, and had working MTP in llama.cpp first.
Ultimately though the real explanation, I think, is Google doesn't care since for their own purposes (in LiteRT-LM), they do bundle them. As far as I know, anyway.
They are more like a single model that has two separate attention head mechanisms.