Even Muse Glimmer (as did GPT-OSS I think) does ~4 sliding window attention layers + 1 full attention layer (like Gemma 4). I’m assuming both labs have good reason to think that gated delta nets are not optimal.
Of course it’s possible the labs just stick with the optimal architecture for large models and GDN is best for smaller models.