here's an example of Qwen-3.6 35B A3B MoE porting my phd code to JAX with only high level guidance from my expertise, newer qwen models share the same noticeable step change in capability as recent Big Ai models
https://github.com/verdverm/pge-jax#note-from-author
are open weights lagging, yes, are they way behind, no
if open weights were so inferior, they would not be >50% of all token processing