The paper mentioned was submitted in May. I think encoder-decoder architectures still make a lot of sense for multimodal models.
There are still many low hanging fruits. I have probably seen dozens of variations of chain-of-thoughts, tree-of-thoughts, graph-of-thoughts, self-ask, self-critique, self-plan, self-reflect, etc.