I think a big part is which model seems to work better with your language/stack. My language is Elixir, which is somewhat niche, and only Claude has been able to produce usable Elixir code so far. None of the other things mentioned in the article mattered, because of this. I wonder if others have this experience that some models just struggle with some languages/stacks?