I’m curious how you guys keep track of each model’s coding capabilities. The landscape keeps changing. I don’t suppose you benchmark all frontier models every other month, right?
Most of the time I don't need what the bench tests and I'm not really giving them completely ambiguous tasks without any refinement.
I only find marginal differences between models at this point and it almost feels like personality quirks in each model than anything.
alias agy="agy --dangerously-skip-permissions"