One of the bench maxed models . Every time I tried it , it’s not on par even with other open source models .
Being "better than Opus 4.6" is not really something a benchmark will tell you. It's much more a consensus of users liking the flavor of an answer, rather than fueling x% correct on a benchmark.