At this point model upgrades do not mean too much for an established use case. I have 36 benchmark scenarios using agents + tools in my app and the results were 30/36 for gpt 6 luna, and 33/36 for gpt 5.6 luna. The benchmark was tuned for gpt 5.6 luna but still apart from slightly reduced cost I will keep the default model to gpt 5.6.