so don't use it at max? The benchmarks suggest that high/xhigh are more than sufficient to be ahead and a whole magnitude below max with regards to token usage. I'd treat that as an outlier and not how verbose the model is in general (QED I know)
mean median
model
5 4.135 4.245
5.5 3.150 2.640
Again seeing how max is a clear outlier, the median cost saving is ~38%, not that far off from the proclaimed 40%.