this reads as a very low quality and probably fully llm written post.
the analysis is very suspicious: “gpt 5 mini had api failures due to wrong temp setting”? wtf?
whatever you used to slop your benchmark didt even take the time to set the temp to 1 (which the docs say is required)