Honestly, given the price 5.6 Luna is pretty fantastic. I replaced GPT 4.1 mini in several projects.
It's also 50% cheaper if you only need it to run sometime in the next 24 hours.
It's also 50% cheaper if you only need it to run sometime in the next 24 hours.
Flex is also easier to get caching to work, there is a little futzing around with OAI's implicit caching but if you do the upfront work you can get haiku quality responses with caching in close to realtime for 1/7th the cost and you don't need to design a polling loop to check for batch completions