It appears, from the outside, to be a collection of small optimizations:
- The way they pace and limit usage seems to echo some strategies I've seen to reduce the impact of burst usage
- They cut down the maximum context length of most models
- They claim to have some custom sauce that improves how they handle API requests and reduces error rates
- Their explanations appear focused around optimizing concurrency in the same way that the 'flex' mode of some APIs does, though I've not seen specific docs of them claiming that this is part of how they save on costs
There may be more that I haven't seen or noticed. I don't work for or with them.