Single prompting a very complex tasks is rare even on frontier models, because it can be done successfully only for specific situations (e.g. you have a very strong verification step the model can iterate on).
Most of my everyday usage is for smaller takes, were you don't really get the benefit of the most expensive models, and my guess is that is the case for the most users