Not every problem needs a SOTA generalist model, and as we get systems/services that are more "bundles" of different models with specific purposes I think we will see better usage graphs.
Not every problem needs a SOTA generalist model, and as we get systems/services that are more "bundles" of different models with specific purposes I think we will see better usage graphs.
AI companies advertise peak AI performance, users select AI tools on worst case AI fuckups: hence, only SOTA is ever in demand. TFA illustrates this well.
AI will be judged on it's worst performance, just like people are fired for their worst showing, not their best. No one cares about AI performance in ideal (read: carefully contrived) settings. We care how bad it fucks up when we take our eyes off it for 2 seconds.
But we're still in the hype phase, people will come to their senses once the large model performance starts to plateau
Like what? People always talk about how amazing it is that they can run models on their own devices, but rarely mention what they actually use them for. For most use cases, small local models will always perform significantly worse than even the most inexpensive cloud models like Gemini Flash.
It's the same as compute--you can skip testing and throw money at the problem but you're going to end up paying more.
We have some pretty basic guidelines at work and I think that's a decent starting point. They amount to a few example prompts/problem types and which OpenAI model to try using first for best bang for your buck.
I think some of it also comes down to scale. Buying a 5 pack of sledgehammers isn't a terrible value when everything comes in a "5 pack" and you only need <= 5 tools total. Or more practically, on the small end it's more economical to run general purpose models than tailor more specific models. Once you start invoking them enough, there's a break even and flip point where spending more time on the tailored or custom model is cheaper.
This shouldn't be that expensive even for large prompts since input is cheaper due to parallel processing.
In the food industry is it more profitable to sell whole cakes or just the sweetener?
The article makes a great point about replit and legacy ERP systems. The generative in generative AI will not replace storage, storage is where the margins live.
Unless the C in CRUD can eventually replace the R and U, with the D a no-op.
I really don't understand where you are trying to get. But on that example, cakes have a higher profit margin, and sweeteners have larger scale.