Here's the chat template I use:
Also, Google and Facebook are still spending like drunken sailors. Nobody has stubbed their toe on hard limitations yet. So yes of course people will figure out how to optimize the cost of AI in their products. Just probably not this year.
There will be people who want to host things on-device. At some point, you could probably do most day-to-day tasks with a Siri-like agent, so you don't necessarily need it to be on a datacenter rack somewhere.
More complex tasks being run quickly opens up a choice: insanely beefy individual devices, on-prem hosting, or cloud hosting, whether that be some data center running FOSS models, or ones from people like Anthropic or OpenAI.
Beefy hardware for individual users? Not cost-effective. Could have people share that hardware by putting it in a data center. Do you want to operate that data center? For proven business cases, sure, why not? If you're still working out what your scale will be, maybe you ask the Googles, Amazons, or Microsofts of the world to rent you the hardware so you don't have wasted or too little capacity.
The real question is, how much value is there in a few companies that talk about how their eventual goal is to create AGI as opposed to just giving you enough intelligence to augment your current workers?
The answer is "probably not enough to justify more than one company having a valuation of over a trillion dollars, and that's generous".
Let's also not forget that even the open models are not getting smaller, they are getting larger. Of course, you can distill them down into something that will fit on smaller compute, but at the end of the day, the data centers of compute, still play a huge role.
There is no loyalty in AI. I can switch to another model at near zero cost. They absolutely do have to demonstrate value. The concept of "good enough" is a misnomer because we're talking about putting these products into the hands of people who need to generate value from them.
To take this argument to an absurd level, if Fable generating a feature cost $100,000 and Sonnet generating the same feature cost $1, it wouldn't be enough that Fable is marginally, significantly, or even dramatically better, because the cost difference is so outsized. For a variety of reasons, the true cost of tokens for frontier models is downplayed at multiple points (whether it's through AI companies running off funding, companies prioritizing AI adoption over AI costs, etc.)
If cost is not an issue at all, the best thing is always preferable regardless of cost. But in most cases, and AI is no exception, eventually cost will be an issue.
Also the decision makers who are signing off on things like ChatGPT Enterprise are at least 18 months behind the curve of what you can actually do with these things and how cheap they can be. They're still trying to figure out how to actually adopt the tech out of a sense of fomo, nevermind making nuanced decisions about hosting an open weights model. I see this firsthand in my own job.
I'm talking about adding facts to a model by modifying engrams or trying to bolster guardrails with J-washing, meanwhile they're still trying to figure out how to best prompt Copilot.
Give it a few years for everyone else to catch up, I'm barely able to catch my breath before there's some new development in the open source/weights space
Not only bigger, but smarter and more capable. From what I can tell, smaller are only getting smarter in very specific areas. There is a subtle difference there, that is extremely important.
> What I can do with an 8b used to require a 32b.
What exactly do you do with an 8b? I usually ask this question and either get no response or it is something that doesn't generate anything of value. So, please surprise me.
I'm using it at work in multiple data classification and redaction pipelines. It's replaced tedious manual labor and opened that staff up to focus on the parts of the work that requires their human intuition, rather than spending time on tedium.
Presumably part of the reason you don't get a response is your hostile approach to asking.
Asking a direct question is not hostile. I'm glad to hear you've found a use for a smaller model that generates value. It gives me hope for the future, that said, I think we are still a long ways away from needing HPC in DC's.
The internet was great before AOL, but the simplification is what brought the masses to it.