They’re giving away these expensive models because it doesn’t hurt Meta’s ad business but reduces risk that competitors could grow moats using closed models. The old “commoditize your complements” play.
TRAINING an LLM requires a lot of compute. Running inference on a pre-trained LLM is less computationally expensive, to the point where you can run LLAMA (cost $$$ to train on Meta's GPU cluster) with CPU-based inference.
While it does appear that a large number of people are using these for RP and... um... similar stuff, I do find code generation to be fairly good, esp with some recent Qwens (from Alibaba). Full disclaimer: I do use this sparingly, either to get the boilerplate, or to generate a complete function/module with clearly defined specs, and I sometimes have to re-factor the output to fit my style and preferences.
I also use various general models (mostly Meta's) fairly regularly to come up with and discuss various business and other ideas, and get general understanding in the areas I want to get more knowledge of (both IT and non-IT), which helps when I need to start digging deeper into the details.
I usually run quantized versions (I have older GPU with lower VRAM).
Wordsmithing and generating illustrations for my blog articles (I prefer Plex, it's actually fairly good with adding various captions to illustrations; the built-in image generator in ChatGPT was still horrible at it when I tried it a few months ago).
Some resume tweaking as part of manual workflow (so multiple iterations with checking back and forth between what LLM gave me and my version).
So it's mostly stuff for my own personal consumption, that I don't necessarily trust the cloud with.
If you have a SaaS idea with an LLM or other generative AI at its core, processing the requests locally is probably not the best choice. Unless you're prototyping, in which case it can help.
For me though, this would be all upside because I have largely explored technical topics with language models that would only be impressive to an employer.
At this point, it is like asking what does someone use a computer for? The use cases are so varied.
I can see how it would be interesting for myself to setup a local model just for the fun of setting it up. When it comes down to it for me though it is just so much easier to pay $20 a month for Sonnet that it isn't even close or really a decision point.
The multimillion (or billion) dollar collections of hardware assemble the datasets that we (people who run LLMs locally) run. The non-open datasets that we-host-LLMs-for-money companies do the same, and their data isn't all that much fancier. Open LLMs are catching up to closed ones, and the competition means everyone wins except for the victims of Nvidia's price gouging.
This is a bit like confusing the resources that go in to making a game with the resources needed to run and play that game. They're worlds different.
The same applies (to smaller extent) to Google, and more recently to Alibaba.
A lot of free (not open-source, mind you) LLMs are either from Meta or built by other smaller outfits by tweaking Meta's LLMs (there are exceptions, like Mistral).
One thing to keep in mind is that while training a model requires A LOT of data, manual intervention, hardware, power, and time, using these LLMs (inference, or forward pass in the neural network) is really not. Right now it typically requires NVidia GPU with lots of VRAM but I have no doubt in just a few months/maybe a year someone will find much easier way to do that without sacrificing the speed.
You're commenting on a post about how the author runs LLMs locally because they find them useful. Do you think they would run them and write an entire post on how they use it if they didn't find it useful? The author is seriously using them.
Can you expand a bit on how you might do this?