You can run a 7b on most modern hardware.How fast will vary.
To run 30-70b models you're getting in the realm of needing 24gb or more of vRAM.
You can run a 7b on most modern hardware.How fast will vary.
To run 30-70b models you're getting in the realm of needing 24gb or more of vRAM.
I can't agree with "very bad". Maybe your standards are set by the best, largest models, but have a little perspective: a modern 7b model is a friggin magical piece of software. Fully in the realm of sci-fi until basically last Tuesday. It can reliably summarize documents, bash a 30 minute rambling voice note into a terse proposal, and give you social counseling at least on par with r/Relationship_Advice. It might not always get facts exactly right but it is smart in a way that computers have never been before. And for all this capability, you can get it running on a computer a decade old, maybe even a Raspberry Pi or a smartphone.
To answer the parent: Download a "gguf" file (blob of weights) of a popular model like Mistral from HugginFace. Git pull and compile llama.cpp. Run ./main -m path/to/gguf -p "prompt"
https://www.reddit.com/r/LocalLLaMA/comments/1cj4det/llama_3...
I haven't used any agent frameworks other than messing around with langchain a bit so I can't speak to how that would effect things.