How to run an LLM on your PC, not in the cloud, in less than 10 minutes
theregister.com
theregister.com
And yes, but there's no lower limit on the memory; it's entirely dependent on the model or kernel size.
Even my M1/16GB gets decent speeds. 7+ tokens/second with llama3
better still you can the use python to call it with langchain chatollama and build anything you want with a little help from claude and chatgpt or codeqwen if you want to do it all locally.
absolute AI don. impress the ladies with that one, you'll be beating them off with a stick when they see it in action.
just need plenty of VRAM after that.