> GPT4All: When you run locally, RAGstack will download and deploy Nomic AI's gpt4all model, which runs on consumer CPUs.
> Falcon-7b: On the cloud, RAGstack deploys Technology Innovation Institute's falcon-7b model onto a GPU-enabled GKE cluster.
> LLama 2: On the cloud, RAGstack can also deploy the 7B paramter version of Meta's Llama 2 model onto a GPU-enabled GKE cluster.
Why not llama2 on dedicated/local hardware? Memory and download size requirements?
Ed: After reading the linked tutorial - it looks like the built docker container will run fine on local/dedicated hardware?
https://www.psychic.dev/post/how-to-deploy-llama-2-to-google...