Noob question, and may be probably being asked at the wrong place.
Is there any way to find out min system requirements for running ollama run commands with different models.
Install Ollama from https://ollama.ai and experiment with it using the command line interface. I mostly use Ollama’s local API from Common Lisp or Racket - so simple to do.
EDIT: if you only have 8G RAM, try some of the 3B models. I suggest using at least 4 bit quantization.
You can easily experiment with smaller models, for example, Mistral 7B or Phi-2 on M1/M2/M3 processors. With more memory, you can run larger models, and better memory bandwidth (M2 Ultra vs. M2 base model) means improved performance (tokens/second).
I have not ran into a llama that won't run, but if it doesn't fit into my GPU you have to count seconds per token instead of tokens per second