Strange. I remember trying to get this to work on a 16gb machine and all of the comments on a github issue mentioning it were saying it needs at least 32 or more.
/edit this was with llama cpp though not ollama
/edit this was with llama cpp though not ollama
Though if your machine can’t keep it all in memory, then speed will still fall off a cliff.
Thanks!