sorry for the extremely dumb question but is it possible to run the 68B model in a 8gb ram computer?
Having said that: if you are feeling incredibly patient you can technically run the 68B parameter model by swapping to disk, although it will not be a pleasant experience (think minutes or hours per token instead of tokens per second)
Additionally worth noting pure CPU inference is much slower than GPU/TPU inference, so the output will be much slower than a ChatGPT-like service even if it does fit in your computer's RAM
With GPU:
VRAM + RAM >= 68/2
Without GPU:
RAM >= 68/2
Technically you can by swapping to disk but it would be too slow to be usable.
What you can do however is use the 7B model with 4bit quantization and use it within 8GB RAM.