There's a float16 version of GPT-J rewritten in C++ here that apparently fits into 16GB...
https://github.com/ggerganov/ggml/tree/master/examples/gpt-j
it's CPU-only and runs inference at "about ~6 words per second" on an M1 MBP
https://github.com/ggerganov/ggml/tree/master/examples/gpt-j
it's CPU-only and runs inference at "about ~6 words per second" on an M1 MBP
These models are obviously nerfed, but you can run them on budget ARM hardware and get legible, paragraph-length responses in less than 10 seconds.