I was able to run 7B on a CPU, inferring several words per second: https://github.com/markasoftware/llama-cpu
But on my machine, it automatically used all 12 available physical cores. Setting OMP_NUM_THREADS=2 for example lets me decrease the number of cores being used, but increasing it to try and use all 24 logical threads has no effect. YMMV.