I'm nothing even close to a knowledgeable expert on this, but I've dabbled enough to have an idea how to answer this.
The special thing about why GPU's are used to train models is that their architecture is directly applicable to parallelize the kind of vector math that goes on under the hood when adjusting weights between nodes in all the layers of a neural network.
The weights files (GGUF, etc) are just data describing how the networks need to be built up to be functional. Think compressed ZIP file versus uncompressed text document.
You can run a lot of models on just a cpu, buts its gonna be slooooooow. For example, I've been tweaking on running a Mixtral8x7b model on my 2019 Intel Macbook Pro with llama.cpp. It works, but it may be running at 1-2 tokens per second at max, and that's even with the limited GPU offloading I figured out how to do.