https://huggingface.co/unsloth/GLM-4.7-GGUF
This user has also done a bunch of good quants:
This user has also done a bunch of good quants:
https://unsloth.ai/docs/basics/unsloth-dynamic-2.0-ggufs
It isn't much until you get down to very small quants.
The flash model in this thread is more than 10x smaller (30B).
https://huggingface.co/models?other=base_model:quantized:zai...
Probably as:
issue to follow: https://github.com/ggml-org/llama.cpp/issues/18931