How are people running this locally? I just checked llama.cpp and it appears unsloth has a version but it hacks a bunch of things to make it work and isn't optimal.
From there I collected the following US providers currently serving GLM 5.2:
- Together (https://www.together.ai/models)
- Fireworks (https://fireworks.ai/models)
- Featherless (https://featherless.ai/models)