Lossless model compression experiment: GLM-5.2 in 25% less memory
brianbell-x.github.io
brianbell-x.github.io
Also, are all that many people handling the BF16 weights directly? GLM-5.2's reference deployment is FP8, and many vendors are even serving at NVFP4 which seems to offer negligible degradation over the FP8 reference deployments.