There's actually an old paper titled Optimal Brain Damage, where they don't try to find optimal quantizations, but optimal sparse versions of a models-- i.e. where some weights are set to zero.
Optimal Brain Damage https://www.researchgate.net/publication/221618539_Optimal_B...
Optimal Brain Compression https://openreview.net/pdf?id=ksVGCOlOEba#:~:text=The%20resu....
TinyVolt’s implementation of it: https://github.com/TinyVolt/optimal-brain-compression
I don’t think many people have implemented such things. You might discover something new experimenting with them.