EDIT: groped -> grouped
EDIT: groped -> grouped
E.g.: Suppose the data has high-order structure but is locally uniform (very common, comes about because of noise-inducing processes). Compute and store centroids. Those are more uniform than your underlying data, and since you don't have many it doesn't really matter anyway. Each vector is stored as a centroid index and a vector offset (SoA, not AoS). The indices are compressible with your favorite entropic integer scheme (if you don't need to preserve order you can do better), and the offsets are now approximately uniform by assumption, so you can use your favorite sphere strategy from the literature.
Here's some work on low-latency neural compression that you might find interesting: https://arxiv.org/abs/2107.03312
Another thing that I'm sure you explored, and I'd love to hear how it went, would be to rearrange the elements in the vectors such that perhaps the denser parts could be more contiguous, and the sparser parts could be more contiguous, on average. That sounds like something that would be easier to compress. Were the distributions such that a rearrangement like this might have been possible? Or were they very evenly distributed?
I.e. could you have rearranged a Gaussian-like distribution into a Poisson-like distribution?
I also remember trying to fit a distribution so that I can generate synthetic data (not for a lack of data, but more for understanding the problem space better). The synthetic data quantized pretty differently - my guess is that it's because of random areas of density and sparsity.
I'm not quite following your exact rearranging idea though. Not sure if the above answers the question.