I was taken back when I saw what was basically zero recall loss in the real world task of finding related topics, by doing the same thing you described where we over capture with binary embeddings, and only use the full (or half) precision on the subset.
Making the storage cost of the index 32 times smaller is the difference of being able to offer this at scale without worrying too much about the overhead.