I look forward to the inevitable probabilistic sub-bit machine learning models :)
I was able to get 0.68 "effective" bits.
The idea is that in each forward pass you add noise to each weight independently drawn from normal distribution, and when you calculate snr, it's sub 1 bit. Points to the idea that a stochastic memory element can be used.