Deep Dive: Nvidia Inference Research Chip Scales to 32 Chiplets
tomshardware.com
tomshardware.com
8-bit precision is good for neural network weights, but not for Bayesian inference, where from my understanding 32-bit floats might be the minimum. I wonder if anyone is working on hardware for accelerating probabilistic programming. Some approaches are harder to accelerate, such as monte carlo sampling (e.g. no u-turn sampler (NUTS)), while variational inference may be better suited to running on this type of custom hardware.
The problem with NN inference / training is that they are eating up the datacenters.
At the same time you can't achieve 20x speedup compared to GPUs if you need 32 bit floats, because in that case your relative energy utilization is not that bad.
Normal GPU's can be used for both learning and inference tasks. Dedicated inference processors have limited numerical accuracy and are more suitable only for inference, but have lower cost and are more power efficient.
But what would be the next step from there?
Could it be that memory chips as we know them get merged with processing and in the case of graphics cards, maybe a logical progression down the line. Instead of a row of ram chips, and a processor chiplet blob, you have those chiplets merged with the ram and form a mesh of both processing and memory.