It's kinda similar to Parallax mapping in a way.
I'd also be interested in what a "voxel-based" GPU would look like, they already support sampling 3d volumes in 3d textures, and the issues they have (namely memory use) seem pretty fundamental and I'm not sure how they can be overcome to match the scale of 3d environments expected today.
Clearly (?) the way the math works out, the polygon-oriented GPUs we have make more sense for 3D graphics. But one can imagine a universe where things would be different.
It's why for display pretty much every modern voxel engine processes the sparse voxel world into a dense surface list - which then maps (reasonably) well to triangles.
I'd argue this difficulty in parallelization is a bigger hurdle than "just" engineering effort.
Much of the way a GPU functions isn't because we decided we loved triangles, but we found things that parallelize well and worked backwards. And then on modern GPUs don't actually devote that much area to dealing with "triangles" - so little is gained that it's often not worth redesigning to remove that capability from data centre "GPUs" that will never see a single triangle in their lives.
Signed Voxel Octrees, and SVO Directed Acyclic Graphs in particular solve this pretty well. Most megavoxel engines are an implementation of this paper: https://graphics.tudelft.nl/Publications-new/2020/CBE20/Modi...
Plenty of examples on https://www.shadertoy.com/
The principle that image generation methods are dictated by available bandwidth goes all the way back to the first graphical computers (for instance the weird color-attribute encoding in the ZX Spectrum is essentially an image compression method to get more graphical fidelity out of a very limited memory bandwidth).
The term VPU is already taken though: https://en.wikipedia.org/wiki/Vision_processing_unit