Ha, ha. Yeah… no.
> there's no mention of its power consumption
Can't be tremendously higher than Pascal for reasons of physics.
> sure, you can get a million TFlops without cache, now try to get that chip to do anything useful
"Without cache" is certainly an exaggeration. It won't have a globally coherent cache hierarchy in the style of CPUs. It certainly will have various on chip memories to hold intermediate results. Neural net workloads are incredibly predictable and homogeneous and are essentially the perfect scenario for hand optimization of data flows to beat automatic caching.
> no CUDA support
You're just being silly now. CUDA isn't a standard, it's proprietary to NVIDIA and this isn't a general purpose processor anyway.