Here's my issue with that. The on-chip SRAM is only 32MB, and the RAM is LPDDR4 rated at only 68GB/s.
Assuming a dot-product (multiply + add) over INT8 data, that's a limitation of 2-operations per 68GB/s (that the RAM moves at). Or 136 GIOPS (Giga-integer8 operations per second). You're limited by RAM, based on what I've seen in the presentation.
Unless their neural net is 32MB and fits entirely in on-chip SRAM. That seems unlikely to me...