That understanding of the system is correct. To make it practical we've implemented a bunch of optimizations to minimize I/O cost. You can see how it performs on inference with BERT here: https://youtu.be/qsOBFQZtsFM?t=69.
The overheads are larger for training compared to inference, and we are implementing more optimizations to approach native performance.