Their compiler is closed source, so it's almost impossible to tell what's going on. But, when I connect using tf.Session(), then .list_devices() shows only CPU available.
However, when I enable the TF_MLC_LOGGING=1 magic variable, it does seem to be printing out messages that indicates it's doing some kind of graph substitution under the hood. Therefore, I assume that this is the intended usage mode.
In other words, there seems to be zero difference between the "CPU" and the "GPU". Normally you can say "Do this on the CPU" while "do that on the GPU." But not with this.
Hopefully they'll open source the code sometime this century so that it's clearer what the heck it's doing. For now, though, it's reasonably fast in whatever this "CPU" mode is -- I only need to run unit tests on my laptop anyway, since all training happens on TPUs. So I ended up happy.
(For the first day or so, I was panicking that I was going to have no working tensorflow whatsoever on my M1 laptop, which would've necessitated a swift return + substitution.)