On Linux it's easy to benchmark this via e.g. `dd if=/dev/urandom of=/dev/null bs=1M`. On my somewhat old server (with Intel Core i5s from 2012), I'm getting 17 MB/s.
Supposing we blame slowness on syscall overhead: at a block size of 16 bytes instead of one megabyte, it slows down to all of 11.6 MB/s, that is, some 700,000 128-bit IDs per second. (And since we're not using SHA-1, all 128 bits are meaningful for collision resistance.) For comparison, Twitter gets about 1% that many tweets per second, globally.
If anyone has plans of scaling to 100 times the size of Twitter on a single 2012-era Core i5, and has figured everything else out already, I'll be happy to audit your userspace PRNG then.