The closest thing to public literature about Colossus and D that we've ever published is the Procella paper, which describes the abstractions that Colossus provides for it. In some ways it's similar to GFS (RPC interface, writes are generally append/overwrite) but many things are completely different now.
You need something that provides the block store abstraction on top of the primitives exposed by Colossus/D. Think of something like what modern SSD do in order to work efficiently with the underlying flash memory.
Then you have to hook that adapter in your virtualization stack (e.g. kvm) so you can boot from the volume and mount it from inside the VM. You could implement a kernel module or do it internally in kvm/qemu somehow, but iSCSI provides a straight-forward way to implement this in user-space: you have a process on your physical machine that speaks iSCSI upstream, and speaks Colossus/D RPC downstream.
(I don't know if they still do this but I have a vague memory of somebody describing the stack of an early version of GCP while I was working there long time ago)
This design has only one hop to the storage node. Low-latency workloads benefit from this design, high-bandwidth workloads sometimes actually benefit from off-loading PD to another host. To do iSCSI with one hop, you need to implement iSCSI interceptor, and basically you would have same design with less flexibility for guest OSes.
The irony, of course, that all this is a lot of legacy technologies needlessly wasting computer power: guest file-system trying to communicate with 4K blocks with “block device”, which goes through multiple layers of queues, then is re-maps to another abstraction, which goes over network to multiple hosts, etc. Not a single cloud customer ever said “we are so excited to manage volume sizes and bandwidth quotas for PD”. Better design would be to implement true data center-level filesystem to better support container workloads and leave PD for legacy cases, but Google’s storage management is so detached from reality, that it’s impossible to do cross-organizational project like this.
2) Then, for VMs you can do FS driver, jump to VMM and booms, you are done, multiple legacy levels of re-packing and redirection are gone. For shared access cases you do NFS/Samba interceptor and then same code path as above.
This system would be highly beneficial not only for Cloud, but for other Google properties as well: it would provide normal posix FS to Borg jobs. Amounts of equilibristics required to use any open-source package is enormous and by this time exceeded costs of developing FS multiple times, ask YouTube, MySQL, package management, etc groups.
Another example of Google’s storage craziness is cross-dc storage. This should be low-level Colossus responsibility. Instead godzillion of teams implement their own, GCS, PD, Spanner, Placer, etc. Crazy.
I will say though, that most of the time I don't want to use open source stuff. Observability is pretty crap, I don't want nor need software that uses write() without fsync(), and the assumptions that most OSS makes about FSs gives me nightmares on Borg.
Some stacks are mostly normal, some are really odd... and then there's storage.
Discord's description of the issue sounds like issues with either zonal or regional SAN storage.