How would you re-architect it if you knew ahead of time it would only sit on SSDs - never on spinning media?
How would you re-architect it if you knew ahead of time it would only sit on SSDs - never on spinning media?
SSDs are, unlike rotational drives, inherently parallel devices. Each drive has a number of flash memory units which are capable of concurrently doing writes. To split the load evenly we divide files up into zones which we call "extents" the number of extents is configurable and varying the number of concurrent extents will give varying performance for a given drive. Normally there's a sweet spot where you have as many extents as it has independent chips. Extents are written to in a very specific manner. They are append only which means that you start at the top of the extent and fill it up with blocks of a predetermined size. The block size can be tuned to optimize performance as well although I'm not sure off the top of my head which properties of the drive determine which block size is best. Once an extent has been filled up top to bottom we stop writing to it and move on to another extent.
Garbage collection: Each block that we write in an extent is actually a version of a logical block in a btree. When blocks are changes they are rewritten and the old block is marked as garbage. This leads to extents which are filled almost entirely with garbage. When we hit a threshold we look for the extent with the most garbage copy all of the non garbage blocks to an active extent and then reclaim the extent. SSDs have their own garbage collection mechanisms that can seriously degrade the performance of the drive if when they misbehave. If we get the extent size right the drive will have a much easier time of garbage collection (I actually think the drive won't have to do anything at all for gc in many cases.