Rsync.net Technical Notes – Q3 2021
rsync.net
rsync.net
Once again, thank you to Allan Jude at Klara Systems for the advice and guidance with the new ZFS "special" vdev for metadata caching that is discussed this quarter ...
Trusting a service provider is really hard in most cases, but you make it easy to trust rsync with posts like this and backing it up with reliable services.
I can't speak for everyone here but I know that many of us, especially me, consider rsync.net to be our life's work.
The solution is not to fix that page but to remove all links to it ... I'll get the scientists working on it immediately.
This makes sense to me (and is a good example of looking at more abstract failure domains in addition to the basic ones we all know and love) -- I'm curious if there's data to support this. rsync.net is in a good position to possibly collect that data.
I have heard smart people confirm that this is a smart and reasonable practice but have never seen any data or supporting figures, etc.
It's basically cost-free and if you don't like other vendors, you can always pair up (current Intel drive) with (one generation ago Intel drive).
As a side effect people generally now stagger SSDs a little to avoid something similar happening (ofc if you have multi machine replication this is less of a issue, but still a total machine loss can hurt due to capacity loss or parts shortages in edge locations, etc)
I've personally not seen a synchronised SSD array happen failure for a long time, but it's hard to know how much of that is because people now plan to avoid them.
[edit]
with the exception of: https://www.engadget.com/2020-03-25-hpe-ssd-bricked-firmware...
I want to clarify - there's the issue of a bad batch wherein their longevity is greatly reduced and they fail in a cluster, etc., etc.
But that is not what we are guarding against ...
Instead, the risk we're thinking about is that there is an actual bug in the firmware that causes a particular workload to brick the drive or destroy it or whatever.
The critical point is that if the drives are mirrored then they experience an identical workload over their lifespan and they could fail literally simultaneously.
So by all means - do indeed guard against bad batches or manufacturing defects by mixing drives. Just understand we're talking about something slightly different here ...
The real magic bullet was changing the "freezing" process where we transform 'borg' the python script into 'borg' the binary executable.
You'll see a full writeup of this in the Q4 technical notes :)
I have a ZFS account - how do those work under the hood? Is it a VM backed by a ZFS volume? How does the overhead compare to a normal account? I suppose using a VM eliminates some advantages of the special device.
I will have to look into this - does your zpool benefit from the metadata cache ? Perhaps not since it is a different zpool ...
Do you think you'll ever get to a point where you run zpools entirely of SSDs? If so, what criteria is important for you? (Raw price per gigabyte? Power usage? MTBF? Something else?)
All of our access is over WAN so disk IO is not that important - it is raw price per GB and even with the "nice" SAS drives we buy ... ~$400 for 16TB is a huge difference vs. ~$700 for 1.8 TB (which is, roughly, the price of the Intel part mentioned in the list of cache drives).
That is, until you fill it up - which you could.
At that point new metadata goes onto the spinning disk vdevs, as it did prior to the cache. As files age in and out, space gets freed on the cache and some new metadata makes it back on there.
So if you fill it up, it behaves more like a cache.
ALSO, you can add multiple metadata caches to a pool ... so if you fill one up, you can add another ...
Too many years of thinking of the aux vdev types like SLOG and L2ARC as caches, has made people confuse this one frequently.
- Chia is a cryptocurrency that made drives ~double in price for a while, but that's mostly stopped, though they're still a little more expensive than they were before. Also, reading between the lines, Chia looks really, really scammy/pyramid-schemey, even by cryptocurrency standards.
- A 3rd party SQLite streaming replication project added SCP as a transport mechanism, so now it works with rsync.net
- Rsync.net supports lots of their-side checksumming tools and they're really like it if you'd stop using computationally expensive ones if all you're doing is verifying data integrity.