Borgbase backups have been unavailable for 3 days
status.borgbase.com
status.borgbase.com
Despite the warning banner the uptime for customers on us10 instance which has been unavailable for almost a week now is showing at 100%.
They are compensating affected customers by crediting 4 weeks subscription but I must say this does make me wonder about their redundancy and recovery architecture.
Opshugs to the folks at Borgbase.
--
> box-us10 offline for storage expansion
> We are currently experiencing a temporary outage on our box-us10 server due to an unplanned expansion process. We added two new hard drives to the server in an effort to enhance capacity and accommodate future growth. However, we did not anticipate that the expansion process would require the server storage to be temporarily offline.
> To prioritize the safety and integrity of your existing data, we have decided to keep the box-us10 server offline until the expansion is successfully completed. It's estimated that about a week of downtime will be needed for the expansion. We understand the inconvenience this may cause and sincerely apologize for any disruption to your services.
> Current progress: 26% (Sunday morning UTC) > Date Created: 2023-08-11 16:17:36 (3 days ago) > Last Updated: 2023-08-13 18:15:53 (15 hours ago)
Sounds like an Elon Musk euphemism after a bad day at SpaceX.
But I don't know the details of what really happened, so I can't make a judgement really. If I could edit my previous comment I would, to be less emphatic that someone did something "wrong".
It could have been that this particular disk array had different firmware or was faulty. All we know is that something unexpected happened.
Yes, the expected behaviour would be for the disks to be initialized. Beyond that, it'd be a configuration setting for what to do with new disks. It could be used to extend the existing LUNs, or added as a new pair in a RAID 10 style setup, or added as new members of a different level RAID.
Then it would be up to the sysadmin to extend the LUN, or divide the newly added space among multiple LUNs.
Different arrays have different behaviour with new disk. Some only add to offline LUNs. Some do everything online with lower performance, as you said. Even some super expensive enterprise kit will be very easy to break by adding new disk.
You can pretty much guarantee it's untested and unlikely to work, when someone doesn't understand how drive arrays work, or test expanding drive arrays before doing it in prod...
edit: Which tells you something about their backups, and it ain't great. Or they'd restore onto another box.
If we are talking about privacy (or rather privacy theatre) then yeah Apple, why not.
They could have priced it at $299 to include the initial 12 months iCloud subscription.
If you only go via the iCloud route, you have zero local copy of your data. If you only only go via local, it isn't safe enough option. So ending up with a combination of the two.
And Apple really wanted services revenue, if it was only just a Time Capsule I doubt Apple would ever make it again.
I can imagine something similar happening here, however there should be a fail over storage setup in case they're selling a service.
One, it is obviously some flavor of raid rebalancing, there’s not a different kind of thing it could be.
Two, they’re admitting they didn’t expect this, that’s not possible if it’s a tested procedure.
ps -- it's not about being bashing these folks. But these type of multiple operational incompetencies strongly call into question their ability to safeguard your bytes and produce them when you need them. It's kinda the whole point of the thing.
That does include their own of course.
"did not anticipate" can be interpreted in different ways too. It may be "didn't know it could happen", or just "estimated it to be so unlikely that emergency downtime was acceptable as a response".
I've seen outside speculations as an inside person before and they're often very likely to be BS. Let's stick to facts and skip the "obviously" and "not possible" until we learn more. What's happening right now may be their prepared and tested procedure. (Or it may not)
(1) Add more drives, start online rebuild
(2) Have one drive fail/URE mid-rebuild
(3) Be forced to switch to offline rebuild because the array was in a strange not-entirely-consistent state even before the failure, even though without the new drives you’d be in a routine online-rebuild situation
If so, possibility (2) could in principle not be detectable in staging, as it could depend on the drive’s age, and having to reread everything—as a rebuild does—is a rather abnormal load. The failure could even predate (1), if it happened to data nobody’d looked at for a long time.
This is less a question about these guys and more in general—is a RAID/ZFS/etc. array/pool/etc. in a vulnerable state after you’ve expanded it?
Far from being an abnormal load, it should be a cron job. Backups need to be verified. Data at rest needs to be periodically checked for bitrot. Reading everything on the array and verifying parity and any checksums is important preventive maintenance and an early warning system.
In situations where you could plausibly get more data, it’s usually OK to admonish people to just go and get some, because whatever the difficulty is doing that, judging the current degree of uncertaintly is probably even harder. In situations where you couldn’t, though (trade or state secrets), there’s nothing to do but to use what data you do have and point out your uncertainty carefully.
(This is a very general argument because what you’ve said is a fully general counterargument to any criticism of anyone from anybody but an insider. As for this particular case, GP’s allegation of lacking ops knowledge seems to be on the charitable side to me: an uncharitable one, as stated elsethread, would be that they had one drive too many fail on them, whether due to incompetence or misfortune, and are forced to do an offline rebuild while choosing to lie about it. It’s the lie part that I’d have the greatest problem with, if it’s actually there.)
> However, we did not anticipate that the expansion process would require the server storage to be temporarily offline.
This is embarrassing. Somebody made a big mistake of the “mistakes like this shouldn’t be made” category. Like not an accident or fat finger or a bug, but engineers not understanding fundamentals of how a system worked and having none of the processes in place to to dry runs in non production or any of the other things that would protect against something like this. Really questionable to trust an organization that makes a mistake like this.
You have no idea what happened.
Protecting against "something like this" often introduces complexity and failure cascades that are much more harmful than a simple system being offline.
In fact, I would consider it a feature if uptime was deliberately sacrificed for simplicity and data integrity.
I wish Manu the very best and will consider it not a failure but a success if this array/subsystem/whatever comes through without data loss.
Everyone makes mistakes and I appreciate they were very upfront about it, but yea. Questioning if I'll be an ongoing customer now.
rsync.net does not support append-only mode, despite advertising it: https://news.ycombinator.com/item?id=32756653
Unavailability is occasionally expected. Don't rely on a single provider to be your sole backup.
Any one have any experience with Hetzner Storage Boxes for this?
I’d like to prepay for a year or longer for critical services so that in case something goes wrong (card expired, didn’t notice reminder emails, temporarily off the grid, temporarily incapacitated), things just don’t go poof.
If you want “perfect architecture” and reliability/availability you just need to pay for it (aws/s3). Or make use of multiple providers.
I find that really funny. AWS has outages regularly. Whole data centers who go dark. Remember the S3 outage because an admin fat-fingerly deleted a large number of servers? ...no, people forget very quickly because of the "No one was ever fired for using AWS" mantra.
The usage graph on each repo is great. I redid my server this year and was accidentally backing up a stale snapshot every night for one of my repos. It would succeed, but without changes. The flat line for usage tipped me off. I’m not sure I would have noticed otherwise since it was a VM image that I could restore and boot, but it wasn’t super obvious it was full of stale data.
This downtime isn’t ideal, but, if downtime is needed to preserve data integrity, I’ll take it over the risk of trying to maintain availability.
Everything is automated and runs on a schedule, but when I need some ad-hoc restore, say, of a single snapshot or folder I can always use KopiaUI.
Been quite happy with it.
[0] - actually two providers, for redundancy, plus local physical storage
Though just this weekend there was some screwiness with sub accounts and ssh keys not working that did resolve itself after 24-48 hours. I didn’t notice anything communicated about it.
This is correct - we believed we had a working recipe that properly sandboxed the 'rclone mount' and 'rclone serve' directives but we couldn't make it work in a way that allowed us to sleep at night.
But now that is all changing...
It appears that we can run rclone serve restic --stdio in a way that doesn't actually serve anything or create sockets, etc., and as soon as we finish our tests we will have it in place and people can lock their accounts with something like this:
restrict,command="rclone serve restic --stdio --append-only backups/my-restic-repo" ssh-rsa ...
... in authorized_keys and achieve proper append-only mode at long last.
Stand by ...
Is there a particular channel where you intend to post your findings? :)
[1]: https://borgbackup.readthedocs.io/en/stable/usage/notes.html...
However I agree that it doesn't look good for a storage company.
I’d like further transparency on this: why didn’t they anticipate that the array would be offline for this expansion? Expanding a storage environment can certainly be a mine-filled exercise for even the most experienced. But why was this expansion unexpectedly an offline operation?
I’m not throwing shade here. I’m genuinely curious because I’ve been considering Borg and Borgbase for an upcoming project.
With enough 22TB spinning drives nowadays you can get in to a scenario where, with the new data being written into the array and the expansion process going on at the same time the rebuild essentially will never complete. This is especially true with lower end CPU servers and without dedicated RAID cards.
It is dangerous because the rebuild stresses the drive which makes a failure more likely, and a failure during the rebuild process is not great, on top of that you've got a months long ETA...
No fun
Can you tell us what solution that was? Something with ZFS or BTRFS? My experience with classic RAID system is, you can't expand them without reinitializing them. (But that comes with obvious warnings about data loss.)
Several cheap, low-power NAS boxes running Linux/Ceph and throw disk at them.
Depending on the data you store you can have 3-way replication of Erasure Coding at your preferred risk/performance level.
Disk failures are painless, box failures don't lose data, expansion and re-balancing scales with the number of disks and pretty quickly the limiting factor is your network (4-5 modern spinning disks can easily saturate a 10G network link).
My secondary backup is borg as well, this one hosted on a Hetzner Storage Box[2]. Having seen some bad practice in this thread, I want to make it clear that I am not duplicating the rsync instance to Hetzner but have created a separate instance entirely to avoid issues with the primary borg instance being replicated. From what I understand this is best practice (although it does mean I have to run the backup process twice--once to rsync.net and then again to Hetzner). Hetzner also provides info[3] on using their service with borg. I pay around 3.50 EUR a month for this service.
I chose rsync.net due to their CEO being active on HN and also because they're based in North America (away from me in the Pacific). Hetzner , I chose because they are based in Europe giving me both operational diversity (operated by two different unrelated companies) and locational diversity (Western US for rsync.net and Northern Europe for Hetzner. Hetzner sharp pricing also helped as well.
I'm sure I looked at BorgBase at one point for my EU-based backup location but something must have put me in Hetzner's direction instead.
I also store photos (not my general backup, for now) in AWS S3 (largely because that makes it easier to deploy my static photo album site which it itself deployed via S3+CloudFront). I use S3 standard storage (for the time being) for a bucket in Sydney (nearest to my location) and then a glacier storage in Ireland for a bucket which is automatically replicated from my Sydney bucket.
Hopefully all this is enough redundancy! But all up I only spend ~ USD20 or so a month for this peace of mind.
[1]: https://www.rsync.net/products/borg.html [2]: https://www.hetzner.com/storage/storage-box [3]: https://community.hetzner.com/tutorials/install-and-configur...
People who can give solid explanations are probably busy right now, and those who aren't can only write what you already seen.
You mean to downplay what actually happened?
First I have a local copy of my data living on my actual hard drive.
Next I use PikaBackup (via borg) to encrypt and sync that to a cloud server I run that has about 200GB of storage for these backups for about $1/mo added cost.
Next I use the Backblaze Cloud tool to synchronize those encrypted backups to B2 using `b2 sync --delete` command. It runs automatically via cron every night. The costs here are about $0.01 per month since I only get charged what I actually use.
Backups are pruned and I can easily control the schedule, copies, etc. If I need to recover a file it mounts as a file system that I can easily navigate using any tools I want, including the command line. I can also mount older or different snapshots.
I used to have back up on local external drives too, but stopped doing that, since the process was manual and I often forgot about it.
Edit: previous comment on same (not a shill, I swear!): https://news.ycombinator.com/item?id=34152369
I get that things happen, but not realising that this was going to be a service impacting event does not inspire continued confidence. Sure, this situation probably won't happen again, but what else don't they understand about their infrastructure?
You automatically test restore? That makes sense but I've never heard of that before, can you describe the process?
1. Make a e.g. 30MB file of random data
2. Copy it to "_reference" file
3. Upload the file to backup service
4. Restore the file from backup service
5. Diff restored file against referencePick a couple random files that should be in the repo, restore them from a random archive, check the md5sums against the source. If the md5sums don't match (or the file can't be found), something is wrong. I am mainly backing up RAW image files, so they should never change.
Basically...
$TEST_FILE=$(ls -p /source_dir | grep -v / | shuf -n1)
$TEST_ARCHIVE=$(borgmatic -c config.file list | shuf -n1)
borgmatic extract yada yada yada
md5sum $TEST_FILE restored_file
You are (maybe) protected against a few disk failures but that's about it.
This FAQ entry seems to confirm this: https://docs.borgbase.com/faq/#which-storage-backend-are-you...
Recently I made the switch to Kopia [2] which seems to have feature parity with Borg (and Restic [3]). It also has a web UI which is way easier to work with than Vorta. And I can easily view, extract and restore individual files or folders from there. This gave me way more confidence about this solution. The only thing I really miss is that I cannot chose different targets for different paths. For instance, with Borg I was able to backup a partial of my Docker appdata to an external source. And I haven't found a way to do this with Kopia. Besides that I'm pretty happy with this solution and I would recommend it.
I don't envy their sysadmin team right now. Good luck, folks.