Gandi loses data, customers told to use their own backups
status.gandi.net
status.gandi.net
I stopped using them a while ago but for a different reason. I used to use their website to check availability/whois for domains that I was interested in buying. If it was available I didn't buy it at the time but until I finished the website/app whatever I was going to put there, this took me a few months obviously. It happened to me that when I was finally ready the domain had already been sold to someone else. This repeated five times during six or so years. Now, I know, "someone else could have thought the same thing" but I find it very hard to believe that it happens so often. These domains were a bit of niche words that were not hot topics at the time, some of them using fairly uncommon TLDs (like .one). Another weird thing is that they were always registered to someone living/or doing business at India, and it was a fairly simple landing page with a "contact me" link. I'm a bit superstitious so I don't think it was a coincidence.
Now, I don't think this is a GANDI problem per se, but my theory is that they share this information (who is looking for which domains) with marketers or something like that, or maybe it was a rogue employee trying to make some money squatting domains. I would have expected this from BigDaddy or similar sharks, but from a company whose motto is "no bullshit" I had much better hope. Anyway, I decided to move (to namecheap if you're wondering) and surprisingly the problem went away.
Edit: Sad to hear of the data loss and for anyone affected. Trusting cloud providers doesn't always work out either.
For example, asdasdahbdajsdbajdbhsbdahsdd.com... not worth the $10
ireallylikechicken.com... maybe worth the $10?
(ireallylikechicken.com is available, go squat it and get rich)
> Creation Date: 2020-01-09T19:37:01Z
Registered with Gandi, ironically enough...
And the domain is meta! Response from HTTP GET
> Location: https://news.ycombinator.com/item?id=22001822
I'll give you 20 for it?
Bitcoin price in USD on January 8 2019: $4004.12
Bitcoin price in USD on January 8 2020: $8045.51
I know it was very shortly stopped once people complained, but it goes to show that it has been done before.
Getting rich off domain's - sounds like a solid business plan!
"Does GoDaddy register domains you search?" (2011) https://news.ycombinator.com/item?id=2326790
Within this, was the comment that reminded me which company it was I was talking about (https://news.ycombinator.com/item?id=2327152) - Network Solutions used to do this - and there is a Wiki link in there which gives the process a name! TIL! :)
They buy domain X and then they put generic advertising, maybe keyword based, on a cheap bulk hosted site. They measure for a few days - is this bringing in lots of revenue? If not, they cancel the purchase, using a "grace period" available to users of the registry in case of mistakes - the purchase is unwound and they are refunded the fees. Domain X is now available again.
In principle this is forbidden for major TLDs but it's still possible and unscrupulous vendors help them do it, albeit it may now attract a fee if you do it enough that the TLD registry detects you tasting.
https://en.wikipedia.org/wiki/Domain_tasting explains about this and related practices.
What if you don't want to cough up $10-$20 on a whim? Would doing a whois (using the NIC's whois site) suffice?
Then again, the squatter would have to know what to search. Isn't it against rules for domain registrars to publish their recent query history [private or public]?
Any more light on this subject would be greatly appreciated!
Never had even 1 of my "saved for later" being squatted, for several years now, so I really have to trust and commend namecheap in that regard.
But as @Jasper_ said, this could be a problem with the domain name registry selling/leaking that info (AKA all their 'is_available' queries), and not the registrar.
It's a single data point, but I instructed a client to search for domains on Namecheap last year since they were undecided. I just didn't want them to use GoDaddy, and I warned them why. They settled on a domain but registered it months afterwards. It was still available.
EDIT: And, to make things worse, each time I was threatened with the "confiscation" of my domain, and the round trip on the tickets was so high that each instance took 2-3 days to resolve. Frustrating as hell.
Disclaimer: I don't work there or have any relationship, except I'm a happy customer
Seems you missed my point though. Both of our anecdotes doesn't really say anything, in terms of if Gandi is good or bad.
Actions speak louder than words. Google famously has a "we don't have bugs, you just don't know how to use it, talk to the hand" policy for example. It is better to learn about the policies due to minor issues rather than major. OP learned of it early on and moved away with little trouble. Others did not learn until now and stayed, and now they are SOL.
This is at the time of the stories of other registrars giving customer second and third chances to guess their PIN, or credit card or whatever mechanism they had, and resulting in domain hijacking.
This is annoying for everyone but the adversary who can just spend $50 to buy a set of fake ids with your info.
Especially since Gandi doesn’t store your old IDs, they aren’t even going to check if the info on the fakes matches the ones you provided previously.
> phone number registered in my name
I can’t imagine this working very well, just give them a number from a country where they can’t verify who owns the number.
Can you elaborate what the "verification" entails? There is an ICANN requirement[1] to validate whois information, although I've only been asked to validate email (at another registar, not ghandi).
[1] https://www.icann.org/resources/pages/approved-with-specs-20...
They're still my go-to provider.
whois mysupergreatcoolappidea.com
into a terminal window and see if you get back a resultIf proven it would be a major blow to their business, so why would they try to snatch pennies from in front of a steam roller?
So I call b.s. on any reports of “the registrar noticed me searching for a domain and registered it”.
It absolutely has happened and quite possibly still does.
https://twitter.com/andreaganduglia/status/12151991477012316...
While I appreciate that there are real people behind these companies that are probably having a really rough time right now, the criticism that Gandi are getting as a company is justified - and if Gandi are truly a "no bullshit" company they need to put something out to their customers asap.
Using memes after permanently losing customer data is extremely disrespectful.
They start thinking they are beyond the normal people and that everything is a joke.
I'm going to look at moving my domain registrations away from them.
Additionally, data recovery is a lot of waiting in most cases, there isn't much to do as your business burns down around you
This post was disrespectful. It's not an excuse, but this is a stressful situation and the thread was getting heated. Either way, I truly regret posting it and it was my decision alone to do so. Please don't take this as representative of the high standard Gandi sets"
"That said, for the sake of transparency, we won't be deleting the tweet -- Julie"
Whatever the context / stakes (doesn't change anything in this case), this is how people should behave in life (not just online).
A company that loses customer data in production is the exact type I would expect to mock their customers using memes.
IMO, the blame lies solely with the CEO, because he is still to retract his statement regarding snapshots not being backups (despite their site selling them as backups to the end-user), and for not accepting the fact that for someone controlling business data that creating backups AND regularly testing them via restores is 100% essential. Culture trickles down, and if the CEO only accepts blame and not the reason for the blame then it's a sign that they won't learn from the problem - and that's the biggest red flag you will ever see in ANY business.
I can only see one way back for them that won't taint their reputation completely. They need to:
* Post a full post-mortem of what happened, how it happened, how they fixed it, and what they're going to do to ensure it never happens again.
* Issue a full apology for the problem. Accept full blame, and accept (including the CEO on Twitter) that Gandi failed to follow accepted industry standards.
* Sit down with the engineers that work at Gandi and hear their grievances. While I doubt that their engineers knew this would happen, I'd be willing to bet that there is at least one person there that had raised the lack of off-site backups and no recovery mechanism. That person needs a promotion, and whatever resources needed to fix Gandi.
* Issue a full refund to those that lost data - not a small discount, as already reported. A discount is a kick in the teeth, whereas a full refund is the start of a real apology for failing the customer. If you go for a meal at a restaurant and find broken glass in your food, the first thing the server will do is give you a full refund, no questions asked, regardless of how expensive your parties order was. Gandi need to take the hit, and live to fight another day.
"Andrea, sorry about that and the incident. If we led you to believe that you had nothing to do on your side when warned multiple times to make your back ups, then we'll have to make it clearer, and stop assuming that it's an industry wide knowledge."
This is one of the worst responses I’ve ever seen from a company, and I’m not being hyperbolic.
The idea that someone would entrust their sole copy(s) of critical business data to a service provider is insane to me. Always keep your own backups.
2. The website says "Snapshots allow you to create a backup copy."
3. He says "No they do not allow snapshots download."
He also states he has his backups, so he's mostly just whining because he's annoyed he has to reupload stuff. Which I get, but again, he should understand what snapshots are and aren't.
Here the right one which state that they are backup: https://docs.gandi.net/en/cloud/volume_management/volume_sna...
Here's the one that you quote (which isn't the same service): https://docs.gandi.net/en/simple_hosting/common_operations/s...
Be careful next time judging with that little knowledge of the issue.
Here's the page as of earlier today. https://web.archive.org/web/20200109194005/https://docs.gand...
> He also states he has his backups, so he's mostly just whining because he's annoyed he has to reupload stuff.
Or they're annoyed that they paid for a service, at the very least billed as backup, only to be told "welp, it's gone".
You can consider it insane, they still sold snapshot as being backup. Insane or not, it doesn't change that's what they sold wrongfully.
Can you point me where that screenshot show what you say it does? The user goes further to specify that you CAN'T download theses snapshots.
Companies should be called out when they lie about what they sell, I hope you understands why it's important.
Especially for random ccTLDs, they're often significantly more expensive than the alternatives.
Random selections for domains: .ru is $1-3 most anywhere else, Gandi is $18.
Gandi might be better than some of the other low touch, self service domain providers but its definitely still in the same ballpark. $18/year still means they're losing money if they ever need to pick up the phone for you. It's not a price point that works with "higher end".
Being a registrar is only a side effect of their business though, not really comparable.
Feels like the CEO has made his money, forgotten the company's roots in the process and is happy for Gandi to be just another generic, overpriced registrar running on auto-pilot.
If they lost all the data, then obviously the only option for customers is to either use their own backups if they have them or accept that the data is permanently lost.
One can criticize their lack of additional redundancy, but don't see what's wrong with the response.
The industry standard is sucking up to them and groveling, and it's led to customers being very badly behaved.
The trouble is no one has a good working alternative to the industry standard.
Gandi certainly doesn't, they're not responding in a well thought out manner, they're losing their cool and getting angry with their customers. That's a quick way to go out of business.
Here I'd just avoid engaging one-on-one at all, just broadcast the situation status.
answering with memes is the absolute opposite of this, specially when your customer has all the reasons to be angry.
I did find HubSpot[1]:
> We may limit or deny your access to support if we determine, in our reasonable discretion, that you are acting, or have acted, in a way that results or has resulted in misuse of support or abuse of HubSpot representatives.
I'm still skeptical because actually enforcing that clause seems like it could lead to an expensive lawsuit. The angriest customers are naturally the most litigious ones, too.
The tone any company hosting customer data should take in the event of data loss is along the lines of 'regretfully... we screwed up... unfortunately... steps we are taking to ensure this doesn't happen again...' i.e. the company should either be humble and apologetic or they should expect to lose a large chunk of their customers after something like this. This isn't merely to say the right thing, it is to demonstrate that they acknowledge this was their issue and something they need to fix going forward rather than a 'sucks to be you' customer issue. This is basic customer relations / crisis management stuff.
I worked at a company of 5K+ people and one of the folks in control of the twitter account(s) would come to me with questions.
Now I applauded them for coming to me for technical questions before posting, that was great, but they absolutely did not have the self awareness / understand what to say / when to say it and etc.
But hey they were tied to a high ranking person (who also had no clue) so they had access to the account.
In my early days I worked PC customer support... I feel like that comes in handy all the time.
That's an AI-hard problem and remains unsolved.
All cloud providers make it absolutely clear, in black & white, that protection of your data is your responsibility, not theirs.
What I find hilarious is that most cloud providers only provide built-in backup functionality for a tiny subset of their services.
Ask Microsoft if you they have a "backup" button for Azure DNS Zones. Or Azure load balancers. Or anything else that isn't a VM disk, App Service, SQL Database, or a Secrets Vault.
I mean, look at this insanity: https://docs.microsoft.com/en-us/azure/backup/backup-azure-f...
"Backup for Azure file shares is in Preview."
After 10 years of operation, this trillion-dollar company has only a use-at-your-own-risk beta for data protection!
Don't be too hasty to point fingers at Ghandi and laugh about how they're unprofessional. Whatever you're using is essentially the same.
Ask yourself this: Could your organisation recover if some malicious admin simply deleted all Azure Resource Manager resources in one go using PowerShell?
This thought occurred to me when I was testing a bulk resource creation script.
My workflow in my lab tenant was:
1) Bulk create hundreds of resources 2) Bulk wipe everything 3) Go to step #1
Turned out, I had some objects with globally unique names that were now conflicting in the production tenant, so I had to wipe my lab.
I had already logged on to the production tenant, and I was so "trigger happy" that I very nearly ran my bulk-erase script against the wrong subscription.
It was a terrifying moment of clarity.
I haven't ever even lost a file on Google Drive, which as far as I know provides no reliability guarantees at all.
The rarity is immaterial, the responsibility for data protection lies with you, not them.
Of course it's material. If a provider has a 0.001% chance of losing some of my data in a year, I'm an idiot for not having backups. If a provider has a 10% chance of losing some of my data in a year, I'm an idiot for not having backups and for using that provider.
GMail is (usually) not an enterprise product and not a paid service, and provides no reliability guarantees. And yet it seems to be pretty damn good in practice.
We have streaming replicas for hot data AND regular snapshots shipped to offsite cold storage, because RAID is not a backup. If we experienced an equivalent event, we'd be fine.
How long will it take you to recover if someone deleted your switch configs, reset the SAN to factory defaults, wiped you firewall rules, deleted you Active Directory accounts (or equivalent), and then ran a secure erase on every every physical server just to raze everything to the ground and salt the earth?
I mean in wall-clock time, how long would it take your team to even figure out what is going on? Where would you start?
Would you recover the switch first, or the server that you use to authenticate to it using RADIUS or LDAP?
How will you securely connect to servers if your CRL and OCSP servers are down?
How will you get access to your passwords if your file server where the key blob is stored is saying "Insert boot disk"?
People think that disaster recovery is for "I deleted a folder".
Disaster recovery is for disasters.
Removing all Azure resources wipes everything. Your vNets... Poof! Your public IPs... Poof! Your internet-facing DNS zone... Poof! Your authentication credentials... Poof! Gone, gone, gone.
How do you plan to restore dynamic IP addresses to their original values?
How do you plan to restore DNS Zones that get assigned to 1 of 10 randomly selected server pools and hence have a 90% chance of requiring a change to the NS server glue records on restore?
Do you even know which order things would have to be restored in to prevent failures during a restore?
Could you possibly work out what is missing if you log on to your cloud portal and see the "Welcome to Azure, to get started click here" splash page?
Get it?
It just occurred to me how much easier it is to wipe everything in the cloud age than the on-prem age. Doing all the things you said for on-prem takes some serious effort. Some, like factory resets, may be impossible without individual physical access. You would probably be discovered and stopped before you can inflict much damage. In the cloud age however, it takes orders of magnitude less time and effort to inflict the same damage.
It is kinda like how much easier it is to steal data now. Before the digital age, stealing as much data as Equifax hack would have required moving truckloads of paper without being discovered. It was simply impossible to pull it off in reality. In the digital age, however, we have accepted massive data leaks as not only possible, but unavoidable.
It's easier for physical facility damage to a single facility (whether hostile action or natural disaster) to wipe everything out in an on-prem setup than in the cloud, where multi-DC redundancy is a click away. But, sure, it's easier to wipe out data without physically destroying equipment in the cloud.
Consider the current tensions between Iran and the US. If Iran decides to retaliate with cyberattacks, major cloud vendors could suddenly have multiple regions go up in smoke concurrently.
They'll just shrug their shoulders and say that it's the customers' responsibility to protect their own data, and that they're just offering platforms for rent.
After all, they both have wings and will both kill you if they fall out of the sky, and I don't see Airbus or Boeing guaranteeing that their planes will never crash, so they must be essentially the same.
That just confirms the parent comment
This mail is a follow-up to the previous email we sent (on January 8th, 2020) on this topic. As a reminder, yesterday, we experienced an incident on a storage unit at our LU-BI1 datacenter, located in Luxembourg.
Despite the replication systems in place, and the combined efforts of our technical teams throughout the night, we were unable to reover the data that was lost on the impacted storage unit.
We sincerely apologize for the inconvenience that this situation has caused. This type of incident is extremely rare in the web hosting industry.
In the event that you have a backup of your data, we suggest that you to use it to recreate your server at a different datacenter.
To help you in this, we have provided you with a promo code that will give you one free month for an instance, so that you can create a new Simple Hosting instance in a different datacenter:
XXXEdit: in fairness, I'm not sure how exactly you would quantify such a loss anyway...
Edit: who knows it may be related to the HPE issue.
https://www.bleepingcomputer.com/news/hardware/hp-warns-that...
Especially when a small set of your customerbase is affected, it won't cost you that much, and "overcompensating" like that means that virtually noone is going to criticize you for quantifying it wrong; instead, the public narrative will be centered around "well, shit happens, they did their best and generously compensated".
From[1]:
> Traffic over the private network does not count against your monthly quota.
I wonder how private addresses are setup by Linode.
[1] https://www.linode.com/docs/platform/billing-and-support/net...
https://www.linode.com/docs/platform/manager/remote-access/#...
To answer your questions: Yes, the backup storage box is in a separate chassis than the host machine that the Linode lives on; they have separate power supplies. The DCs themselves also have some sort of fire suppression. I don't know what would happen if there was an explosion.
1. Power delivery systems that bring power to the buildings - see issues at 111 8th Ave failures during Sandy.
2. Power systems inside the data center. Blast radius there is rather nasty. See the infamous Internap blow up around 2015(?).
3. Fire suppression/firefighting protocols.
Obviously hosting providers do not make it easy to extract your data because that's their vendor lock.
Gandi has never explicitly said they never had their own backups, just that they don't offer backups as a service. It's entirely possible that they did have backups, but couldn't recover/restore them.
And to "trust marginally more" simply means:
gandi_cost_per_month + P(gandi_fails_per_month) * cost_recovery
<
alt_cost_per_month + P(alt_fails_per_month) * cost_recoveryWorse than a bad incident there is only bad management of the following situation.
Why would they include that sentence? Are they trying to imply it is rare for them because it is rare for the industry? Are they saying they are not as good as the industry, so customers should move to other providers? Or are they trying to show they apply the same inattention to their customer communication as they apply to their data backup/recovery practices?
This kind of data loss should simply never happen. It’s one thing to say “it will take us up to 30 days to restore your data because our fast recovery options aren’t working and we have to bring up cold archives”, it’s entirely another to say “your data is gone, tough”.
I read it as: "This type of incident is extremely rare in the web hosting industry, because apparently the overwhelming majority of our competitors aren't capable of fucking up as badly as we just did."
Doesn't inspire confidence at all, IMO.
They're a French company; it may be a non-native speaker not catching the implication.
It's also possibly an editing error, e.g. they started writing something like, "these types of incidents are extremely rare and when they happen etc" and most of it was dropped without considering how that changed the implication.
I read this as "so maybe you should consider one of the other web hosting companies that doesn't have problems like this."
If you're in such a situation, The easiest is to do filesystem level backups with something like zfs and ship the backups to a third-party system that only has write/append-only semantics (better yet, use a write-once-read-many (WORM) disk to really guarantee it.).While there will still be _some_ data loss, it'll let you recover since the last snapshot.
If you don't have zfs, a database backup that runs the db dump script and scp/sftps it to a server running as a cronjob can also be an immediate remedy while you get your shit together (and by that I mean buy yourself a product with an immaculate reputation like aurora or cockroachdb to manage the db for you)
Harder but better would be to tee the log of the changestream (all distributed systems have such a log) to a third-party system. This is ideal because if it's done synchronously it'll let you recover since the last committed transaction.
And of course, test your backups, because backups are subject to code rot as well.
You could replay the (combined,sorted,agumented) changefeed in-order, or shard it on the table's primary key to ensure per-key monotonicity when applying the streams in parallel threads/transactions/nodes.
Their response was awful and rude and completely unprofessional. I never got my domain back.
Based on that experience, this incident doesn’t surprise me at all.
The courts are too expensive. The culture of taking pride in one's work maybe is disappearing.
For the most crucial parts of doing business/living life, we are required to trust someone else. For example, I can't just go and make my own cell phone tower or ICANN.
And yet I can't even trust those entities to get it right.
There's got to be a measurable (negative) economic impact.
Do they have that in their terms? Independently of that, do they have a history of doing that?
Everyone's website says they're "no bullshit". It's all bullshit.
But at the end of the day it's a cheap provider with, ahem, French-style support so I'm not sure what people were expecting out of them.
Gandi is the only domain registrar I’ve had an issue with.
Avoid GoDaddy at all costs!
I consider GoDaddy to be one of the worst companies in existence, as bad as anyone else you can think of, as free of corruption as current ICANN and as fraudulent as registerfly. Clients looking at available domains have found them immediately registered and squatted at {hundreds}% markup. Their incompetence lost me a few domains, and several freelance clients reported similar -- all of whom were paying vastly over the odds for what they were getting. GoDaddy make Gandi look an exemplar of ideal behaviour for behaving as people are reporting in this HN post.
Their previous CEO had domain squatting and a complete lack of personal ethics as sidelines. That's quite apart from their horrific upsells making a simply renewal a 22 page nightmare of deeply dark patterned "no" clicks against atrocious value "offers".
Keeping registrar and hosting separated seems like a good idea.
Edit; I use (and have been for a very long time) namecheap for registration and (recently) Cloudflare for DNS. I used to host all DNS myself, but that became a bit of a pain with many domains as that's definitely not my core business.
Their response to this is exceptionally poor. To say essentially "this could happen to any other web host" it nonsense. I've never had this happen with any of the providers I've used for hosting and I'd be very angry if I had just lost an entire VPS. The fact that they've lost all snapshots as well (which are advertised as backups of the underlying volume) is unforgiveable.
I've used support twice when transferring domains into Porkbun and they were good. I transferred a domain out and has no issues. Their 2FA options are really good. They frequently have the best prices around (tld-list.com).
I had one weird incidence with trying to host wireguard on it, I couldn't get it to work reliably even after changing MCU to suit them or trying other fixes.
[1] https://docs.aws.amazon.com/Route53/latest/DeveloperGuide/re...
My machine going away because you had hardware issues isn't my problem, and I'll spend my money on a more competent company.
Always have your own offsite backups.
In some extreme cases, a concert of bad luck may coincides to ruin things despite multiple levels of redundancies. But that's extremely rare, especially nowadays. However DO is much larger now than it used to be, so the odds of hearing about extreme accidents increase.
disclaimer: I used to work there.
Good thing we only kept caching servers in Digital Ocean, so those were easily recreated, but that always kept me away from DO, personally.
In fairness to them, though, DO do not claim to keep backup of the servers, as far as I know.
[0] https://news.gandi.net/en/2019/02/futureofgandi-the-adventur...
That's a very strange blogpost.
> we have found a new investor in Montefiore Investment, who have replaced our former shareholder!
Am I missing something?
I've used them for many years and had several complex support interactions with them.
Their customer service policy is very "API-like" in that you get exactly the t&c you paid for and nothing more. Hand-holding and soothing noises are not included in the t&c. They fuck up you get a refund, you fuck up they'll tell you exactly that. Outside that they're very casual relaxed humans to communicate with.
I find that far more trustworthy (in the mathematical sense) than a "slick" twitter feed.
Politness does not imply trustworthiness.
The last time I tried buying a domain through them, they took my money and then demanded "identification" via government ID (citing some bullshit in their ToS). I refused, so they closed my account and took the domain with them.
Based on that, I'm not surprised at all by their CEO's response to this incident[0]:
>If we led you to believe that you had nothing to do on your side when warned multiple times to make your back ups, then we'll have to make it clearer, and stop assuming that it's an industry wide knowledge.
[0]: https://twitter.com/StephanGandi/status/1215287619938062342?...
After more than 8 years with Gandi, not had a single issue with them.
And honestly, if you don't keep data of stuff you host on a server provider like this, you kind of get what you deserve...
Pretending that the cloud is permanent in infallible is extremely dangerous. I would seriously question the competence of any sysadmin relying on this as a base principle.
Sure, they screwed up, but this stuff happens. We should actually be happy it happens "only" on a "small-ish" provider like Gandi and not an entire AZ at Amazon.
Can't wait for that shoe to drop, I'll bring the popcorn, if there's anything left of civilization then...
As far as I understand correctly they only made snapshots on the same machine, which is why there's trouble to begin with.
Considering they're currently "reminding" customers that backups are an industry standard right after losing data due to missing backups I wouldn't just shrug it off.
The fun, of course, starts that one time when it does not work and you realize that no one looked at the corner case that bit you.
Backups aren't free. Replication isn't free. DR isn't free. If a customer isn't paying a premium for them, they aren't getting them. Read the terms of service.
See full thread. Snapshots are marketed as backups.
Intelligent people can argue all day about whether a snapshot should be considered a backup or not, but it won't change the fact that a snapshot doesn't provide any protection from a failure in the underlying storage and it's ridiculously foolish for the owner of data to solely rely on snapshots as their backup strategy.
That depends on how snapshot storage is implemented by the hosting provider. They can use different storage for it, or tapes or whatever. On AWS I can easily have my snapshots on Glacier or copy them to a different data center.
Having cloud provider X say they moved the bits from one place to another should not be considered a backup by anyone, regardless of what they advertise.
- Was this mostly a power loss or a data loss?
- If data loss, did this affect EBS (which has had a claimed annual failure rate of 0.2% - 0.5% or so if I remember) or S3 (much lower failure rate). Remember, EBS WILL have volumes go bad - that's in the docs, they recommend snapshots, aws backup manager etc if you need higher durability.
Sure, let's blame the victims here; that's effective and helpful.
Give me a break... It's not like anyone died here. There's a reason I host my own shit. Problems happen, errors are made, and data is lost. It's also your responsibility to deal with data permanence, even if your provider has all the promises in the world.
A company violates their agreement with you in a way that costs you time, money, and potentially business, and you're not the victim?
They are fools on their side for failing to preserve user data, but you end up being the bigger fool for trusting them to do this for you without preserving a backup plan yourself.
for purposes of keeping your data safe, your cloud provider is just one, single, copy of your data. all of their redundancies and backups and whatnot are for _their_ convenience, not yours, regardless of the marketing copy.
(they can decide to intentionally delete your data because they think you didn't pay. no amount of RAID and georedundant backups on their part will help you then.)
While I agree that everyone should have their own off-site backups, this does come across as incredibly crass victim blaming.
With that having been said, everyone please stop assuming your data is safe. It’s never safe, but it’s extremely not safe single homed somewhere. Make backups. Anything that’s saved locally on one machine only? Consider it gone until it’s backed up.
Cloud providers may be able to give you better assurances, but if you really care about data give it at least 2 independent homes. I’ve lost data more than I care to admit. BuyVM lost one of my VPSes years ago. Who’s fault was it really?
When you are ready to stop kidding yourself about your data, check out some backup solutions. I particularly like Borg Backup:
https://github.com/borgbackup/borg
And if you do not have network attached storage anywhere there are services that provide it as a service.
(Note: I think needless to say it’s also a good idea to back your NAS up to other places too, although I haven’t gotten into this practice yet. Synology supposedly has a lot of features around this.)
As a former hosting engineer, at the risk of pissing on everyone's outrage parade, but unless an explicit guarantee of a backup is included in your plan's contract, or you can pay for backups as a bolt-on, then if you've lost data it's your fault for not planning for this scenario.
And I mean proper backups where you get, for example, twenty eight days of hourly backups and you can pick a specific version of file to recover in that 28 period. And where those backups are stored on different hardware or off-site. We offered this as a bolt-on (in-site and off-site). Tt was 20 quid a year for in-site, the off-site was a bit more. But a great many customers chose not to pay for this add-on, even despite the great big red bold warning text explaining that unless they paid for this add-on we made no guarantees about the permanence of their data in the event of a storage problem. Guess what....
Now that's not saying we didn't take snapshots of the hosting environment, but they were for internal use and to allow us to recover quickly in the event of something unexpected going wrong, but now and again stuff breaks.
Sure, it's unfortunate some lump of storage hardware has failed and whatever mirrors they may have had have been taken out as well. They possibly could have done better but shit happens sometime.
You shouldn't rely on an "implied backup" from your service provider, if you want that then you're going to be paying a shedload more for hosting your Wordpress and Woocommerce site. It's up to you to make sure absolutely sure your data is safe if it's critical to the day-to-day running of your business.
Edit: ok, so this is tucked away in their docs (thanks to itake below):
https://docs.gandi.net/en/simple_hosting/common_operations/s...
But it does say:
> Snapshots do not make a backup of your databases. If you would like to perform a backup of your databases, we recommend you perform an export, or launch a dump script via crontab.
The bottom line...is it guaranteed in your contract? Always check. And as per my follow up comment, those plan prices are are just too cheap for that facility to be taken seriously for business continuity. They're a convenience to quickly recover a version of a file, not a serious backup.
https://docs.gandi.net/en/simple_hosting/common_operations/s...
They are supposed to be providing backups.
There are different kinds of backups here:
* the ones that are part of the offer, where the provider gives you a convenient way to recover from your mistakes, this is a feature they provide when their services are operational (in this case, the snapshots feature).
* the ones they put in place to mitigate incidents and maintain their SLOs. If you accidentally delete a file, you don't have access to them, they are useless to you. These backups are a mean to reach their service level objectives. Nobody can offer you 100% guarantee that they won't lose your data in an SLO. If someones promises you this, just... don't believe it.
(edit: formatting, typo, mention snapshots in case 1)
Also:
> Snapshots do not make a backup of your databases. If you would like to perform a backup of your databases, we recommend you perform an export, or launch a dump script via crontab.
For those plan prices if I was running anything mission critical there then I'd be making darned sure I was squirting copies of my site's dynamic data to somewhere else on a regular basis (and you should also be able to re-deploy your code from local). Even if there was a guarantee, I'd still have a backstop in place. Never underestimate the chance of a good cockup.
You have to be foolish to assume you get proper, actual backups for the price of a Coca-Cola can.
To be honest this is not a good way to do hosting business. If you provide a service called "Simple Hosting", putting backup requirement on customer (when it is your fault) is pretty unfair.
PS: I think price of the product shouldn't effect minimum requirements.
Then they need to pay more for their "Simple" hosting.
> I think price of the product shouldn't effect minimum requirements.
See above. Sigh.
- Some airline is selling plane ticket and insurance on website. (insurance covers change of plans, rebooking etc, and even if you don't fly that flight, they are booking you another one same day)
- Then when a flight got canceled, telling customers "we rarely cancel flights, please use your insurance. (you should have bought insurance)"
PS: Simple hosting [0] I am referring seems like managed hosting.
Do most people actually do this? I never do.
If an airline cancelled my flight, did not provide alternative arrangements, and cited some legal fine print instead... then I would be very upset. They might be legally in the right, but that wouldn't prevent me from taking my business elsewhere.
Can anyone recommend a domain registrar "equivalent" of a Fastmail or Letsencrypt or DNSMadeEasy i.e. truly no bullshit, geek friendly and polished at the same time ?
I'm not too bothered about price. I just want a well run outfit that has a wide selection of TLDs and ccTLDs (and ideally isn't a mega corp like google but is big enough that I don't have to worry about them disappearing overnight).
I've got a few bookmarked, but I haven't tried them: Porkbn, Nuage, Hover, and Namesilo.
in another comment on this thread, someone pointed to cloudflares new registrar offering, which also seems good.
AWS Shared Responsibilities [1]
Flipping a switch that says "Backup" does not mean you are handing your responsibility to them. At most, they will fail to meet their SLA, write you a check for according to the TOS and be done with it. At best, you'll be able to bitch about it on Twitter, possibly threaten a lawsuit (you read the ToS?) and still be in the same position because you did not share the responsibility of securing your data.
[0] https://docs.microsoft.com/en-us/azure/security/fundamentals...
[1] https://aws.amazon.com/compliance/shared-responsibility-mode...
Why are they speaking of the "industry" as a whole when they are to blame?
It's even crazier they are not even explaining the source of the data loss and why the "replication systems" didn't help.
IHMO they are trying to sweep this event under the carpet. They should instead explain why they should be trusted in the future and why this would not occur again.
Let's say a typical admin of a small shop wants to backup his postgres database.
The first thing he'll use is probably pg_dumpall which he'll output to a storage.
No replication involved. The backup is just a bunch of sql statements to recover the last known state of the database. It's a different kind of format however, which -by definition- isn't a replica anymore.
(And this process has several caveat's-one of which is that it can produce unusable dumps in some rare cases and isn't complete. users, triggers etc aren't dumped iirc.. could be wrong there)
You probably should look up the definition of “replication”.
All non-trivial replication has to cross machine boundaries. To transmit to another machine, you have to use a serial format since there are no pointers on the wire. So insisting that a replica must be the same format prohibits the concept of replication in practice.
How is it a copy if it can't recover the original in some cases?
We're not talking about a compressed archive here, it's a (incomplete) step-by-step instruction to recreate the data. If anything out of norm happens, it's gonna fail-possibly silently.
Do you think a gun is not a gun if it sometimes jams?
> We're not talking about a compressed archive here
I think we have two camps. Mine is considering "copy", "backup", "replica" to be broad categories that are distinguished by simple mathematical or technical properties. For instance, I'd consider a device that copied a single bit to be "copying," even though it's arguably just a wire.
The other camp has very specific products and tasks in mind. A replica is associated with distributed computing, while a backup is something a systems administrator makes as part of disaster recovery.
but a pgdump is like a step by step guide for building the gun, leaving out a lot of the process... how can you honestly call that a gun?
but i guess we'll have to agree to disagree. which proves that it was a discussion we shouldn't have started i guess.
Triggers are dumped, users need pg_dumpall (as they live across multiple databases, same with tablespaces).
> a backup, or data backup is a copy of computer data taken and stored elsewhere so that it may be used to restore the original after a data loss event
Since a "replica" is a copy, that seems technically correct.
The original claim was "a backup is a replication of the live dataset, although, usually out of sync to be useful when the main dataset goes bad."
The only thing that makes a replica special is that it's in sync. Once you add the caveat that it's out of sync, it's just a copy.
Now, granted, it'd be a huge pain to track down all the people who had copies of the 1,500 different repos, and try to find as up-to-date as possible of a version of each, but I doubt they got anywhere close to potentially losing all their source code.
Incidentally this shows why it's a good idea to sync your repo to GitHub, even if the canonical repo is elsewhere: in addition to the usual reasons of incentivizing some contributors by giving them "GitHub credit", and increasing visibility of your project's code, GitHub can serve as a backup!
Also, on a side-note, 1,500 separate repositories?! That sounds way overkill. I wonder if they'd benefit from having a monorepo.
No it doesn't. Github has at least 20 million public repositories. Would they benefit by combining them into a monorepo?
And yes, a monorepo is usually the best approach in most cases for a project or even an entire company.
They're 100% virtualized and keep backups of all those machines. In addition, you can purchase a package so that YOUR backups are automatically backed up to 2 different datacenters. Between the two of those solutions, there would be a way forward.
I don't really know what Gandi is so I can't speak to them directly, but this is a solvable problem.
Why would anyone believe a motto is anything other than a marketing device? It is only believable if people follow it contrary to pragmatism. Any company is eventually going to have a fair share of people who believe being pragmatic is more important than their motto. And in Western culture at least it's usually considered rude to bring up the "big guns" and have a fundamental values discussion when everybody just wants the meetings to end and to start making more money.
Fujitsu – “The possibilities are infinite” ... "The Ways in Which we can Screw this Up Are Infinite"
Intel – “Leap Ahead” and “Sponsors of Tomorrow”. "We've got to protect our entrenched position".
LG – “Life's Good”. "Life is Actually Objectively Bad".
Google - "Don't be Evil". "How We Actually Make Money is Evil But Our Mission is Good".
Above all, "no bullshit" is our golden rule—to treat our users how we want to be treated. It's a promise to respect your rights and to level with you about our shortcomings.
https://www.gandi.net/en-US/no-bullshit
ex: https://twitter.com/andreaganduglia/status/12151991477012316... (thanks op)
We will listen to you, and be honest in our replies, even if it means you won’t always like what we say.
They are actively treating their customers like shit, and that tone starts at the top. No bullshit does not give creative license to be assholes to people that are panicked because of something you directly caused.
https://github.com/StackExchange/dnscontrol
(Terraform users have a similar benefit)
This is basic stuff.
those outages costed millions of euros, and he never picked up his phone at night, once I asked him why he never picks up, he told me:
"I used to be a general surgeon, when someone calls me people die. Relax, nobody is dying during our outages."
now I think I am taking myself(and my work) too seriously.
Hence he's not a surgeon anymore.
I've done it before, but I'll recommend to you Scott Adams' book, The Dilbert Principle for some light reading about forces like that at work.
> we have a problem to import zfs pool on the unit storage
I really want to know what went wrong to a) break ZFS b) prevent recovery from backup.
* https://news.gandi.net/en/2019/09/exporters-detect-micro-inc...
> Gandi’s storage infrastructure consists of two environments: one for IaaS and one for PaaS. Both are based on FreeBSD-based storage units (filers), that stock each volume (disk) as though it were a ZFS volume.
* https://news.gandi.net/en/2019/03/tracking-a-storage-issue-l...
* https://www.bsdcan.org/2016/schedule/attachments/351_FreeBSD...
No mention of what they're doing for backups / "replication systems", unsurprisingly/unfortunately. I'm anxious to know what the failure mode for `zfs send | zfs receive` replication is here?
Updated on Thursday, 9:58 PM +0200: we're not sure we will be able to provide the data but we were able to recover a version of the filesystem from right before the crash
I have been using them for DNS and some minor hosting for a long time and I will stay with them. I think it's important to avoid the monoculture/centralisation which is otherwise happening.
Sure Gandi has their flaws, they are humans.
I expect they do will a proper post-mortem on what went wrong and how they managed to fix it. Seems they were using ZFS and relied on it a bit too much. Or if they indeed managed to restore the last snapshot, then their only error might have been the classic one of underestimating how long restoring/investigating several terabytes take even on modern HW.
3TB at Gandi costs $6 + you get compute with it. 3TB of bandwidth at AWS might be $270.
Has anyone tried this instead of using cloudfront etc? Get 100 $6 hosts and pump out content for your ipv6 connecting clients etc?
Since Gandi is mostly known for domain registrations and DNS, I'm curious if you (as an individual who hosts websites/online services somewhere on the web) backup your site's DNS records periodically (or whenever they're changed). What if your authoritative name server lost data and all the caches of those records across geographies expire while you're asleep/away? If you do back these up regularly, how do you do it in an automated way on a *nix system? I found this article [1] when I searched about this, but it's not a simple shell script. The scripts that I did find on some of the Stackexchange sites seemed to have specific subdomain names hardcoded.
[1]: http://www.programblings.com/2012/07/23/do-you-back-up-your-...
After many many years of experience with systems, I made sure we had as many possible ways to recover user data as we could. The initial solution was a large Postgres database for all the metadata/indices and s3 for the actual storage.
Despite much pushback we built in little things like an individual meta file on the file system for each file we stored. That way, if we lost the Postgres dB for any reason, we could create a script to rebuild the dB and restore access avoiding massive counts of orphaned files. A simple and probably stupid solution but...
Well guess what - the DB got corrupted and after some ado, we restored all access and none of our customers lost anything.
No it’s not full backups but...
with modern container hosting you really should be able to make your own. even with cheap VPS hosting. There is no reason to live in a world where a server goes down you lose anything anymore.
Lost 3 sites built with WordPress. Will rebuild as static sites repo separate from host, no more database, lesson learned.
Set a retention policy I.e even if someone ran some delete command it wouldn’t delete. Someone with retention lock permissions is the only one that can remove the locks And delete.
There is cold storage and other things even cheaper. But cloud object prices are pretty cheap per GB it’s ridiculous.
I think they make most of the margins on bandwidth.
Losing customer data. All customer data is pretty ridiculous. I can understand downtime. I can understand losing a day of changes. But everything? That’s just unacceptable business.
Considering how critical email is for me, seems like I won’t be trusting their MX servers to process all my inbound mail anymore and will soon be looking for another solution that works well with Gmail (don’t want to pay for GSuite), and possibly also transfer my domain to another registrar.
That support tweet is such bad taste.
unless you run all the MXes (and can prove otherwise): you're likely having emails dropped all the time already
an example from 5 minutes ago from one of my MX'es (which only forwards, after heavy greylisting and spam filtering):
Jan 9 18:05:59 mail postfix/smtp[26197]: to=<ABC@gmail.com>, orig_to=<XYZ>, relay=alt1.gmail-smtp-in.l.google.com[209.85.233.26]:25, delay=25000, delays=25000/0.01/1.6/0.16, dsn=4.7.0, status=deferred (host alt1.gmail-smtp-in.l.google.com[209.85.233.26] said: 421-4.7.0 Our system has detected that this message is 421-4.7.0 suspicious due to the very low reputation of the sending domain. To 421-4.7.0 best protect our users from spam, the message has been blocked. 421-4.7.0 Please visit 421 4.7.0 https://support.google.com/mail/answer/188131 for more information. ABC - gsmtp (in reply to end of DATA command))
that's not my reputation (which is high), that's the reputation of the sender's From address
and it doesn't send it to the spam folder, it just delays the email forever until my MX gives up
sending to outside gmail I suspect it will count against you (though major providers likely have special treatment for Google's MXes)
If you're in Europe, they're cheap for many European countries' domains.
Back 10-15 years ago they were special because it felt like a hacker kind of company. They gave free WHOIS privacy, what seemed like good DNS control/UI at the time. But it was the WHOIS privacy that got me onto them.
I still use them because they're around half the price for .co.uk than many registrars - and many others I've used have become more rubbish than Gandi has.
All my DNS is hosted elsewhere now, and I never understood why Gandi introduced hosting et al. I've never used it and never would, it seemed a terrible diversification for a good domain registrar.
I think majority of people buy domains for hosting websites so it makes sense they would want to setup one using one click WordPress or something similar.
https://news.gandi.net/en/2020/01/major-incident-on-our-host...
A site of mine is also hosted as their PAAS at Luxembourg, but was luckily not effected. Probably my site was on another storage unit ("on one of our ZFS storage units").
PS I also always thought that the snapshots were backups.
Is there a registrar you would recommend as an alternative, I don't need DNS, nameservers and glue records and I'm ok.
The main selling points are stability, transparency and simplicity. I don't care if it's not the cheapest.
Will likely transfer those to elsewhere after this. Probably Name.com, I guess.
> Updated on Thursday, 9:58 PM +0200:
> we're not sure we will be able to provide the data but we were able to recover a version of the filesystem from right before the crash
Maybe it's not all gone.
Is that a lot of data? That sounds like a very small filer that could have easily been backed up.
>we have a problem to import zfs pool on the unit storage. Our engineers are still working on it.