Cloudflare silently deleted my DNS records
txti.es
txti.es
BTW If you, dear reader, ever find yourself so frustrated with Cloudflare that you feel like your only recourse is a blog post... my email is jgc@cloudflare.com and I’m happy to hear from people.
It's likely a good learning for all.
You can post updates with any relevant information. Probably goes without saying, but if the issue has to do with my billing or address please don't post specific details without asking me first.
I will link to this comment from TFA for verification. (Edit: added to the bottom. If you need more verification you have my email.)
Edit2: I see that the domain is back in my account and listed as "Pending Nameserver Update". I don't think that's because of something I did.
I do this sort of thing all the time. Sure, it’s unfortunate this is #1 on HN, but shrug. Fixing the problem and figured out what happened is important.
What is the best email to reach out to you?
Thank you.
I remember about 4-5 months ago I spent like 2 weeks going back and forth with Stripe's regular email support trying to understand their docs for SCA.
I kept getting a new rep who repeated the same things the previous reps were saying, which also had no bearing on what I was asking. It was basically a copy / paste from a script loop.
Then something negative about Stripe was on the HN front page and I happened to comment about a bad experience with the new SCA docs.
Within a few hours I was put in contact with a lead developer from Stripe who went as far as creating custom flow charts for my use case that wasn't covered in the docs and it was a pleasant experience, where "pleasant" felt like the person receiving the email was reading the words I wrote instead of just skimming them and pasting a boilerplate response.
But it only happened because of the HN comment. If that thread never appeared on HN, I'm not even sure I would still be using Stripe.
Crazy how far down the bar has gone.
I switched a personal domain to them, but it won't be long before I get my employer to move their resources over to them, assuming I see this all play out well.
Maybe it's a cloudflare issue, maybe it's an honest mistake, maybe they are bad at customer service... either way, it's not great to shame a company before you even open a support ticket or talk to someone to find out.
To combat, hey my issue is always a Sev1 ticket, one can probably institute something like, here is a red button and if you click it, we will charge 100$. If it is indeed an issue that caused you to lose 90%(say) traffic and it was our fault, we will return the money.
I would trust Cloudflare to pay me back, and I'm already putting something on the line. If it really was my fault this is going to be very embarrassing.
Edit: In case it isn't clear from context I'm the OP so I'm pretty sure I know what I feel.
The vast majority of support tickets are the customers 'fault'. They'll be things like setting up DNS records incorrectly, believing they don't need to pay a renewal with their old registrar just because they moved their DNS over to cloudflare, frustrations over cached resources, etc.
Essentially I felt that this was alright because when I filed a ticket I was informed that I should expect a long wait and that they recommend that their non-business customers post publicly on their support forum for crowdsourced support because that leads to faster replies. I was unable to log into that forum, and I suspect that may be because the way they set up SSO between the forum login and their main login may have failed in Firefox (with all tracking prevention and ad blockers disabled).
I felt that if a company invites me to ask for support publicly on their forum to save on customer support costs it's reasonable to talk about the issue in another public place.
A fair number of questions aren't unique - product questions, how to use an API, etc. Someone may have asked a similar question, in which case you'll find an answer, find it faster than it'll take to hear back from the support team, and it deflects an unnecessary (already answered) ticket. That should be a win all around.
Now if you do have a novel question or something account-specific, by all means, open a ticket. There you'll get replies from people who can look up your account and give you specific answers.
The ombudsman tip in this post doesn't make a whole lot of sense when the normal support process wasn't really given a chance before making the blog post.
1. If you write into support as an Enterprise customer, you're basically getting an automatic response.
2. Enterprise customers can dial a direct line to support and have this under investigation within minutes.
3. Enterprise customer have a dedicated account manager, and in most cases a Solutions Engineer helped get everything setup. Usually Enterprise domains get locked, so they cannot get caught by linting type services. The likelihood of a domain being removed under those conditions is very low.
4. The most common cause of this type of thing is the Name Servers no longer pointing to Cloudflare, no one noticing for a while. Cloudflare periodically checks to see if a domain is still using CF name servers, and if they aren't they get moved as the assumption is they are no longer on Cloudflare. I don't know that this is what happened here, but it's easily the most common issue on lower tier plans. Enterprise companies often have people monitoring their infra, thus they catch this before Cloudflare conducts the removal of the domain.
Source: used to work on the Cloudflare support team.
He _tried_ to request help in their recommended public forum, but could not due to login issues.
Seems entirely reasonable to me to then go and ask about out in a different public forum?
> I spent some time thinking about if it was fair for me to post this on the same day as I filed a support ticket with Cloudflare.
Technology companies need an "ombudsman" - a contact that customers can go to when the normal tech support processes have failed.
The Ombudsman must not be part of the technology companies ordinary support processes, it must be entirely separate, and have highest level authority to demand action within the company.
To avoid the Ombudsman being overused, you could give it a price of say $20, which is always refunded when the case is resolved.
HN constantly has front page posts from people for whom big tech companies have support processes have failed but there is simply no other recourse unless you have "a friend in the business".
It just doesn't work to have some random Cloudflare person offer their email address as some post disaster issue resolution process on social media. Formalise it with an official Ombudsman and maybe then companies like Cloudflare might avoid HN front page bad publicity.
I had an issue at "one of the biggest tech companies" that went on for days and days in which tech support kept telling me I had set up something wrong, until eventually I emailed one of the top managers who I happen to "know" at that company - it was fixed within hours. That "contact a friend in the business who can actually get things done" is a necessary part of a large support organisation and it simply does not exist yet in any tech company that I know of.
The Ombudsman role is there to get things fixed when all else has failed, and before the angry customer posts to social media.
On the other side of the equation sometimes I do want to complain for the sake of complaining without being harassed by some support account.
Especially hate getting obviously automated responses for daring to mention company names even if it would actually escalate to a human.
It shouldn't matter, and it should not be required, that someone "known and important" within an organisation decides to start doing hands on tech support in social media following a PR disaster.
If "jgc" is actually someone important within this company then maybe after fixing this issue, they can then go fiox their tech support by setting up and ombudsman and get their PR disasters off the front page of HN.
What’s a recent development is the complete lack of support when shit goes south. Back when you were interacting with real reps you had people that could see when stuff was obviously wrong and escalate appropriately.
From what I read this is nothing to indicate the process failed (so far) just that the user decided to skip to the head of the line by writing a blog post internet style to get something resolved and attention. Failed is not 'I didn't get a reply or find what I needed as quickly as I think it should happen so now let me complain publicly so I get a reply'.
> To avoid the Ombudsman being overused, you could give it a price of say $20, which is always refunded when the case is resolved.
In theory nice but first it would be a 'deposit' and also opens up a host of new issues as far as the money being paid back and how that would be done and so on.
From my perspective Cloudflare's process did fail. Assuming I didn't do something insanely dumb and what I think happened did happen, I would consider that a failure on its own even if the support was perfect afterwards.
Copying from elsewhere:
I address this in TFA.
Essentially I felt that this was alright because when I filed a ticket I was informed that I should expect a long wait and that they recommend that their non-business customers post publicly on their support forum for crowdsourced support because that leads to faster replies. I was unable to log into that forum, and I suspect that may be because the way they set up SSO between the forum login and their main login may have failed in Firefox (with all tracking prevention and ad blockers disabled).
I felt that if a company invites me to ask for support publicly on their forum to save on customer support costs it's reasonable to talk about the issue in another public place.
(I think in my case it was adding google metrics from the apps page.)
All companies should have that!
The "shit filtering" workload would be tremendous..
Not saying it doesn't work, just that a lot of people wouldn't follow the process, and anyone who sets this up should be prepared for a lot of triage work.
Source: Someone who emails them from time to time to get impasses resolved.
Now, it's just another level of very poor, very scripted support.
I'm fairly sure Amazon has taken the (perhaps wise, in a business sense) approach of not caring if a small percentage of users leave, due to support issues.
The cost of keeping customers with certain support issues, greatly outweighing supporting them.
This is why you have to hunt madly around Amazon's webpage to find contact info, why all forms of help point away from contacting a person, including their chat being bots now, until you move outside of their scripts.
Obviously, Bezos may not read those emails but his aides and assistants who do have access to his inbox and act on the emails on his behalf do inherit his complete authority.
Some refs:
https://news.ycombinator.com/item?id=16341154
https://news.ycombinator.com/item?id=17193363
https://news.ycombinator.com/item?id=22286350
https://news.ycombinator.com/item?id=9356182
Very few people will complain about their lifetime Google ban after Google employs appropriate personnel to evaluate such cases.
Today very few people complain about Steam's refund policy, after Valve rewrote their refund policy to actually include refunds, after a judge ended their decade-long crime spree that saw an estimated 20,000 Australians robbed and an unknown quantity globally.
Any requests that haven't gone through the proper process get auto-rejected.
The current go-to move is to tweet a complaint at the company's Twitter account. This is surprisingly effective across multiple industries and actually was something my wife did that helped resolve a time-sensitive AirBnB issue.
Arguably this does block out poorer people from receiving "special" customer service, but there are not really other things people are willing to lose (or put up as collateral) for this type of service. I can't really ship Cloudflare my toaster or car until they resolve my case.
Maybe it could be something that is given to someone when their ticket is closed (or maybe after the first tech response... it depends on the company/ corporate structure).
That way the ombudsman has something to work with, and would slow down the barrage that would occur by having a such a public contact point.
I'm never a fan of 'pay then get refunded' for something that's not your fault, and is entirely out of your control.
If it is refundable, make it $300. That is high enough that only people and businesses with a showstopper situation will use it.
You're missing the part where some people, that actually would need this support, literally wouldn't be able to find that money because of the difference in purchasing power of the local currency.
Phillipines is a good example; Average yearly salary is somewhere around 12k USD, Median is 8k USD.
That means if someone is doing a tech startup there, 300$ is somewhere between a quarter and a half of some employee's pay for a -month-. And given the potential for tight margin/cash flow of startups even in the US... it would still price a lot of smaller players out of the market.
You need an adjustable amount that is based on the annual income/revenue of the person/entity making the request.
Make it high enough to be non-trivial, but not so high that it blocks all effective usage of the safety valve.
Now, if you can solve that problem, I've got some bridges for you in Arizona.
This is a good way to filter the full level of importance. Most places that need a problem solved, now, dammit will be willing to pay.
And, if you want to be nice, you can refund it. Or, if the person was a jerk, you can keep it.
So, all support systems should have a triage type system with a "nurse" having a constant eye on every new case that comes into the support system. When there's an emergency, such as the one associated with this post, then it should be forwarded to the ombudsman or some other emergency team immediately.
What if the cost was put onto the business instead of the consumer and the business just hired support people who are all Ombudsmans by default?
Instead of focusing on copy / pasting boilerplate scripts and answering as many tickets as possible, they should focus on the problem the customer is having by default and do everything possible to reduce the number of incoming questions by fixing bugs, making a better product, improving their docs, etc..
I personally do around the clock email support for 35,000+ people who sign up to my programming courses and support isn't bogging me down. Relative to the number of minutes I'm awake, support is one of the least business related time consuming things I do per day, but I send individually personalized in-depth answers to everyone who asks me questions -- usually within an hour or less.
I agree that in theory you can accomplish the same thing by making the product foolproof, but I don't think you can accomplish that for consumer facing products, and that doing so may not be a worth while trade-off. Additionally focusing on issues that greatly impact people rather than small things that cause friction with the product may (or may not! if it causes lack of retention) be worth more.
I can understand why GP suggested company pays $20.
Simplifying, that's "if (false)", since ignoring the customers in the limit gets you a failed business and bankruptcy.
And even if it is just a sufficiently large company, if the customer is not a large payer and the issue is non-trivial to solve, the the cost of losing that customer could very well be less than the cost of fixing the issue.
Assume the company has no way of distinguishing "valid" or "high impact" cases from other cases. This means that in order for the company to handle cases they need the sum of fully treating every case to be greater than the sum of the cost of fully treating every case. This is almost certainly never going to be true, unless each of your cases is high cost to ignore (think enterprise support).
So you need to funnel down the cases. Typically this is done with low tier support that tries to suggest fixes and such. You can also offer high value added support to give customers the ability to pay for a support plan. This $20 proposal is like that but on a more ad-hoc basis.
In technical support, one of the largest problems you have to deal with is all of the idiots that don't know how to use the product and won't look up the documentation or learn the product on their own. These people spam the shit out of support all day for the most basic shit. Seriously, I was working support at an anti-virus company and 85% of the calls into the paid business tier support were requests for us to do installs or basic application configuration for them. I know that not everybody knows how to configure a firewall, but "where do I download the install file for x" is literally Google-able.
The idea of an ombudsman is ultimately a bad one in this case because the problem isn't with the product or the support or the documentation. The problem is people Always think their issue is special and the most capable person should help them. Our support tier was paid, so more often than not we did what they asked, but the free customers would just get routed to a sales rep because free customers are even dumber and more entitled than the paying ones.
This whole situation is farcicle to me. Yet another free customer blew a minor issue out of proportion and had to apologize when it turned out the system hadn't failed him, he was just a free loading mooch all along.
Support that powerful are going to basically be devs. With dev salary expectations.
If someone uses your product or service and has an issue, the first contact they have with your company is through support.
A crappy support experience could easily be the difference between having a life time customer valued at thousands of dollars, or your customer feeling frustrated and going to another company, netting you $0. That has a butterfly effect too because having 100 happy customers who are praising your support could lead to many new sales from organic recommendations. Having a bunch of customers who felt neglected by your support could yield a situation where your company is now on the front page of HN for having bad support or worse.
For whatever reasons, bigger companies focus more on measurable metrics like "tickets closed per hour", where the emphasis is on things that aren't important because measuring customer happiness is pretty subjective and doesn't translate well to employee evaluation scores.
I know if I ever got the point where I would need to delegate support assistance to someone, you can be sure I would pay them amazingly well, at least equal to a developer's salary because I never want anyone to ever feel like they get ignored or have a low quality experience with my products.
"Metrics" are an unhappy middleground for everyone.
I've got no idea what the cause is. It's probably a mix of poor management, ineffective metrics, low salaries and a work culture in India which isn't synonymous with quality. The bottom line is, even the same company can be doing support wonderfully and terribly at the same time.
So if you end up hiring someone who produces a bunch of 1-2 star ratings, you just fire them and try again with someone else until you get someone who produces mostly 3s+.
In the end you would think this would result in better support but it never does. You just end up turning over a lot of support employees, and the user experience for the customer is still bad because you need to first get through a robot menu, then talk to an entry level support who will happily let you pour your soul out on the details of the problem for 5 minutes, and then at the end they are like "oh, I can't do that, but I can forward you to someone who can".
And now you get to be put on hold again, anticipating by coincidence you'll probably lose the connection, and if you're lucky now you get someone who is capable of understanding the problem and you get to re-tell the whole situation again.
Before you know it, with wait times included you're 20 minutes into this and you just barely got to the point where you might get help for the issue. Businesses could solve this problem, but they don't. Instead of hiring better support folks for more money, they put the burden on the customer to have to wade through a bunch of BS and essentially train their entry level support for free -- and it's worse than free too because you're paying that company money for their service and you're trading time from your life to do it.
I'd be the first to pay to get them explain to me some of their misterious weirdnesses.
This isn't so dissimilar from the method used in pathology to deliver consistency in results across multiple labs and assay methods. You identify some boundary cases. This is definitely normal, this is definitely cancer but this is borderline, repeat a few times. You replicate and get all the labs to mark your samples. Then you can identify labs (or assay techniques) that aren't reliably putting things in the right categories and demand they improve or stop.
A big part of this is how empowered you make your various levels. The SLAs I've always worked with, created, and worked under measured time on calls and customer satisfaction was always one of the more important measurements.
If a customer is willing to put up 10k to get their issue resolved, it's probably an issue worth resolving.
HN actually acts somewhat like a crowd-sourced ombudsman. People who have an issue write a description and post it to HN. If enough people find it compelling, it makes the front page. Once it makes the front page, someone in authority at the involved tech company will see it, and try their best to resolve it.
I am a tech executive in my company and my direct e-mail address is at the bottom of every Web site. Yes, this means I deal with routing all sorts of problems BUT I know instantly when there is a problem. Any problem. And I can actually help. But really it helps me with my job. It's win/win/win.
Plus, poor customer services is simply inexcusable -- you have to treat people the way you want to be treated. You're letting yourself and everyone else down otherwise. There is ethics and morals in business and they are important.
It doesn't have to be my way but there are definitely ways to do it, do it well and not break the bank.
But that was like 10 years ago. Not sure if it's still good.
Your solution basically boils down to "companies are failing at escalating support issues well, so they should escalate support issues well."
At least in the case of telecom companies they all seem to have executive response teams when the usual channels fail. Send an email to any of their Exec team members (CEO, president, CTO, etc) and it nearly always gets assigned to a special team that solves the problem. I can confirm this out of personal experience with at least AT&T, Comcast, and T-Mobile.
Basically the way the TIO works like this: 1. The customer must first go through normal support until they get stuck. 2. The TIO sends the Company a nice letter requesting that they get their act together. 3. The Company then has 10 days to fix the issue or the TIO will escalate. 4. Either the problem got solved or the TIO keeps nagging until it is.
Every time the TIO has to mail the company it also charges the company for having to get involved. level 1: $31 level 2: $260 level 3: $475 level 4: $2250
If an issue gets to the end then the TIO can direct companies to implement solutions costing up to $50,000.
This process heavily incentivises companies to actually provide functional support. I don't think your idea of charging the customer $20 to get help with support would achieve anything.
I wanted to point out something that may have been missed by many - which is that the OP in this case is a free user. i.e. he is NOT a "customer". He is a user of free services that Cloudflare is providing. It even says so in his post that he's a free account holder and that he's not entitled to tech support, hence why he kicked up a fuss online. My company is a Cloudflare Enterprise customer - when we write in, we get a response typically within an hour by their swat team.
The idea of an "Ombudsman" in this case wouldn't be an actual ombudsman in the spirit of advocating for the customer, but would be more of a "one time support fee" kind of deal. Which I think is fair, but in this case I'm not even sure the OP would have paid it. The OP had a viable option which would have been to upgrade to a Business account and get priority support.
I run a popular platform with 1.5 million registered users and the vast majority of them are free users. We also have the same problem. Most of our support is just swamped with free users mostly asking questions about how to do things (and it's in the documentation, they just don't read or search for it, it's too simple to just email support or blast away on Twitter and @ mention us). I've even had to withdraw from Twitter and Facebook entirely just because I was everyone's "ombudsman" (even if they weren't paying customers).
He registered his domain through Cloudflare Registrar, a paid service. I'd be concerned if my domain registrar told me to head to the forums for support for tld level issues that only the registrar can solve.
Also I think there’s an issue with what the expectations for support is for. The Cloudflare service itself - the DNS, DDOS protection, ssl, CDN, etc are all premium services with a free tier for kicking the tires. As far as I understood in the post, the OP was on that free tier but had registered the domain on Cloudflare (which was probably the added complication).
Sidebar - AWS offers no support even if you’re a paying customer - support is an additional product you have to add on and it is a percentage of your overall usage bill, even if you don’t contact support...
Well the blog post isn't really primarily about his specific issue, it's about the systems being severely lacking. A ticket won't resolve that issue.
So basically you give the ombudsman $20, and they keep the $20 if they fail to resolve the problem?
Pretty much every technology company already have "ombudsman" for their important clients.
> That "contact a friend in the business who can actually get things done" is a necessary part of a large support organisation and it simply does not exist yet in any tech company that I know of.
I can assure you that every tech company has these "friends" available for their most important clients/customers.
Whether it makes sense for companies to make these contacts available to every customer is another matter.
The machine / IP combination has been working for years and then one day it just stopped. No emails from ec2-abuse, nothing in my support dashboard. After troubleshooting for a while I bit the bullet and ordered business support.
After cycling through three support agents over 48 hours Amazon said that they were filtering the IP because of spam complaints but neither the Abuse department nor my support liaison was able to see the block - it has been applied at the network level and the network team did not communicate the changes to the other departments.
I told them that I would not be paying for a support plan (over $1000) to resolve a problem that I did not have any notifications or alerts or any way to troubleshoot for myself.
It required talking to two manager and almost a month before they finally agreed that the billing should be refunded. After lots of support tickets, emails, and phone calls I got the money back - but man... An ombudsman would have been amazing.
I know that people will think it's great that you are doing this and I also know that you think it's good (for you) to have a feel for the issues that frustrate every day users. But I think it's not a great use of a company execs time and I am not even sure it's a good way to deploy resources at Cloudflare.
The reason is people will tend to (as a rule) do as little as they can themselves but then use as a hammer the court of public opinion to get something resolved.
You say 'ever find yourself so frustrated with Cloudflare' but you know that in itself is different for different people. What will happen is you will get people using you as a help desk and then after you don't help them as quickly as they think you should they will then follow up with a post, comment or story about how you did nothing.
Separately if someone is posting publicly about an issue (as this person is) and if you can verify that it's actually coming from the customer (I mean who says it is actually?) I don't think you need them to say it's ok to resolve online. In fact to me it's the opposite. You take the time to reach out publicly and you take what follows good or bad even calling you out (the customer yes you can do that by the way) if you think they didn't put the appropriate effort into finding an answer.
Everyone optimizes for the worst case. They think “if I give out my email I’ll get tons of useless email”. I can assure you I get 10x the crap sent to me on LinkedIn than via direct emails from customers or others.
I deleted my Linkedin account well over a year ago and I still get emails from them saying my profile is being viewed. Tossers.
As a counter example, Jeff Bezos (whose time may be worth more than anybody else's) famously audits his email for customer complaints and occasionally derails an organization for a day or two in order to figure out what happened. He stands behind this practice and has said that he often picks out cases where the anecdotal complaint is counter to data that he's been presented, and that more often than not the anecdotes are correct and find a shortcoming in the data. IMO it also demonstrates a culture of caring and following up about anecdotes to others whose time is worth less than his own.
Is that why Amazon is growing more and more notorious for selling fraudulent items over the years?
Second, a bunch of honest questions:
Did you consult to your supervisor (or anyone with authority) to be able to bypass the support process (if there is any) like this? If so what was the response? If the response was negative, how did you convince people? After things resolve, can you kindly post how many spam or unrelated emails you receive so that it will be an example to the industry?
I'd like to put my skepticism on hold and blindly believe that your post is a reflection of pure concern and not just a PR stunt for damage control.
From my perspective, a developer from the company the article is about arrives in the comment section. I ask an honest question about the communication process internal to the company. I am told this developer is the CTO and is supervised by the CEO. Then I state the fact that I would appreciate if we also knew what the CEO thinks about this situation.
I don't literally expect answers from the people I addressed and that's OK, they are probably too busy working on the subject matter. If they take time and write one, that is really admirable and I'd be thankful.
What I don't understand is why it is considered inappropriate to be curious about these topics, as judged by the amount of downvotes I received.
Maybe the tone I intended for my post got lost in translation, since I'm not a native English speaker.
Cloudflare is a solid company with a good reputation and has done a lot for general welfare of the internet. The staff there generally care about customers and it seems like they're trying to figure out what went wrong here, and of course preventing negative PR is always a bonus.
There's really not much of a process when you're a top officer at a company, you have the freedom and responsibility to exercise your judgement and do whatever it takes for the good of the organization.
"I'd like to hear CEO's opinion" probably doesn't mean "I wonder what the CEO thinks about public communication by his/her colleagues inside internet forums". I'm reminded English is hard even when I think I have a basic understanding of common speech patterns (parlance?).
Your solution basically boils down to "companies are failing at escalated support issues well, so they should escalate support issues well."
For me and my of my customers, having your "entire cloud deleted" is like... the #1 nightmare scenario.
So why does this capability/function even exist for active accounts at CloudFlare? It sounds like the OP fell victim to what is essentially a regular process.
Or to put it another way: No amount of explanation or assurance is ever going to make me feel comfortable with my doctor having a handgun as one of his medical instruments.
Your doctor has a syringe as a medical instrument. They can kill if not used correctly.
This is the doctor seeing a patient that hasn't moved in a few hours and executing them:
https://community.cloudflare.com/t/does-cloudflare-automatic...
Oh wait, he was just sleeping? Ah, well... "oops".
There is exactly one reason you'd delete customer data without their active involvement: They're overdue on their bill and have ignored all warnings and communications. That's it. There are no other reasons.
Feel free to disable things if they're a security risk, but Do. Not. Delete. Customer. Data.
Ever.
Unless the mechanism exists for legal reasons, i.e. they are forced to take some sites down by the government or law suits.
It appeared to me as though they deleted my domain registration from their system. That's mostly why I freaked out. I wasn't sure if I'd even be able to transfer my registration easily.
However, on the advice of someone from HN, I found that if I went through the flow of linking an external domain to my cloudflare account the registration details came back.
I certainly don’t know what all the DNs entries on my 15 domains are, so if they disappear it’s not going to be easy to get it all back.
I hope for your sake that right now you are fixing that problem?
Whether this was a bug or a rare protective mechanisme, there will be times when your DNS provider makes a mistake and removes records. You mentioned in your post your DNS isn't hard to reproduce, but how certain are you that _all_ records are restored? How long do you have to fight DNS issues before it's OK?
I built DNS Spy [1] for this exact occasion. It monitors your DNS for any changes made, keeps a version of all DNS records (current & former) and allows you to restore/download a BIND9 zone file for your zone. You can easily import this into any commercial DNS provider or in your own BIND9/PowerDNS setup.
I would love to hear feedback on how DNS Spy could be improved when DNS disasters like these occur!
If you're a business, whether it's a SaaS or "just" a marketing website for your brick & mortar store, I think it's crucial to have back-ups. Most people think of backups as files, database dumps, previous versions, etc of their website. But the configuration data (in the form of DNS) isn't often considered.
You're tech-savvy and can restore your DNS records because you know yoru servers' IP address and your MX records, but who else could do the same?
1) You can't use it after the fact.
2) It's very specialized.
People are not going to set up dozens and dozens of services to monitor for really rare things. It should be part of general purpose monitoring suite.
A solution to this is to keep DNS under version control with eg Terraform and deployed by CI. master is then authoritative
But as a general reminder to everyone (I think this is an unfortunately common problem from a number of companies): If this is how your company handles account issues, you're probably wrong. Whether it's automated or manual, a user should be able to access all of their own information even when you decide to no longer provide them service. And you should test and retest the ability for people who you now deny service to transfer out.
I actively looked for someone we could pay money to, so we are their customer (as opposed to being a free tier user, effectively a cost)
The winner was DNSimple[1], who do exactly 1 thing, and they do it extremely well. And they are small enough to not take themselves too seriously[2], which I really appreciate.
Oh and their normal support channel is email, and everyone in the company takes a turn. I tested out their support before signing up and quickly heard back from a competent engineer, so they passed that test too.
[1] https://dnsimple.com [2] https://dnsimple.com/dnsound <— bonkers
I haven't needed to talk to them much, but one time I tried to add a .ninja domain, and there backend wouldn't handle it. I emailed them to report the problem at 4:49 p.m. I got an email at 7:09 p.m. the same day (2 hours 20 minutes later later) asking me to try adding it again. [1] When a free service fixes your problem in a few hours, they get +1 gold star from me.
[1] I just checked my email to look up the actual times. This was on Mar 15, 2017.
For example, you can run your own DNS server on a VPS or something, and HE will AXFR the zones from your VPS and serve them authoritatively.
This allows you to run a hidden master, for example, which I can imagine some HN folks being interested in.
Running my own dns looks more and more reasonable though.
Though I think the post would benefit from some citations to improve its relevance/usefulness otherwise it is little better than personal opinion/conjecture.
Unless you are specifically questioning the relevance of hosting spammers, on which case: If that is true (again, some examples would be helpful here) and you intend to host your own mail servers via their services not just the MX records pointing to other mail services, you could find yourself blocked by association at some point. False positives are a big problem in this area and can be much admin to clear up.
Came home, internet works fine. Everything looked just fine.
Back at work next day still can't connect. So I tried pinging, and I immediately see that the ip my home hostname resolves to is not what my ISP has. So I go to nslookup and try a DNS server I know (another local ISP), and it resolves to what I expected.
A bit of checking later I find that at work they've started using OpenDNS, and OpenDNS has blocked all of no-ip due to malware and spam.
So yeah, could be relevant.
I host all my personal stuff there, including something that updates their DNS via an API. They've been great to me.
Even now, like you, I get the occasional (ie. once per year) named segfaults or it randomly stops responding over TCP.
Few SaaS products are less effort than one reboot per year, but still worth it IMO.
Two years ago there was a moment where I was close to working for them too so I always try to use their products where I see fit. :)
CloudFlare is selling domains at cost. That means they are not making any money from being a domain registrar, which means they will do everything to keep the cost of doing it as low possible to themselves. This means lack of customer service and use of ML dragnets for "anomalous" behavior.
Sure, I could spend $lots to get a dedicated account rep from MarkMonitor or CSC but that's not really feasible for my personal site.
Are there really any registrars that hit a reasonable price point for individuals and offer service beyond bargain basement? Because if so I'm doing some transfers this weekend.
From previous research, at least, most domain registrars have ticket support at best. I did move all my "less important" domains to Cloudflare for cost savings recently, but they have my most important domains.
I'll speak for myself and say that all my domains have been with Hover for well over a decade now, and the times I've had to deal with their customer service, they've been excellent.
In fact, I even had to call them once, and I got a human almost immediately, and that human was able to resolve my issue while I was on the phone.. I don't recall the exact issue and I'm sure it wasn't anything major, but it was still nice.
So yeah, Hover. They're nice. And I think their prices are decent?
I tried several of the lower price registrar's back in the day, and they all sucked in their own way, despite me not needed anything except the thing to just stay registered.
One or two would change the price of their domain privacy, most renew the privacy for like 3 dollars and then send you the renewal email that your domain needs to be renewed, one of them used to charge me separately like 80 cents from some weird Canadian shell company...
I actually have a domain still with probably the biggest "cheap" provider, and they now have a thing where you are supposed to keep a deposit in your account to cover automatic renewals. Just charge my damn credit card guys, please.
So I'm saying namesilo all the way. Only one that hasn't ever pulled any shenanigans on me.
Not true. CF claims it is $8.03. Let’s say it’s $7.85 + .18 tax or something.
You don’t deal with the registry. You deal with a registrar. Just as they can charge a markup, they can give a discount.
Registration could be free. It would mean a loss of $7.85/$8.03 per year. Maybe they make it up by selling ads on your domain. Maybe they use it as a loss leader service.
There is no floor.
Even CF’s at cost pricing loses them money. It’s not free to maintain the services.
(1) https://news.ycombinator.com/item?id=21700139 - Sinkholed
(2) https://news.ycombinator.com/item?id=19322966 - I lost my domain and everything that goes with it
No different than this story where the author's DNS records were deleted because of so called "anomaly".
Here are so many more stories: https://news.ycombinator.com/item?id=21710939
DNS was a good idea but now there are organizations that have the power to arbitrarily take control and even remove your domain names and records. We really need to come up with a peer-to-peer solution and take back control of the naming system from these authorities.
Does anyone here have experience with running their own DNS servers for their domains?
[1]: https://ens.domains/
Of course, if the registry (i.e. the TLD) wants your domain gone, you are out of luck whatever you do. If this is a concern then you should pick a TLD with what you consider reasonable management. There are a lot of ccTLDs and gTLDs to choose from.
Therefore, what you absolutely shouldn’t do is to pick whatever domain registrar is either cheapest or largest, and pick whatever domain name which happens to look cool and be available. Both are recipies for potential disaster.
% dig ns yp.to +short
uz5jmyqz3gz2bhnuzg0rr0cml9u8pntyhn2jhtqn04yt3sm5h235c1.yp.to.
%
If your provider requires more than one server, just make something up, within reason, of course: % dig @tonic.to yp.to ns
…
;; AUTHORITY SECTION:
yp.to. 86400 IN NS uz5jmyqz3gz2bhnuzg0rr0cml9u8pntyhn2jhtqn04yt3sm5h235c1.ns.yp.to.
yp.to. 86400 IN NS uz5jmyqz3gz2bhnuzg0rr0cml9u8pntyhn2jhtqn04yt3sm5h235c1.yp.to.
;; ADDITIONAL SECTION:
uz5jmyqz3gz2bhnuzg0rr0cml9u8pntyhn2jhtqn04yt3sm5h235c1.yp.to. 86400 IN A 131.193.32.108
uz5jmyqz3gz2bhnuzg0rr0cml9u8pntyhn2jhtqn04yt3sm5h235c1.ns.yp.to. 86400 IN A 131.193.32.109The second story you posted is about a user who forgot to renew their domain and did not wish to pay the overly-inflated fee to re-register it while it was in the grace period.
I hold no love for any registrar that jacks up rates for getting back an expired domain and agree that they should have sent a reminder email, but describing this as someone "losing their domain through no fault of their own" is, frankly, incredibly misleading.
The user:
1) forgot to renew their domain 2) had full right to recover their domain but objected to the price 3) had full right to transfer the domain out to another registrar for the original 15EUR price and 4) eventually got back full control of the domain
Zero communication in my case as well.
https://community.cloudflare.com/search?q=127.0.0.1%20audit
Every related incident seems to be due to either nameservers temporarily/incidentally chanced away from CF (and CF's service not re-checking it perhaps) or the registration billing failing (which doesn't look to be the case since registration expires 2021[0]). The latest change to the domain was about a week ago[0], so if that was when it was transferred to CF, it might be the first scenario.
> Because Cloudflare deleted my domain registration I can't change the status from clientTransferProhibited through their dashboard so I don't think I can even leave.
Unless something else happened, deleting the zone from your account doesn't affect the registration. Re-adding the domain will instantly allow you to view the registration info and likely transfer away; this would only not work if the zone is banned for some reason.
"Your domain registration configuration depends on your DNS zone configuration" is a very strange way to do things.
> Every related incident seems to be due to either nameservers temporarily/incidentally chanced away from CF (and CF's service not re-checking it perhaps) or the registration billing failing (which doesn't look to be the case since registration expires 2021[0]).
The changes a week ago involves adding and deleting TXT and A records only. Cloudflare manages the nameservers I use as my registrar and I never changed them from the default. I just confirmed all of that in the Cloudflare audit log.
> Unless something else happened, deleting the zone from your account doesn't affect the registration. Re-adding the domain will instantly allow you to view the registration info and likely transfer away; this would only not work if the zone is banned for some reason.
Thank you so much! Trying that now.
btw here are the dates referred in whois for your domain:
Updated Date: 2020-02-24T20:32:34Z
Creation Date: 2015-06-09T15:34:37Z
Registry Expiry Date: 2021-06-09T15:34:37Z
looks like a change has been done even more recently.I'm into issues like this more and more, where you run into some strange behavior on a website and you wonder "How did this ever make it into production?", then you open the website in Chrome and the flows work fine. I worry that Firefox is becoming less and less viable.
Just had a similar case today: My Mom tried to order something online on her old Android tablet - and it didn't work. She blamed the tablet for it, saying "It's just too old, it doesn't work correctly anymore! I used to be able to order stuff on this website". I had to explain to her that her tablet is still working fine, it's just the website that is broken because it's not supporting her device (or browser) anymore. Shockingly, she listed quite a few websites, which she has used for years, which have stopped working for her in the past few months and years; all of these she mentioned as evidence that the problem must be her tablet - not the websites. When I opened two of the sites she mentioned, I wasn't too surprised to find very shiny, very modern single-page applications (with service workers registered and even WebAssembly used on one of them)..
So when you are creating a modern web app, please don't just test in Chrome on your new MacBook Pro. Think about your Mom. Ask yourself: "Is this still gonna work on her crappy old device?"
The manufacturer should be required to support it for the full lifetime of the device. Especially since your mom uses it to order stuff, which usually includes some pretty security sensitive information. I think you are putting the burden on the wrong party.
If it is somehow checking for support for SameSite, Secure, CSP or any of the other mechanisms that have been implemented in the last years then it might fail. Or they might be using mechanisms that work around the problem that those three are supposed to help since they are not available in older clients, but just don't have the resources to test the random android 4.12 version that you use. I think it should have a proper error message if that is the case.
But I feel like you are pointing the finger in the wrong direction. I try to build my apps without extraneous fads, but keeping a webapp secure (in other words keeping up to date with the latest protections) does not mean "submitting a form", and it does not mean letting any old client lacking the required protections through.
It also does not mean doing "WASM compiled redux reducers in ES6 module workers authenticating over JWT to send gRPC commands to a kafka broker talking with ingressrouting over anycast and a internal service mesh with mTLS3.9 auth using curve9999.9, token binding and Wireguard to secure internal communications over a VPC-less multi-cloud k8s cluster that uses Multi-Raft, Single-Paxos to have a single, distributed, disputably non-consistent CRDT-consensus algo over blockchain RS-232".
So, yeah, I'm not for fads over usability in tech. But I'm also not for supporting insecure clients just because the manufacturer of those clients doesn't give a shit.
Are you kidding me? If you're looking for shiny stuff to add to your resume, yeah, you can't possibly support those! If you're an HTML5 game developer, yeah, gotta use the latest and greatest. But if you're in the business of selling shoes, why do you need anything newer than Android 4.1 in order to process the transactions?!
A lot of basic security was missing back then, asking anyone to enter any security-sensitive info on those devices now is like telling someone to use WinXP and IE8 for online banking. Besides that the great variety of devices that shipped on versions v4.1-v4.4 and with their different hacks/customizations it is simply not feasible to ask developers to support them years after many of them cannot be had to even test a website on, let alone debug it.
You're very brave considering that Cloudflare doesn't even have U2F yet Google and Amazon do.
Cloudflare does support standard TOTP-based 2FA like most people use for Amazon and Google. So whether or not the lack of U2F support should matter depends on whether you actually use it elsewhere anyways.
Cloudflare's Domains service is new, and some of it's management tools are lacking, but I also moved most of my domains to it over the last year for cost savings. I'm thrilled with it, but I'm still keeping a few of my most critical domains with GoDaddy. (Hate them all you want, but GoDaddy hasn't screwed up my domains in well over a decade.)
You may be able to save a lot of money without risking your primary domain that you route email through.
Also for your core domains, do not let the registrar and dns provider be the same entity.
Also, don't decide on not migrating just because of one bad experience. None of them are perfect, though vigilance is wise.
(I know am probably preaching to the choir :) )
Consolidating domain registration, DNS hosting, and site hosting under the same account is a terrible idea and a big risk IMO. Always ask yourself “what if this account gets banned?” All of the big tech companies have automated systems that could turn you into an outlier with no support.
1. Setup up monitoring on your critical domains. UptimeRobot and Hetrixtools are good starters with generous free tier. You should know when your website/email/dns isn't working.
2. Don't tie your domain registration with your DNS provider. You lose everything if something goes wrong with your account.
3. Be able to jump ship easily, have backups of your zone, already know where you will transfer to.
Are there any open source status pages/monitor programs that have build-in checks for HTTPS, DNS records (ipv4/6), arbitrary port checks, etc? I'd rather just setup a status page/alert app on a $5 minimal DO/Vultr node and self-host/support/contribute to a FOSS program than use a commercial provider.
I'd probably recommend using one of the gateways (or a more fully-featured service like Pagerduty) for more serious businesses, but for personal use (or where an outage detected the next day isn't crippling), it's remarkably useful.
Remember that Linode/Vultr/etc don't run their own datacenters, they share datacenters and sometimes downtime events can exist outside of datacenters.
Nagios. Or its descendant with a better configuration language, Icinga2. They're fairly easy to do a minimal install and configure in a container or on a VM.
</opinion>
Maybe one fits what you are looking for.
https://github.com/skx/overseer/
Handles SSL-checks, DNS-checks, SMTP-checks, & etc. Runs a thousand-checks every two minutes for me, give or take. Pluggable output via a redis-queue.
Lesson learned :)
> Don't tie your domain registration with your DNS provider. You lose everything if something goes wrong with your account.
I don't see how that helps. How do I recover from my registrar deleting/disabling my account even if DNS is somewhere else? I think there's still only one failure point and the lesson is that I need to pay that failure point more money.
> Be able to jump ship easily, have backups of your zone,
Luckily I have that
> already know where you will transfer to.
Any suggestions? Ironically I recently moved from Google Domains to Cloudflare because I was worried about issues with opaque support. I've learned my lesson picking based on cost alone, but I'm a college student who can't afford something too heavy-duty.
Your outage was a DNS outage, not a registrar outage. If you still had control of the domain you could update your name servers to another provider, import your backed up records and get the site back online without talking to CloudFlare.
I believe it was both.
If I have a registrar outage I'm hosed. If I don't have a registrar outage and do have a DNS outage I can recover with a little work. But in the only case I can recover my registrar was reliable, so why didn't I just have them do DNS as well?
Because they have just proved being uncapable of doing it? Because redundancy? Because you shouldn't keep all your eggs in the same basket.
I've been self hosting for at least 15 years and did not have any huge problems like the domain becoming non resolvable. I would never host my DNS on my registrar's infrastructure. It's being sloppy and lazy and it gets you embarassed.
If you haven't already, you might consider checking the CloudTrail logs for the account in question to see if there were any API commands related to the zone.
From reading the linked helpdoc, apparently your entire domain can get removed from cloudflare if your register stops reporting cloudflare's servers for the ns records.
The mere idea of having to re-enter all hundred or so of our dns records using cloudflares 1.2 second delay at every step of the way add dns record interface because namecheap bugged out for a few seconds is horrifying.
There is no way to export all of these, there is no way to import or mass add and i don't think they can lean on the api to save them here.
Dns records are data, dns records are sometimes important unbacked up customer data. Cloudflare does not offer a way for customers to back this data up, nor a way to restore or recover from a backup but it acts very callous with this data, deleting it in automated systems based on data from 3rd party providers.
Not a good look.
Anywho, the help article does not state it it "takes days", so it does not have to "take days". Otherwise known as listening to the documentation not the observed behavior. Programming 101 here.
So much business focus goes into the onboarding experience, and since you assume all of the people your service terminates are "probably bad people anyways", not a lot of thought goes into offboarding, or ideally, appeals.
I didn't give much thought to it, as I wasn't using CF for anything in production at the time, but sad to see that it also seems to happen to other people.
For instance, a while back I forgot to renew one of my side project domains so it briefly expired for maybe a day or two. Got this email from Cloudflare saying
> Your DNS records will be completely removed from our system in 7 days.
> ...
> Once you have completed this change, click the “Recheck Nameservers” button in your Cloudflare dashboard to ensure your domain stays active on Cloudflare.
I promptly renewed, except there's no "Recheck Nameservers" button anywhere, and the dashboard still read "Moved" for maybe a day. Eventually the problem was just gone, but the communication worried me that entire time.
(I do appreciate Cloudflare's service, though.)
This sounds like a plot of a japanese horror movie.
My website was down for yak shaving when this happened, but before then I had DDOS protection turned off.
I get it - these are free services. You should factor that into every decision. But the risk is real even if you pay for an account. I’ve been slowly moving away from Gmail to a custom domain, but something like loosing DNS records and not being able to restore them quickly is even worse.
Back up everything that can be backed up, don’t rely on a single provider and always have a continuity plan!
I don't get the point of 'shoot first ask questions later' type approach. Obviously it would pay to get some kind of affirmative reply from Cloudflare prior to a post which everyone here with incomplete information speculates and wastes time on (like I am doing).
Also Cloudflare did not 'delete my (the) domain. It deleted the dns records. There is a difference and no I am not being pedantic either. How would 'the internet' know why this was done there could be any number of good or bad reasons.
Lastly the domain is not expired and as such the registrar is required (per ICANN) to supply an auth code so someone can transfer out. Or to allow the customer to change the primary and secondary dns to another dns provider. There is zero (legitimately) that allows cloudflare as either a dns provider or a registrar to lock the domain up pretty much (other than for a legal court order) just for some reason they might decide to do that.
> Also Cloudflare did not 'delete my (the) domain. It deleted the dns records. There is a difference and no I am not being pedantic either.
Thanks. You're absolutely right. I meant delete their record of the domain as it shows up in the UI of their dashboard.
> How would 'the internet' know why this was done there could be any number of good or bad reasons.
For many reasons luckily HN isn't 'the internet'. I've already gotten some good suggestions.
> Lastly the domain is not expired and as such the registrar is required (per ICANN) to supply an auth code so someone can transfer out. Or to allow the customer to change the primary and secondary dns to another dns provider. There is zero (legitimately) that allows cloudflare as either a dns provider or a registrar to lock the domain up pretty much (other than for a legal court order) just for some reason they might decide to do that.
I know. Again, I guess I was insufficiently specific. Cloudflare has warned me to expect long wait times before I can talk to a customer support rep. My question was if there's a way to transfer out without needing to wait on a slow support loop.
At first I thought you were talking about Cloudflare shooting first, but apparently not.
Now we can track, record, and audit that service provider's promises...
That way, the service provider can't use an all-too-easy excuse like "we can't find that record in our database -- so you must not have ever created one..."
An assumption of false-payment led to them suspending "300-500" accounts (mostly UK based). I am still of the opinion something far more sinister is at play... and this doesn't comfort me.
That's when I moved the couple of handfuls of domains I had left at GoDaddy over to Hover. It's more expensive, but the Hover interface is better, and I trust Hover (Tucows) more (well, I trust GoDaddy less).
I don't trust Cloudflare one bit, and I think everyone should question whether their attempt to re-centralize everything is beneficial to the planet.
There are two major problems here: one, the problem itself, which is the deletion of DNS for apparently no good reason, and two, which is the bigger problem, is that it's incredibly difficult to talk to a human about what happened, so there's no assurance it won't happen again.
If people want things to be reliable, we've got to stop using companies with which we cannot communicate.
As a self-hosting enthusiast, something like Cloudflare is one of the best chances of having a plan that competes with "just hosting it in the cloud".
Cloudflare will only help once your server has gone down with their "Always On" thingy, if you have that enabled. They don't cache HTML by default.
I'm talking about their rather political move to re-centralize DNS by shoehorning themselves in to Firefox via DoH, for instance. Their unwillingness to be transparent makes this all the more frightening. Add to that their blatant desire to make money at the cost of doing the right thing (and I'm talking about unambiguous things - is someone going to argue that freedom of speech allows people run a phishing site of your bank?), and you've got a scenario where once they reach critical mass, they will be exercising their position to the detriment of everyone who isn't paying them, similarly to how Gmail, through doing and not communicating, say "screw you" to many small email services.
When people who don't use large providers have email issues with Gmail, lots of people have knee-jerk reactions saying that everything should move to the big providers, that people and small businesses should not host their own email, and so on. This is NOT the way the Internet should work, and we should never allow Gmail to just arbitrarily do whatever they want, then accept it as the new normal.
If you have more than a dozen megabits of outgoing bandwidth, you can easily host a blog from your home network which can handle a front paging here. Just don't expect to dynamically generate a new copy of the site for every visitor, and if your bandwidth is tight, then host your images on a static server off of your network. Cloudflare is not necessary - perhaps it's easier, but it isn't necessarily best to blindly trust a company that wants to become a monopoly.
IPFS can work as a CDN, at least for static content that users are willing to seed. This is especially relevant to the "blog post hits frontpage on HN" case. Of course, dynamic content is not quite as easy.
good idea
i do this for all the domains i use/manage
this post has been a good reminder to check them :)
imho about audit log: since they "delete" everything, nothing is left in the zone/domain.
thus, initial log (127.0.0.1/creation) comes up. kind of feature of the bug/logic error.
imagine if one day your bank decided to close your entire bank account without telling you...lol.
When/why did they remove that one? Have you got any source?
You don't now.
Why do successful technology companies have a right to have a proportionally large influence on the public political debate? Is it good for society to allow successful technology companies to have such a large degree of control over something so incredibly vital, merely because they were effective at running a particular type of business?
I'd switch to an ISP that promised they'd take no money from Nazis
Cloudflare also didn't even bother telling that lie anymore when the dozens of sites they censored afterwards including 8chan were systematically barred from basic commerce.