An update from Linode about the recent DDoS attacks
status.linode.com
status.linode.com
When our customers are emailing and tweeting us and they just want to know when we are going to be up, and all we can say is "We have no idea, we don't know why this is happening or what's really going on", that's pretty much the definition of a worst case scenario from a customer service standpoint.
As someone whose business relies on Linode currently to function, I am sympathetic to Linode's plight...this is the equivalent of someone coming and setting off a bomb in your factory; not exactly something that you can always plan for even if you have prevention measures in place. But they would have kept a lot more of my sympathy long-term if they would have communicated better with their customers in the first place...
EDIT: And it looks like the attackers decided to start things back up again, as Linode.com is unavailable...
The timing of the DDoS was pretty interesting too, happening when not everyone is available.
> And it looks like the attackers decided to start things back up again, as Linode.com is unavailable...
They're watching our status page for updates and starting new attacks when we resolve previous ones. There's been an almost 1:1 correlation lately.
This always rings hollow to me, and yet I hear it over and over.
A company like Linode surely has at least a dozen people who can be on call in a situation like this, probably much more. All it takes is for one engineer or even a product person... Heck a technically-minded support person could listen in on the war room meetings and get enough information to post something better than "we're fixing it".
5 minutes of blogging every six hours would be plenty.
And yet, people always claim it's impossible and there's no time. Frankly I find it frightening... If you are coding that fast that no one on your team has five minutes to step aside, take a breather after six hours of coding and summarize what the team just spent the last six hours doing, I shudder to think what kind of panicked alarmist interventions you are making.
There's no excuse for silence. There just isn't. It's a gross failure on the part of management to prioritize the responsibilities they have to your customers customers. Full stop. Sure, the engineers can't be expected to remember to tap out to blog. But if that's the extent of the accountability structures you can assemble during a crisis, that is a serious organizational failing, particularly for an organization the size of Linode.
Perhaps its time to consider some failover at another host. Same goes for anyone solely dependent on anyone.
CloudFlare doesn't offer VPS instances. Google and Amazon are in a league of their own and, last I checked, don't offer the similar services at the same prices with the same level of support and same controls. Amazon might send you a nice bill at the end of the attack depending on what type it was.
DigitalOcean, as an example, routinely receives DDoS attacks and will just nullroute your VPS automatically until the attack ends.
What happens with 1/10 second's database records not replicated? That could be a successful transaction on one, and no record of it on the other?
You can use a multi-master setup and strict consistency checks, though.
Say I have one server setup on Linode London and one duplicated on DO Amsterdam. How do I quickly re-route the traffic from my two main load balancers London to Amsterdam without much propagation time?
* This is for an answer for the average "Hey I run a website or two" question. If you're asking and you are dealing with financial transactions or ad networks, well, you're asking in the wrong place and you should probably hire someone that knows how BGP anycasting works.
Eh, usually you just bgp anycast the DNS server. If you anycast just your DNS server, you are pretty much in the same pickle as if you host your dns with someone else who runs a reliable dns server (and have them do the anycast) - the pickle is that there's a minimum ttl.
Now you can set up a redirector server and anycast that, and like 3xx redirect all your requests (or more often just the first request) and that's better, really, 'cause that redirect can be changed quickly and on a per-request (or more often per-session) basis... but while I know a thousand companies that will sell you anycast DNS service, I don't know anyone who will sell someone with a vps customer budget an anycast'd redirector service.
Because they physically route the traffic, you can change backends pretty quickly and have traffic going to the new destination within seconds or minutes. Using DNS for end-users just has too much unpredictable latency, but it works well for routing the backend.
Around the holidays, network engineers are the only ones who don't really take time off. The sorts of people who might say "hey, we need to give some clarity to customers" are less available than the people whose time is spent firefighting.
We've had a few false-starts, but we've finally gotten all of our datacenters dropping all traffic to these critical IPs at the edges. This should stop the most serious outages, so it's a big step.
I do feel a certain amount of loyalty to linode, as they've been an excellent service the past few years, and I see them battling to put out these fires.
I'm hoping as linode grows in size they can put in place more sophisticated measures to guard against this - DDoS is a problem everyone faces. Its tough, but when the dust settles, it could be an opportunity to innovate.
1 - Attack mitigation was mostly successful. As I thought and they have confirmed, the attack vectors evolved continuously.
2 - They had to deal with this over Xmas. Anyone familiar with such a job knows what this means in terms of human resources, knowledge distribution, organization of technical response and communication with 3rd parties.
3 - Linode is not Nagios. If you don't monitor your own infrastructure don't expect Linode to SMS you because your site might be down. Linode resources were focused on fighting the DDoS, as they should, and provided regular updates through their status site, as is expected. Everything else is nice-to-have, but no a must-have.
4 - In line with what others said, I had 7 hours downtime in my London VPS. That is an uptime of 96% in the last 7 days. Considering restless DDoS ongoing over holidays, I'd say that is pretty good.
I'm sorry, but what happens to Linode sucks, but it is an eventuality anyone with assets depending on this service should have counted with, because it can happen everywhere. Cannot blame Linode if your HA strategy does not exist, or you never thought of a way to gracefully fail over to a second provider if your business depends on >96% availability.
2) Netops people don't get Christmas off (1). Our teams are working all year, all day, every day. Netops isn't HR or marketing. It's not super difficult to arrange a conference call with your transit providers, DDoS mitigators, etc. This is old hat to them.
3) Completely agree.
4) The downtime was a lot worse if you had multiple instances in more datacenters. Or put another way, the bigger the Linode customer you were, the worse it was.
(1) I've had a router die every Christmas or New Years for the past 6 years. Never had one die outside those windows. They must angry at me for what I've put them through.
Fuck you. I will continue to be a Linode customer. Not sure what your goals might be but you will not succeed.
Frankly, and I am going to be politically incorrect here, these are the kinds of cases where I wish there was a "special forces" kind of task force to hunt down these pieces of shit and put them out of their misery.
This amounts to financial terrorism of the worst kind. It affects small and large companies and creates untold losses across the board. It is entirely unproductive. The world would be a better place if the pieces of shit who engage in this sort of financial terrorism simply didn't exist.
Happy New Year.
Linode folks: I'm renting another server next week. Don't need it. Just want to support your effort and, in a tiny way, help mitigate losses. I might just give it to the kids in the robotic team I mentor so they can play around in a real server environment.
Oh, it's certainly in the capability of the FBI and the NSA to hunt them down. Even child fuckers inside the Darkweb got busted.
The problem is priority: unless either child porn or a huge US company is involved, the three-letter-agencies don't give a shit about this kind of crime.
This is a problem that has a solution, but the people in the best position to handle it have shown no interest in doing so for a decade.
ISPs can't be expected to solve this problem: the ones that have tried discover that it takes huge efforts to get a hacked customer cleaned up, and then they frequently just get re-infected. ISPs exist in a highly price-sensitive and competitive environment: providing dedicated, personalised security services to customers who don't even care or can't be blamed (e.g. hacked router) is incompatible with charging a competitive price.
I don't know what the solution is. Just saying "no" next time someone wants to use crap programming languages and frameworks that systematically lead to exploits would be a good start. No more using C in embedded devices that can't be remotely upgraded.
We need to evolve the internet past these types of attacks and build systems to protect everyone from this kind of terrorism, rather than just get angry at people who abuse well known about loopholes.
Your anger, while justified, is not helpful.
Example: US government and all necessarily parties start to shut down --as in kill off-- all packets coming in from China every time there's a large DDoS attack. If it stops, we know the source. And we treat them to a full day of packet blocking to see if the Chinese government might become inspired enough to do something about it. Next time, it's two days. Or maybe a week.
This is the kind of thing where pain has to be felt at a high enough level for those who can do something about it to intercede and put an end to it.
And, yes, in the end, it's a technological solution. I won't claim to be informed enough to fully understand what has to be changed in terms of protocols, hardware or topology in order to make DDoS a thing of the past. I simply haven't spent any time studying the problem. I'm sure there are lots of people very well versed in the subject who can pitch in.
At some level though, I think this has to be criminalized and prosecuted with great consequences for the perpetrators in order to put an end to it. I say this because it is likely that whatever new technology is instituted there will be holes bad actors can exploit.
Know that you're in good company here and that we're rooting for you.
I just want to say that I switched this summer to Linode from DO for my personal sites, and I've been very happy with your services. I plan on launching a few more things in the coming year, and these attacks have galvanized my resolve to continue supporting you through all of this. Thanks for everything you guys are doing, and keep up the good work!
I know it's unlikely to ever happen given the nature of these attacks, but I really do hope the perpetrators are found, brought to justice, and locked away for a long, long, long time.
Tell them that. Then say what happened to your business and customers was the digital version of it. They should get the idea.
What upsets me the most is that customers haven't heard a single word from Linode. If you weren't watching their status page (and have server monitoring on your servers), you'd be clueless.
The least they could do is to email affected customers about what's happening, and the time frame they need to fix it.
But one week of continuous issues is just not good enough.
Happy new year to everyone. And looking forward to a great 2016 at Linode.
2) Your transit providers should be defending their infrastructure. I've never seen a transit provider allow an attacker to take out their /30 serials or IX addresses. This is their network after all. If attackers try to hit the serial between customer and provider, you just readdress the serial to RFC1918 space. You don't really need a routable address there other than to make traceroutes easy to read. If they attack farther upstream in the provider's network, you just add ACL's at the provider edge. Nothing external will ever need to reach a provider's core. This is basic, basic stuff.
Next time, don't only run your network on house bandwidth (HE, TelX, etc). Or in other words, caveat emptor.
People dont need to jump ship they need plans in place to deal with problems like this, even just a rsync.net account.
There's no intelligent discussion left here. It's devolved into a sympathy vote. "You said something mean about Linode. I love Linode. Downvotes for you."
If Linode had posted the update to NANOG, it would have been more productive. I don't often say that.
Or are you referring to xconnects inside their network? That's up to them to work out and I've never seen a provider just abandon their network while under attack.
I wonder what size investment this is taking, and what the end-game is for the bad actor. Unless Linode's mitigation tactics are increasing the bad actor's costs, what's to stop the bad actor from continuing the attacks until Linode goes out of business?
I imagine there are a few thousand linode fans who'd be happy to help fight back.
I know I migth come late, but being one of your very satisfied Customers, and having experienced this type of issue multiple times with other providers that didn't even bothered to even ackowledge that there was an issue, I can say that I wil remain with you regardless.
Also to those saying "I have all my business running at Linode so this is unacceptable", I only says this: You get what you pay for, and for a VPS Service you won't find better than Linode, and if you have something critical running for YOUR Clients, than it is YOUR responsability to ensure resiliency against this type of situation. Linode is a VPS provider after all, and the reason why you are making money out of someone who doesn't know enougth to go to the VPS hoster themselfes.
Good luck making a profitable business and milking your Customers running on AWS or AZURE. You'd be broke and in debth at the first DDOS and Over-Bandwith charge from any of them.
I work at a service provider myself, and I understand what you guys had to deal with the last 10 days, and you have my full support.
It seems a bit unfair to have this fall on Alex's shoulders. I could be way off base, happy to be put right. I'm sorry to hear about your ruined holidays. Hopefully you'll get some time off soon :)
And yet, we still get DDoS attacks. Why?
Edit. This is in no way a criticism of linode. The worst outcome is if we all end up with one monopoly supplier. I have deliberately avoided using the big player in this space as I want support diversity. This makes my job harder, but it is better for us all if we don't put all our eggs in the one basket.
We have much the same stack at Reverb.com and I focus more on creating a disaster recovery plan than tackling a mirror of our infra. I'm focusing on bringing up our stack in another region of AWS if we would ever need to rather than mirroring everything.
Sometimes a mitigation plan is far more useful than building in excessively complicated redundancy.
Are you suggesting to have a slave ready to become master in the 2nd datacenter? If that's the case, my main area of uncertainty is: if you promote a slave to master in a 2nd datacenter and that new master accepts writes, you then have a brutal problem of propagating writes to the old master when he comes back online.
The whole thing may exceed my capability right now but I think having a disaster recovery is far better than a mirror now that I think through it.
I'm not saying something off-the-shelf will work or out of the box. Just seems like there would be a clear path for admins or developers of such software to integrate it with... something [1]... that did that.
[1] Example: http://sector.sourceforge.net/
Unless you're a bank, you run asynchronous with watchdog scripts and accept the very very small risk that you lose a few transactions within the ~100ms window of async latency.
It's all risk analysis. For most of us, hours of downtime is riskier than losing a tiny sliver of data for a small number of customers.
My system is dynamic, but once a user is allocated to a node it is easy to replicate their data over the network. The key was I designed the system with easy replication in mind - flat files rather than databases, etc.
https://news.ycombinator.com/item?id=10822904
EDIT to add: Good work on Unison btw.
I think the real key to creating a simple and low cost reliable system is to have this in mind when you design your application. I could have used a database just as easily as a flat file approach, but this would have made things much more complicated to replicate between the nodes.
"I think the real key to creating a simple and low cost reliable system is to have this in mind when you design your application. I could have used a database just as easily as a flat file approach, but this would have made things much more complicated to replicate between the nodes."
...is a great point. It echos the comments and thinking of Bernstein in his paper on lessons learned from Qmail.
http://cr.yp.to/qmail/qmailsec-20071101.pdf
Particularly, section 4.5 "Reusing the filesystem."
It doesn't always have to be complicated. Chances are It'll probably feel more like grunt work than rocket science.
Look up Disaster Recovery Metrics. You set an RTO (recovery time objective -- how long it takes to recover) and RPO (how much history is lost when recovering) based on what's needed for your project and feasible to implement.
Feasibility of disaster recovery strategy = Cost to implement strategy + Cost of losing data back to the strategy RPO + Cost of losing business during strategy RTO < Cost to business of riding out probable disaster
Pick the fastest/cheapest feasible strategy, implement it. Then once your ass is covered, pick the one with the greatest expected value and implement that. If you wait until you can design and are given resources to implement a 5-nines available strategy with sub-milisecond RTO and RPO (which seems to be a popular tendency of developers), disaster will strike before it happens.
And, y'know, because it would be really cool if I managed to ;-)
1. IBM mainframes.
2. HP NonStop systems.
3. OpenVMS clusters on SMP machines.
Give them the prices of systems and rare labor for each. Then, show them the price of a three to four 9's solution with Linux that has three to four less 0's. Mention how many extra bonuses, err business investments, they could make with such savings. At this point, they might treat you like an idiot for not buying the cheaper route. Maybe even enthusiastic about your business savvy approach. ;)
If you're person B (under attack) it's pretty difficult to track through all of that to person C. You'd need a lot of cooperation from people (likely in many different countries) who really just want to go back to their normal business. They're likely also charging for the traffic, so they're not really that bothered, and they're each only seeing a small proportion of what person B is seeing so they don't see it as much of a problem (so aren't likely to be inclined to get involved).
It says here that it's roughly 10% of machines. It's safe to assume at least half are infected. But aren't they mostly third world?
https://www.netmarketshare.com/operating-system-market-share...