I won free load testing
fasterthanli.me
fasterthanli.me
Well, that’s nice and all, but if a fly.io customer were attacked with 3.1GB/s throughput, according to the lowest outbound bandwidth price of $0.02/GB [1] they’d be burning at least $3.72/min. 6 times that if attacked from India. That would be a lot less fun.
[1] https://fly.io/docs/about/pricing/
Edit: They mentioned they waive charges as a result of attacks: https://community.fly.io/t/about-rate-limiting/156/4
For the time being, I've decided that getting insights into "how well the platform worked for me" was more valuable for me, so my video platform is staying there, but I've initiated multiple discussions about pricing and I intend to keep doing so until I'm happy with the answer.
That said, bandwidth does cost money. There are three ways we could handle it:
1. Charge as little as we can get away with, be transparent about it, eat the cost when someone has a negative experience.
2. Charge for bandwidth capacity, but don't meter it (ie: give VMs unmetered 100mb interfaces).
2. Don't charge for it, call it unlimited, put a hidden cap in place, and restrict what kinds of apps can run on the platform.
We opted for #1. I could give you a lot of post hoc reasoning, but the reality is that it's what I'd prefer as a customer. Companies that promise "unlimited bandwidth" feel a little slimy to me. Amos wouldn't be able to run his video hosting on one of those platforms.
Unmetered interfaces with restricted throughput do seem pretty nice. I've used a bunch of services like that. The cost to get started is high, though. And in my experience, the quality is poor. I've never had worse network performance than when I was paying for an unmetered network connection. Which makes sense, because these people attract all the users who want cheap, unmetered bandwidth. And they can't really afford to build enough upstream network capacity to handle all of them.
I don't love surprise expenses any more than you do. I do think we picked the least bad option, though, even though it puts some people off.
Of course, if you get hit with something slightly more targeted, this defense is worthless.
As a treat, this is a testament to a logic error I made in the caching code (inserted uncachable versions of pages into the cache for a little while). Enjoy!
My twitter timeline also has tons of neat little crates pop up every month, and I have to thank my colleagues at Netlify (back when they were doing Rust) and now at fly, for showing me some of those I showcase in the article.
Also, Contabo Asia (AS141995) is misclassified as in Germany. Although they are German, that AS is exclusively used for their Singaporean operations.
S-tier implementations include: firewall rules or a BPF program, or a VPN-based approach (like Cloudflare Tunnel). The way I did it is fine for small-scale attacks like that one, but a large enough attack will have you spend too much time on syscalls and waste valuable kernel resources.
I'd love to read a write-up about how these different approaches perform in practice, because this is largely gut feeling / the popular wisdom that "the sooner you block, the better".
I wouldn't expect this to be all that significant if already using TLS. So in the context of only allowing cloudflare to establish connections to the service, and be 100% fronted by cloudflare, mTLS isn't much more expensive than already using TLS, which we all do every day. And my read of the blog post, is most of the expense / optimizations are in the time it takes the server to generate and serve a page.
> Especially when a server is potentially under duress?
Well, it's a tradeoff. The desired property is only cloudflare can make requests to the actual web server, so compared to IP whitelisting there might be some tradeoffs. I think the biggest one is maintenance, IP whitelisting requires staying ontop of any changes from cloudflare. I'm sure cloudflare is good at pre-announcing new IPs, PoPs, etc, but there is a small risk to missing this, especially if it's a set it and forget approach. Although to be fair, mTLS also has this depending on when the root is set to expire.
> Can you cache the verification so it wouldn't have to be done each time?
Sort of, but it depends on some support on both the client and server, so it depends on whether both sides have opted into it. The terms you're looking for to do more research are TLS session resumption, which does involve it's own security concerns.
More effective at this level, is clients will re-use connections. So as long as the client/server support Connection: keep-alive, you do the mTLS once and run hundreds or thousands of queries in the single TLS verification.
My 2 cents is the biggest downside of the mTLS approach is you may still leak the location of the webserver. If the web server is on the same domain or a domain known to be used by the target, some certificate scans may turn it up. IIRC server cert is exchanged before client certificate, so it's possible to leak this without presenting a client certificate. I'd have to double check the spec to be sure though.
Although to be fair, this problem also exists in the specific implementation of whitelisting from the OP, as it appears to be code embedded in the web server to drop the connection if not from an approved source. So someone scanning for certs / servers could still learn about the server from the TLS connection, and do a simple flood the server out of existence type attack, or some more limited resource type attacks on number of active connections, etc. Ideally, with this type of whitelisting you would want to do it in iptables or the host firewall, so that you just blackhole the unapproved IPs and reveal nothing about the existence of the server.
- I didn't look at the presented source code close enough, but the code in the article does appear to update the cloudflare IPs in a loop. So in the presented code it's not an issue to get out of date, but for anyone replicating this, would be something to consider. - Also of note, there is a reason security minded folks don't like using the IP address for identity, as the whitelist likely covers far more than necessary, and is treated with less scrutiny than something like a TLS certificate. For a personal blog though I don't think this is really a concern. But might be a consideration for a company with something to protect.
I tend to be paranoid about exposing things to the Internet, so just put my raw servers behind Envoy. I have tuned that to do rate limiting, circuit breaking (stop sending requests to an upstream when it returns too many errors), idle connection termination, and to shed load when a certain amount of memory is in use, so without any additional configuration for a new service behind the proxy it's somewhat difficult to get the proxy and other services to not respond at all.
I'm guessing that in a real attack, the rate limiting service is a weak link. I use a custom rate limit service to aggregate rate limits across a /24 (and hacked that together in an evening), and that is likely the first thing to blow up and erroneously deny service to legitimate users. (I'm sure I have it set up to fail closed, which will be annoying.)
I had a hard time ever generating enough load to test any of this for the static serving path. I just set up a mirror of my production environment on my workstation, limited the critical services (Envoy + nginx + rate limit + Redis) to some low amount of CPUs, and then had 31 workers generate synthetic load. I was able to get circuit breakers to open to at least prove that that code works, but I somehow think that I'll run out of network bandwidth before I run out of memory to keep track of open streams. Difficult to load test when the upstream can respond to most requests out of memory.
Would be interesting to dig into it more. But for those of you reading this and thinking "I'm going to launch an attack right now", I will just turn off the site if I go over my bandwidth quota. Clone the config repo, host everything locally, run your tests, and send me the results ;)
I helped develop Tunnel over the last ~3 years, and I love it, but it's definitely overkill if you just want to serve some static files on the edge.
(Please delete if this goes against HN policy, I'm trying not to be a shill here, just help tristor avoid spending an hour configuring page rules and cache policies)
https://en.wikipedia.org/wiki/Intrusion_Countermeasures_Elec...
Site Reliability Engineering is a fascinating problem space.
if let Some(net) = ip_nets.load()
.iter()
.find(|net| net.contains(&addr.ip()))
ip_nets is a 'HashSet<IpNet>' but it should be a radix/patricia tree.Something like https://lib.rs/crates/iprange
(Keep in mind this happened during the attack, so compromises)
Some of them are very intentional: https://www.torproject.org/
> Yes, yes, I know, I should add anchor links for headers.
This is a rare case where blind people using screen readers have it (a little) easier. Every serious screen reader I know of has a command to skip to the next heading. It's too bad most sighted web users don't have a similar feature handy.
For our business use case WordPress is a necessity and so switching to a true static solution simple isn't feasible (we acquire and merge other content sites. 99% of content sites being sold are built on WordPress so being in the ecosystem is critical for this reason and many others)
Very interesting... I also use NNW and FreshRSS
Beatrice [0] - A web server with built-in connection limits and thread limits. It's async but supports non-async request handlers.
fair-rate-limiter [1] - In theory, one could use this to shed most of the load from DDoS attacking nodes.
Setting aside the fact that headless chrome or other browser testbeds do a good job at hiding their presence, what could be the vector for a botnet infection if this were true? Extensions?
Could even be just a user with a few beefy machines and a lot of proxies.
It doesn't look like headless chrome, or headed chrome, or any kind of chrome... because there's just requests for one file and no other resources. Chrome would do a lot of other stuff.
This is a wget loop or some other minimal agent with a changed user agent string.
Noob question: was that caused by the article getting posted to HN? Or was it really an attack?
And humor:
>> Because it doesn't return an AddrStream but instead a Pin<Box<TimeoutWriter<TimeoutReader<TcpStream>>>>...
>> Gesundheit.
I got a fly app and can't proxy it through cloudflare cause that doesn't work.
But I created a cert, added the a and aaaa records to CF, then after it was verified, turned back on the CF proxy.
However, my app stopped working until I turned off the proxy
And realistically, there's only so much they can do about someone running Tor exit nodes / an open proxy on their infra. Everyone in the cloud space has been fighting that off (and miners) for years, it's one arms race among many.
I think what is happening here is that there are lots of free hosts that let you send traffic to websites (in my case, volume didn't matter, just people signing up for free trials to get a little bit of free compute), and there is really no way for cloud providers to reduce the volume in a meaningful way. They are not necessarily serving malicious customers, rather their legitimate customers have gotten hacked and are now the attack vector. Or, their business is hosting, and people using THEIR free trials are using the free trials for abuse. (Consider if you just want a new IP address with which to sign up for some web service how easy it is to use something like the CircleCI "free for open source" plan to do that.)
If I ever started my own cloud provider, one thing I'd want to get under control is a good view of traffic leaving the cloud provider. Probably more than ports + bits per second; actually proxy the HTTPS or whatever. That way, if someone starts abusing other people's stuff, there is at least a point where I can rate limit it ("kill all video downloads to notable Rust personality's website because they asked me to") while hacked customers get their stuff cleaned up. This is a hard problem, balancing security and good Internet citizenship, but something I'd want to spend some time on.
Anyway, TL;DR, DigitalOcean shows up on your radar because they are pretty popular. Lots of Linux VPSes equals lots of insecure Linux VPSes, which is the perfect point for launching another attack. There is only so much the cloud provider can do, but doing more would certainly be nice.
If this topic gets flagged then we’ll know if it was go specific I suppose.
EDIT: Why am I flagged?
We all have our own particular axes to grind but even me, someone who is language agnostic, is getting tired of the tone of those submissions.
The problem is that he’s writing interesting, informed commentary on topics that are hot button issues for people who are a tad over-sensitive about their career choices :)
You can be a happy Go programmer, whilst also recognising that the language has limitations & things it’s really not good at. That would be the mature, honest response to these articles. What is not a mature, honest reponse is DDOSing the guy because you (the generic you, not you specifically hnlmorg!) don’t like his opinions.
The problem is the author preemptively shrugged off those responses by ostensibly accusing those coders of having Stockholm Syndrome.
It was content like that which caused the issues. Regardless of whether it’s eloquent trolling or just a genuine but passionate piece, it was touching on an already hot topic with poor consideration about how it would be received.
And that’s fine for a personal blog. But when you then see your articles explode online and then proceed to write follow up pieces in the same tone and intended for the same audience (regardless of whether he directly submitted it to HN), it’s harder to dismiss as someone not trying to exploit flame wars to booster their own blogs traffic.
I guess they succeeded in that too; albeit a DDoS attack wasn’t quite what they intended.
To be clear, I don’t agree with the DDoS attack. Nor do I believe they deserved it (what they actually deserved was just for the articles to get flagged and forgotten) but I can still blame the author for the arguments on here when they saw the existing discourse and decided to write follow up pieces of equally antagonistic tones.
> I can still blame the author for the arguments on here when they saw the existing discourse and decided to write follow up pieces of equally antagonistic tones.
seems very victim blaming to me.
Amos’ articles are opinionated, well written & amusing rants. If a bunch of immature Go programmers can’t take a spot of criticism directed at their favourite language then that says a lot more about them than it does about anyone else.
Given HN is the victim in that context, that would mean my statement is the literal opposite of victim blaming.
Funny because he works for fly.io which explain all the Rust thing but doesn't fly.io use a lot of Go as well indirectly?
https://fasterthanli.me/articles/lies-we-tell-ourselves-to-k...
"It may well be that Go is not adequate for production services unless your shop is literally made up of Go experts (Tailscale) or you have infinite money to spend on engineering costs (Google)."
You should tell that to the thousand of compagnies running Go just fine in production.
Your productivity in any language is directly proportional to your time and effort investment in it. It's in your best interests to pick the language that's likely to thrive and spend time learning it's ins and outs. On the flip side, betting on a horse that doesn't win could mean the loss of months or years of effort. This is why people evangelize the platforms they're invested in - convincing other people to join improves the health of the platform, increasing their return on investment.
This evangelizing can sometimes become contentious if others perceive it as an attack on their platform. People defend their language mostly because they don't want to see it lose popularity. If it did, their language's viability is threatened and their investment is in jeopardy. It's also partly because they've spent so long on it that it's become a part of their identity.
A person who thinks of themselves as a “Go developer” rather than a “developer” is going to take that article personally.
E.g.: the "east vs west" fight in Ukraine is a great example. One side is invested in the democratic/capitalistic model, the other in the autocratic/central-planning model. The war isn't just over a patch of land on the border of Europe.
If people are willing to start a shooting war with actual blood, violence, and death over a mere "ideology", then it shouldn't come as a surprise that developers are willing to go to a "war of words" over their favourite language or platform.
As an aside, I very much enjoyed "A half hour to learn Rust" from fasterthanli.me as it got me to actually start writing simple stuff in the language so I'm a little biased in favour of the author, but the critique didn't seem to be particularly harsh. As someone who's primarily a Java dev I'm used to seeing much more biting (and often inaccurate) condemnation of my own preferred tool!
[1] https://fasterthanli.me/articles/a-half-hour-to-learn-rust
Reading what dang wrote makes it seem really sensible.
Links to dang's comments:
https://news.ycombinator.com/item?id=31228950
https://news.ycombinator.com/item?id=31227642
https://news.ycombinator.com/item?id=31227584
https://news.ycombinator.com/item?id=31117569
The downweighting is intended for when overly generic or offtopic subthreads are stuck at the top of a page, choking out more interesting or on-topic conversation.
(Unfortunately it caused some trouble a while back because one such downweighted subthread was critical of a YC startup, and a much more important rule is that we moderate HN less when YC or YC startups are the topic. But it was a straightforward mistake—the user who marked the thread generic just didn't realize that the topic was YC related.)
Since it's an experiment, we want to keep an eye on it. Currently I have the software send an email each time a downweight is applied, so we can review them as part of going through the regular HN inbox. When I saw this one, it was such a perfect example that I came to the thread and auto-collapsed it as well. So that's its current status: downweighted and auto-collapsed.
Btw, basically none of these situations get created intentionally. They're a tragedy-of-the-commons thing where each person does what they do innocently but it all adds up to something suboptimal. That's pretty much what moderation exists for: to jiggle the system when it gets stuck in one of its failure modes.
The dots that I didn't connect was that the post was a reply to the response on HN and since the post was controversial, it (being a HN comment, in a sense) should have been "more thoughtful and substantive, not less".
Perhaps the headline was the worst part of the article? Like you say - "Lies we tell ourselves to keep using [X]" could have just been "Common reasons to use [X], and why I disagree with them".
Thanks for pointing it out!
I was hoping for the discussion here to be around similar experiences, how DDoS really has become a commodity, how other folks's websites is architected, how various platforms fare against these threats (and what it costs), whether folks were are of Cloudflare's default caching policies, or its DDoS protection.
Anything BUT rehashing the discussions of the past few days: everybody is over that. If you do want to discuss those, you can move that conversation to my subreddit, or my twitter, or heck, send me an e-mail.
(Also, and I've stated this elsewhere many times: the Go community doesn't have to answer for that. This is one individual having well-timed fun, let's please PLEASE leave it at that)
Pin<Box<TimeoutWriter<TimeoutReader<TcpStream>>>>? Gesundheit indeed.
Great writing as always.
Bigger sites might handle that number of requests every second. Hand coded and highly optimized services can handle that number of cached small requests every minute on one machine. After all, that's only an egress rate of a few GBits. And your homepage certainly ought to be both cached and small.
I can think of maybe 10 sites that get anywhere near that amount of traffic and they have millions and millions of dollars of infrastructure behind them. 34 million requests per second is an absurd amount! Nobody even knew how to handle that level of traffic 10 years ago.