I'm blocking connections from AWS to my on-prem services
consulting.m3047.net
consulting.m3047.net
https://docs.aws.amazon.com/vpc/latest/userguide/aws-ip-rang...
https://www.microsoft.com/en-us/download/details.aspx?id=565...
https://support.google.com/a/answer/10026322?product_name=Un...
etc etc...
Perhaps a better use of your time (because de-balkanization of the internet would be more of an academic exercise than a mass market reality; consumers have already self-selected into social media and crappy LLM outputs which is a far bigger matter and driver than ephemeral resources or reverse DNS) would be making your TLS work in a way that doesn't allow for easy abuse. So automatic redirects, HSTS, and a certificate that is standards-compliant.
Tangent, but last I checked, it was all the big tech giants pushing it down our throat. Unless it's your full time job to find loopholes and workarounds, there is no reasonable way for consumers to opt-out.
While it may seem more useful to aggregate the ranges in some points of view it'd be significantly less useful from other points of view. E.g. those who want to whitelist any IP ranges matching a specific DC, service, availability zone, or country.
You can always aggregate the detailed list but you can't do the inverse on your own.
[Edit] I think this might be where I found it [1]
#!/usr/bin/perl
use strict;
use warnings;
use Net::CIDR::Lite;
my $ipv4String='[0-9]{1,3}\.[0-9]{1,3}\.[0-9]{1,3}\.[0-9]{1,3}';
if(defined $ARGV[0] && $ARGV[0] eq '-h'){
print "usage: $0
This script summarizes your IP classes (if possible). Input IPs with mask one per line. End with CTRL+D. Optionally, redirect a file to stdin like so:
$0 < cidr.txt ";
exit;
}
print "Enter IP/Mask one per line (1.2.3.0/24). End with CTRL+D.\n";
my $cidr =Net::CIDR::Lite->new;
while(<>){
if(/($ipv4String\/[0-9]{1,2})/){
my $item=$1;
$cidr->add($item);
}
else{
print "Ignoring previous line.\n";
}
}
my @cidr_list = $cidr->list;
print "======Aggregated IP list:======\n";
foreach my $item(@cidr_list){
print "$item\n";
}
[1] - https://adrianpopagh.blogspot.com/2008/03/route-summarizatio...I find it a shocking that people still expose internal web services (e.g. Gitlab) openly to the Internet, in my opinion you should at least have one additional layer of protection through a VPN or similar mechanism so that your services aren't discoverable from the public Internet.
I only expose SSH from a single bastion host, which is the only host that's publicly reachable, something that I'd like to get rid off in the future as well by adding a VPN layer on top.
What VPN software would you use? Personally I've never found anything I consider as trustworthy as OpenSSH.
Personally, I find that having to set up an OIDC provider is too much overhead for a VPN. In a corporate setting, you likely have something already, but for individuals or small teams it's too much extra work.
So to give an example if I enter http://geder in my browser I want that to resolve to 100.100.5.10 regardless of if I am on my home network (where geder is) or if I am on a train.
From my perspective half the reason to use tailscale is that it replaces why I'd want mDNS with less bugs.
Under what condition, the root pw being "admin"?
- put ssh on a port not 22
- only allow key-based logins
- don't allow root logins
- keep software up to date
Not that I am an expert... So please tell me if I have a hole somewhere in my setup.
- don't have a guessable password
With extremely rare exceptions (the NSA might), attackers don't have some magic sauce that breaks SSH. Even the recent SSH vulnerability was very hard to actually exploit (but you should have updated ASAP anyway). Their strength is that they just guess passwords all day long on the whole internet. If one server has "admin", or "root", or "1234", they'll get in instantly. If one server has "alcatrazquinine" they'll get in less instantly. If one server has "XgMTaJR35a7gSpXTD2T", they won't ever get in. This is secure from all the people scanning ssh keys. Well, don't use that exact password I just published.
Key authentication is preferred for two reasons. One is that if you accidentally connect to the wrong server you won't transmit your password to that server. The other is that you can store your key in a file and use it automatically so that just typing "ssh myserver" gets you all the way to a shell prompt. That's very convenient.
Not allowing root logins can make sense for auditing reasons (so you can see which user logged in and then used sudo), but if this is just your private server, there isn't really much reason to avoid it. If it makes you feel better, just pretend your name is "root". It also makes sense if you subscribe to the philosophy of "typing sudo in front of every command helps prevent mistakes," which I don't.
Using a port other than 22 can remove provide a very slight decrease in bandwidth and CPU load, and a bigger decrease in log file output, from processing failed logins by scanners. If these things actually matter to you, go ahead. I promise they don't. Doing it for security is either paranoia or security theater, depending on whether your password is "XgMTaJR35a7gSpXTD2T" or "1234".
Keeping software up to date: of course.
[0]: http://netpatterns.blogspot.com/2016/01/the-rising-sophistic...
If you will expose something (even some deep hidden web component) make sure that it is not a door to your data and infrastructure.
What had changed a bit since then is from where that traffic comes, and deciding if its right to receive/block it or not. Legal servers/services, end user side proxies, cloud providers, the amount of crawlers had increased a lot, and so on.
Legitimate question. I assume there might be a benefit, I'm just not sure what it is.
SSH allows to do so many things, has so many config options. To the opposite, Wireguard is very simple, allows basically one thing (authenticate by a key, then pass packets), and is much harder to misconfigure.
(OpenVPN, on the other hand, does not have this advantage.)
You can achieve a somehow similar result running by OpenSSH as `ssh -N -D`: do not run anything on the remote end, work as a socks5 proxy.
ssh -w -N sets up a network interface and does not run a shell
It still runs over TCP, so it's not ideal. TCP-over-TCP is a recognized antipattern that causes extra retransmissions, wasted bandwidth and delays.
It only bothers you because you're looking directly at it in your console. If you bring an overly sensitive Geiger counter on your aeroplane ride, you might be alarmed, but if you're not aware of it, it won't hurt you at all.
I am also surprised by internal services being exposed to the internet, but that's for two reasons: (1) I don't trust most programs' authentication systems, and (2) I don't want people to know which services I'm using internally - not from a technical security standpoint, but sometimes just privacy. But things that are supposed to be on the Internet, that I trust to have a strong front door (or be suitably sandboxed) - they can be on the Internet all day.
Port scanning is just the Internet equivalent of walking around the city taking notes on whose lights are on. Stalkerish? Maybe a bit. But it's public info.
By the way: ssh -w makes a VPN tunnel interface, but it won't auto-configure the rest of the VPN like actual VPN products do.
As far as I noticed, ping with a spoofed source address is the only actual abuse mentioned in the article. It should go without saying that you can't tell if a spoofed ping packet came from AWS, because the source address is the address the spoofer wants you to send a reply to, not the spoofer's address. And a much less invasive mitigation would be rate-limiting pings to, say, 10 per second.
While the Internet is becoming balkanized this is mostly because of social media siloing itself to generate advertising and data revenue and to extract profit from AI training data (e.g. the Reddit/Google exclusivity deal) rather than because of providers blocking IP ranges.
I certainly don't understand the rational mindset behind blocking certain providers over some pings and then complaining about IP connectivity becoming balkanized. The balkanization is caused by the ones doing the blocking.
FWIW, data center IP addresses are already being treated as second class citizens by major content/service providers, and this has become an escalating barrier to self hosting. I am honestly not sure what the author is trying to accomplish.
Could you please expand on this a bit?
1. totally blocked by some services (especially those related to copyright, like almost all the streaming services), 2. treated as suspicious by lots of CDNs (so you would get captchas more frequently; have stricter rate control, etc.)
Also, what qualifies as a data center?
The reply above yours is mostly correct though I have to admit that “data center IP” could be a bit of a misnomer when it comes to IP reputation. There are essentially 4 categories:
- Residential landline connections are the most mundane but are also least restricted because this is where your average users are found. The odds of bad actors on the same network is fairly low, and most ISPs will overlook minor transgressions to not incur additional customer support costs.
- Mobile data connections are often behind CG-NAT. Blocking entire IP range tends to generate a lot of false positives so it doesn’t happen very often.
- Institutional IP ranges (such as 17.0.0.0/8 or any org that maintains their own ASN) tends to get a pass as well because they tends to have their own IT and networking department to take collective responsibility if something untoward was to happen.
- This leaves public could and hosting services on the lowest tier because these networks have very low barrier of entry for bad actors . Connections from these IP addresses are also far more likely to be bots and scrapers than a human user so most TDS systems are all too happy to block them.
And even if I saw your site on here, say, and liked it, if I don't bookmark it immediately, I'd still go to a search engine to try to re-find it in the future if I wanted to go back. Not to say you need to be doing paid ads or trying to raise your SEO or whatever, but I've had times where I remember some unique phrase from an article I read years ago and Google can use that to find the original source.
A large proportion of the resources I host on-prem are just that. Stuff that's too large for email, or isn't static, or may get updated. Unless people host their own email server (like I do), it's going to e.g. Gmrgle anyway if you email it to them. Maybe it gets blocked, maybe it gets fubared.
Uploading it to an on-prem server and sending the link in an email is no more trouble and any issues are easier to debug.
This wasn't some huge technical lift for me to implement. Trust me on that. I got tired of Amazon stinking up my logs and decided that since I can't discriminate based on reliable information about the services being hosted there which have a legitimate need to reach out, I just don't need their help. Really, I'm helping them by ensuring no spurious pongs or SYN/ACKs come from me. See? I'm helping the best I can.
If you think this is heavy-handed and arbitrary, take a close look at email and domain reputation providers sometime.
Need a version of nc which does multicast and is written in python? Well, you can't get that from an Amazon address anymore... unless you've mirrored it. How many people care? How many people care about precinct-level voting patterns for King County Washington from roughly 2005-2009?
It's not a "grand narrative", I've been playing with the internet since it was possible to do so legally. If explaining that history is grandiose for you, that's you. I host on-prem for my convenience and nobody else's. Enjoy the article... or not.
I have not seen individual blogs or small enthusiast sites in search results for quite some time.
If I want wikipedia or stackoverflow I can just search those sites directly. I'd like an option to exclude all the "usual suspects" and see some more long tail stuff.
Why wouldn’t (even small) ISPs run their CGNAT gateway in their own IP space?
Running a CGNAT gateway in the cloud would lead to a lot of problems. I think your subscribers wouldn’t be able to watch Netflix and the like, since they wouldn’t be on a residential IP. It would probably also lead to more anti-bot CAPTCHAs from Cloudflare and Google.
Are there actually any known examples of this?
(Ideally you'd make then switch to TCP by truncating UDP responses to specific clients but that sounds like a hassle to set up so it's understandable to skip that.)
The idea of "I just want the legitimate traffic" is a simple one, but the implementation of the idea has very little to do with "I will just block the big bad cloud!".
Blocking huge IP ranges is knocking yourself half offline, and it doesn't even stop you being "attacked". I'd start blocking if and only if there is some actual problem for your server (e.g. excessive CPU or bandwidth usage), not just because big bad scary cloud.
Me, because i would like to read all of the syslog without meaningless noise.
I don't think your take makes any sense whatsoever. Beyond the puerile "I'll block you too", what exactly do you hope to achieve with this nonsense?
For example, the org might be self-hosting WireGuard or another VPN solution on a cloud provider and people are connecting through that so their outgoing IP address comes from a cloud provider.
All that said, it's trivial to use proxies or VPNs to bypass any blocks.
Maybe for you sure. The large drove of people flocking to serverless these days suggests that even most technical people don't want anything to do with their own networking or infrastructure.
The desire to limit the noise and only allow a "small circle of friends" is also appealing.
But I do that for specific services, not my domains in general. Mumble server: only open to the 3-4 countries that my friends are in, and none of the 'cloud providers'. Tech blog: world+dog can see it.
I am firmly in the 'We all benefit from shared knowledge' camp. So if my notes on modem init strings for my 300-baud C64 modem can help one other person; they won't go through the same pain I went through, and the world will be a better place.
I get the desire, for many reasons. That's cool. You do you.
Amazon is too large to ignore. I understand that a lot of ICMP and SYN traffic is garbage. I'd be happy to help out and block it (and I do have mitigations in place); in fact I do, by default. That's part of why Amazon is a PITA because they trigger my mitigations "at scale". Amazon doesn't help sort the wheat from the chaff: "send a PCAP (for a ping issue)". I don't learn anything by sending Amazon stuff and hearing nothing. I don't need their good traffic any more than I need their bad traffic... or the traffic which is spoofed which is attacking them.
If they can't see fit to help me help them, I don't need any of it. I'm just keeping my life simple.
It's not a technical problem as in overwhelming any resource. It overwhelms me, for starters. Secondly my existing mitigations suggested it, I resisted the move for some of the reasons people here are saying it's a bad idea; I finally concluded the benefits outweighed the costs and I'd try it and see what happens.
So far, so good. Once the fire burns out it should be great.
Why do you publish a blog, if not for people to find?
I can turn that back on for you if you miss it. I saved the Content Imposition Disorder Study Group website, it's still there.
Maybe I'm missing something obvious, but if the author believes the ping traffic is being spoofed, how could they know AWS is the source?
From experience I've seen AWS be the source of overwhelming traffic so many times that we in some cases resorted to the same solution, blocking AWS completely.
I don't know if AWS doesn't care or is just slow to react. Maybe reporting is to difficult, I don't know.
So the obvious you're missing is: AWS IS a huge source of "bad" traffic and getting a misbehaving customer shutdown is too hard, while renting insane amounts of capacity is too easy for bad actors.
Edit: I've almost never seen GCP or Azure being the source of the same amount of crazy traffic.
Data scraping? DDOS attacks? Bandwidth trouble? Security?
I don't think anyone will miss my stuff if they're part of the small minority of people accessing the internet through a VPN hosted in large data centres.
The biggest challenge for implementing this will probably be figuring out how to block inbound connections but keep outbound connections working. I'm sure there's a good nftables rule I can come up with eventually.
nah, there's connection tracking for that.
ctstate new srcip <amazon-set> drop client server
-------- SYN --------->
<-------- SYN+ACK -----I opted for CIDR aggregation and rate limiting of data center ISPs in nginx for one of my frontends. There are reasonable limits for normal IPs too. Not all of us have the capacity or desire to scale.
I publish telemetry outing the worst offenders. What is Amazon doing for me, as a non-customer, to support our presumably shared goal? So maybe it's not really a shared goal.
Here's an analogy which is very visceral for me: I have a chronic condition, which has never been "root caused" in my case by the school of western allopathic medicine. They put me on (several) drugs, I've taken them for a decade.
Due to the ongoing enshittification of medical care over the past several years, I ended up having to see a naturopathic physician (with prescribing privileges, and not covered by my insurance) to get the prescriptions renewed because the western docs simply had better things to do than schedule and keep an appointment. Dude does the things a western doc should do, looks at the labs, and says with all seriousness: have you considered maybe you're allergic to drug "Y"?
Holy fuck well that causes all kinds of problems identifying a replacement therapy, but that's not the point of the story here.
Point is they've never done root cause, blamed the increasingly worse side effects on the condition they never root caused, put me on more drugs for the side effects, and told me to just live with it because that's what happens when you get older: this is what happens when you get trapped in an ecosystem.
So I have identified a replacement therapy, it maybe doesn't work so well for the symptom drug Y was supposed to mitigate but it works. On the other hand, here are the side effects: constellation of symptoms of the actual condition (drug Y was intended to mitigate only one of them) has virtually disappeared, I sleep better, I have more energy, and I've lost nearly 20 pounds in the last six months.
In summary, it's not that hard to run my own server (I have the skills, knowledge and specialized bespoke tools to support doing it my way). It keeps my professional skills sharp and gives me real practical intelligence about the fight in the streets.
This isn't a technical problem, it's a legal/social problem.
I'm going to going out on a limb and guess that all of this traffic that isn't related directly to AWS, but its customers. You can set PTRs for your allocated elastic IPs with a request to support. But then again nobody is going to do it because... it doesn't matter. It may have mattered when you were hosting with a block that you actually truly owned, before the ICANN times, but no more. No one cares. Everything is ephemeral, so why should the reverse matter when things get cycled through addresses multiple times per day? If you're seeing excessive anything, then it's probably time to reach out to the abuse contact published in the whois. Let me help you with that:
OrgAbuseHandle: AEA8-ARIN
OrgAbuseName: Amazon EC2 Abuse
OrgAbuseEmail: trustandsafety@support.aws.com
OrgAbuseRef: https://rdap.arin.net/registry/entity/AEA8-ARIN
Comment: All abuse reports MUST include:
Comment: * src IP
Comment: * dest IP (your IP)
Comment: * dest port
Comment: * Accurate date/timestamp and timezone of activity
Comment: * Intensity/frequency (short log extracts)
Comment: * Your contact details (phone and email) Without these we will be unable to identify the correct owner of the IP address at that point in time.
Use modern features built in to modern versions of common packages and products: rate limiting, redirects, filters, and on and on. If you're just blocking to block to make some sort of statement into the void, you're just hastening that balkanization.The end result of my efforts was:
1) No feedback at all from my reports to Amazon - not even an acknowledgement that my report had been received
2) The spam continued unabated for weeks until I finally had enough and just blocked the entire Amazon SES service
That was a few years ago and maybe they are more responsive now. They sure as hell weren't responsive back then.What precisely is his problem with Amazon?