A distributed peer to peer list of bad actor IP addresses and phone numbers
sentrypeer.org
sentrypeer.org
I spent a bit more than a year as the most senior dev on the team that owned the sign up page for a cloud computing company. We faced bot attacks constantly. They had a variety of reasons to attack, but International revenue sharing fraud (IRSF) was a major one. In short, bots would convince our sign up flow to make verification phone calls to numbers that charge a lot of money but don't actually exist. (Note: whatever "why don't you just" you're about to reply with, we had an entire team of people doing nothing but trying to stop this for months and years- we tried that and it didn't succeed, or there were business reasons it was not feasible.)
I switched companies. In my new role, I happened to be on a team in the same org as their sign up page. Chatting with them, I learned that they were not facing similar attacks, but identical ones. The countries involved, the patterns of the attack, everything- this was the same attacker, going up against two completely unrelated companies.
I don't know if this system is the right implementation- I need to do a deeper dive on this- but I do know that something like this is needed so that companies can coordinate their defenses.
I mean if credit cards can do it...
I love that you added this. So many armchair experts with brilliant solutions to problems they know nothing about.
(Some SIP providers allow you to disallow expensive numbers to be called and/or state the minute costs before the call starts (giving you time to hang up.))
but can't legislate away the basic 1980s phone scam.
Wonder why their priorities are so skewed.
Except that clearly is a number which exists? One that answers, GP isn't fighting 1980's phreakers with a bluebox, they are using Twilio, Signalwire, etc to make automated verification calls/sms for a very large company as far as I understood it. The call being answered is baked entirely into your programmatic logic.
Unless I'm missing some niche telco semantics or lingo here? Or they are rolling their own?
To embrace the "don't question it, you don't understand" narrative, I'm actually quite stunned someone would let this go on for such a long time, if this is corruption, crime or some other sort of telecom fraud aided and abetted by foreign governments/companies why not just turn off the tap and ban the country code if it was such a cost sink? Why not use some of the many alternative (and significantly more secure) verification methods for your customers?
There's literally no other solution for customer verification other than an automated phone call into third world countries with corrupt telcos?
We clearly aren't talking about the engineering team here.
Note that these documents frequently list some numbers at ridiculously high prices (eg. €25 per minute), presumably to deter such fraud.
It wouldn't surprise me if the list only applies to consumer plans, or (as you say) high-balls stuff in certain countries.
Good luck collecting your scam cash from my personal line! The phone company doesn't have a mechanism to pass the cost on to me, so I'm guessing the telcos already figure out how to block this at the network level.
Has anyone tried calling some of the bad actor numbers from a burner phone?
As it was explained to me (anyone feel free to correct me if I'm wrong), a shady/criminal organization takes control of a piece of the phone network. They pick a range a phone numbers for a country that is usually expensive to call, maybe a range of numbers that doesn't actually exist in that country, and they use the phone network to "advertise" a slightly cheaper route to connect phone calls to those numbers. Instead if it costing $1, it costs $0.80.
When someone makes a call to those numbers, they charge the cheaper rate, and then don't actually connect them to a real phone. Every time a call is made, they just get money.
The "S" is for "shared". They offer other organizations "for every call you generate to this range of numbers, we give you a cut". Let's call it 20 cents.
Now picture how many organizations are vulnerable to this. Every "call me back when a customer service rep is available". Every "call me to let me type in a code to verify I'm human". Every "leave a number our sales team can reach you at". They're all being attacked by these sorts of things constantly.
The cost to get a bot to trick the system into making a phone call is much smaller than the revenue generated. As long as that is true, the attacker wins.
The automated drive by ones trying to spam posts in every form or are probing for word press exploits or what have you.
The ones where someone is specifically taking time out of their day to tailor their attack to your specific page.
The former is easy to deal with the latter is far harder to sort out.
Humans always investigate thoroughly before serving themselves.
None of these IPs could be zombies; here come internet Supermen to save us!
The main problem is that it's just too hard to map different kinds of abusive/non-abusive actions to a shared scale. The model here seems to be basically a boolean OR: if anyone flagged an IP as bad, it's marked as bad in the data set. Let's say that you're running a shop and detect a bunch of fraudulent transactions for expensive items, and add the IPs to the list. What can somebody else using this list for say spam-filtering comments to a web forum do with your data points in isolation? Not much, because the action you're protecting is rare, high impact, and unlikely to have FPs so the threshold where you'd mark the IP as abusive would be very different than for the spam-filter use case.
At a minimum you'd need all the client of the IP reputation signal to also share the positive reputation signals, not just the negative, and to export some forms of volumes or ratios of good and bad traffic. But it'll still be really hard to combine those reports to a single verdict that's generally useful.
A secondary problem is that different kinds of abuse don't even correlate particularly well. Let's say that you've got an IP address sending SMTP spam; the odds are that this IP will not be doing ssh credential stuffing, credit card fraud, mass-scraping, traffic pumping, warez distribution, or DDOS attacks.
The classic SMTP IP blocklists work only because everyone using them is using them for the same purpose, and is as such on a shared scale, and there is a high likelihood of the same abusive actor attacking multiple different organizations. Your example would fit that as well, and doing reputation for that one domain would actually tractable unlike federated cross-domain IP reputation.
Or if someone added my phone number as a prank, a revenge or an attack on me?
The only way I’ve seen this somewhat work is to have a complex system that pulls apart the connection info, and then you use a combination of data science + threat intel + good ol’ reversing to make decisions of if something is malicious or not. Then you need multiple teams to run and tune these functions as attacks change.
[Edit] I am running tcpdump right now and it didn't even take 15 seconds to start seeing queries for sip SRV records.
Run some test on the list before using it, to make sure your own assets isn’t mentioned in there. If you don’t, you end up creating a DOS vulnerability in your system :D
ShouldIAnswer is probably more oriented toward end users but maybe both projects can still benefit from each others?
I suppose if you see a lot of requests coming out of a network you can just block ips based on the network portion of the address.
challenge accepted
Depending on how this project is structured, it wouldn’t have to process your GDPR request.
Edit: I was only partially correct. According to Article 2, Point 2C, the regulation (GDPR) does not apply to the processing of personal data “ by a natural person in the course of a purely personal or household activity”
So if this project can be viewed as a personal project (which I suppose it could, in theory…), then it wouldn’t have to comply.
Please note that that is not what I’ve meant. I’ve updated my comment with a bit more info.
I think this is a poor choice (from a security perspective). It should be written in Go or rust. C programs (exposed to the network) are dangerous even when written by experienced developers.
Really, in 2022, everything Internet facing should be written in a memory safe language, running as a normal user (no root) and have a strong MAC policy applied. Anything else is too risky.