Show HN: Tiny, fast, and free API to geolocate IP addresses
github.com
github.com
const realClientIpAddress = (req.headers['x-forwarded-for'] || req.ip || "").split(',')
const ip = realClientIpAddress[realClientIpAddress.length - 1]
X-Forwarded-For is appended-to for every proxy the request passes through. You want the first IP address, not the last one.Example, if your app was on Heroku behind Cloudflare, the request will look like this:
IP: <Heroku's load balancer addr>
X-Forwarded-For: <Real user addr>, <Cloudflare addr>
Your code, as written, will be geolocating the Cloudflare node.
So it’s the opposite of what the normal behavior is, meaning the real client IP is guaranteed to be the last in the list. I probably should have a condition to get the IP based on this logic only if the app is hosted in heroku, then use the standard express way otherwise.
The last entry in the list is simply the IP address of whatever is making the request to Heroku. In my example, that's Cloudflare, being a proxy. Heroku simply appends the originating IP address (coming from a proxy) to the header which is what you would expect.
The Stack Overflow answer is addressing X-Forwarded-For spoofing, something you don't care about for geoip lookup. Someone could prefix 8.8.8.8 to the header before making their request, thus "X-Forwarded-For: 8.8.8.8, <Real IP>, <Cloudflare IP>", and it's inconsequential that your service will return results for 8.8.8.8 instead of <Real IP>.
The SO answer is wrong that this is Heroku-specific behavior. Heroku is simply appending the originating IP address to the header.
Obviously you can push a route to production that logs req.headers to see for yourself to nip this in the bud.
IP: <Heroku's load balancer addr>
X-Forwarded-For: <Spoofed addr>, <Real user addr>, <Cloudflare addr>
So you either need to know the number of proxies that you trust that are in between you and the user, or you need to know the IP addresses of those trusted proxies, in order to determine which parts of X-Forwaded-For to trust.
For just about every geoip use case, there is a better solution. Namely, almost every modern phone and desktop is capable of providing it's location, and is more accurate than any geoip database.
The main issue is that people can make their device lie about it's location, so if you're using geoip for security (say you're a streaming service) then that's about the only valid use case, and that only exits because studios still want to live in a world where borders matter.
They don't need to be perfect and its certainly better than nothing.
You need user permission to get that.
Just because it requires effort does not mean it can't be done in an open source way.
2) IP Spidering via traceroutes / RIPE/etc data.
3) Agreements with third parties that have IP/Address mapping due to data supplied from users. [least accurate]
That'd be my guess anyway.
But there's also a NSA patent on this topic, "Method for geolocating logical network addresses" (filed in 2000).
https://patents.google.com/patent/US6947978B2/en?oq=6%2c947%...
Not anymore, maybe they are reading HN as well :-)
2023-09-15 - Adjusted expiration
> 2019-12-09 - Application status is Expired - Fee Related
2005-09-20 - Publication of US6947978B2
2005-09-20 - Application granted
2002-07-04 - Publication of US20020087666A1
2000-12-29 - Assigned to GOVERNMENT OF THE UNITED STATES, AS REPRESENTED BY DIR. NAT. SECURITY AGENCY, THE NSA GENERAL COUNSEL (IP&T)
2000-12-29 - Priority to US09/752,898
2000-12-29 - Application filed by National Security AgencyBTW, do you use geolite or geolite2 db? The former is getting deprecated next month.
A bigger problem seems to be that many forget to continuously sync their IP DB with their provider. Your targeting is only as good as your IP -> Geo map.
My team built a tool for testing GeoIP implementations here: https://www.geoscreenshot.com to get around the issue of testing if it works.
For corporation, there is another form of targeting (account based targeting) that relies on IP ranges. I believe DemandBase covers this specific use case.
(not affiliated, just a fan)
Example: It is not fit for security postures (in theory). One can dump all the CURRENT v4 routes being advertised out of China and block them via blackholes/firewalls/etc. However immediately after that a rogue operator could hijack a non-China affiliated prefix, use it for badness, and then release the hijacked prefix.
Most Geolocation services that are static (point in time) will not detect the above scenario. BGP-based monitoring services will, but that's a step up $$$ wise.
There's a tradeoff. No, geolocation isn't perfect, and it's often oversold, but it can be useful in a security context. Simple (admittedly reductionist) example: say I have no admin-types in China, and don't expect to. It's pretty simple operationally to grab the 'China range', block port 22/tcp (or whatever the hackers are after today), reduce my risk surface area by a billion IPs or so, get that noise out of the logs (maybe collect statistics on the rule for trending/anomalies), and then have more bandwidth to spot the edge cases where a hijacked block is coming after me. Far from a 100% solution, but maybe a 90% solution. Another tool in the toolbag. Your risk model may vary.