From side project to 250M daily requests
medium.com
medium.com
Having the data locally allows for potentially synchronous lookups, or at least for lookups with an availability guarantee which makes your code much simpler to reason about.
Plus, you are guaranteed to get the information, independent of the availability of a third party.
And finally for $100 per month I get access to a weekly updated locally available database (that's what Maxmind charges for the city-level database) no matter how many queries I'm going to issue locally.
Flat-fee access to a local database would even allow to bulk post-process web server log files (yeah - they still exist and don't come with the usual privacy issues surrounding third-party analytics providers) within a reasonable time- and cost frame.
Not everything that can be an external service has to be an external service and for geolocation I definitely cannot see any advantage to not having this data stored locally.
This isn't big-data and will easily fit any amount of RAM (if it even needs to), this doesn't require a costly sys-admin team, this doesn't require any hardware knowledge. This is about fetching a file and putting it somewhere on the server. You are doing this daily with your daily web-browsing.
Now, to OP, I'm very happy for you and I appreciate the service you are offering and I'm very happy that you are solving an issue some people are having. I don't want to belittle this at all.
I'm just saying that while I might personally err a bit too much on "doing it on my own", I absolutely cannot see any justification to do IP geolocation with an external dependency.
We use the same dataset that IP info started with (from MaxMind), although it sounds like they have many different sources now. It's worked fine for us, matches over 99% of the IP addresses we look up, and generally responds in less than 5 milliseconds - basically only limited by network latency.
Right now this is published to a private Docker hub repository, for no good reason, but if people would find it useful we can make it public.
My thoughts are:
1) We're not modifying the maxmind DB. We download their MMDB file and leave it completely untouched. Our code basically does something like this:
if ip in custom_db
return custom_db[ip]
else
return maxmind_db[ip]
2) Our database is openly available via our API. Anyone can access it without needing to signup etc.They're the principles I've been operating under, but I'm not a lawyer, and I'll definitely go consult with one to get a definitive answer on where I stand with this. If it turns out that we're not meeting the license terms then we'll certainly make whatever changes necessary to ensure we do.