However, ideally, we should have a PoP in Seychelles. Some of our recent expansions have been in offshore territories.
380 karma · joined October 27, 2022
Things I am responsible for at IPinfo:
- https://is.gd/devrel - https://ipinfo.io/community - https://ipinfo.io/probe-network - https://ipinfo.io/developers - https://ipinfo.io/lite
However, ideally, we should have a PoP in Seychelles. Some of our recent expansions have been in offshore territories.
First, we use active measurement for IP geolocation as our primary source. Our main data source is our ProbeNet infrastructure, which consists of 1,200 PoPs across 500 cities. Through ProbeNet, we run active measurements to every IP address and routable IPv6. Because we run both ping and traceroute, we typically have a very strong sense of where an IP address is located.
However, ProbeNet is not our only data pipeline. We process several dozen additional data sources. For example, unrouted or unassigned IP addresses do not generate active measurements, so we must rely on other forms of data.
In reality, we must “fallback” to alternative location evidence when active measurement is unavailable. Even though we manage, expand, and maintain a very large server network and a highly complex data pipeline, some IPs require us to rely on what the ASN operator reports.
It’s fair to say that no ASN other than AS131279 (https://ipinfo.io/AS131279 ) is located in North Korea. However, if someone asks for evidence showing that a particular IP address is not located in North Korea, that is extremely difficult to prove in isolation. We prefer not to rely on a null-island methodology and instead choose a location based on a hierarchy of hints.
For unrouted or unassigned IP addresses, geolocation can point to random locations, and in such cases we often must rely on the ASN operator’s data. I’ve seen this happen, even among well-established ASN operators. Some assign random placeholder locations to unrouted and unassigned ranges particularly for IPv6 IP addresses.
More context:
https://community.ipinfo.io/t/the-north-korean-gamers-on-ste...
https://community.ipinfo.io/t/why-is-this-orange-com-ip-rang...
Maintaining an IP geolocation database requires some upkeep. You have to download the database regularly (in our case, daily) to keep the data fresh, and you need a system in place to make it useful.
That’s why we created a dedicated API tier that offers unlimited requests. The data is being used by many open-source projects, so we’re simply doing our part to support them by providing both the data and the API infrastructure service. Last year, we processed over 2 trillion API requests across all our API services. There are many projects, Open Source and Enterprise, that are making billions of requests daily, and they are on a free tier plan.
I manage approximately 1,200 PoPs that are part of IPinfo's ProbeNet platform, which is used to generate our internet measurement data. We, in turn, use this data to produce the IP geolocation data you utilize. The issue is that the servers we manage can be moved to different locations or not be in the advertised locations. We operate these PoPs across 500 cities.
Whenever the PoP fails to meet certain physical location checks, I run basic diagnostic tests to determine where the server is located and where it is not. Aside from running ping and traceroute operations to the target servers, I can run traceroutes to certain IP addresses whose paths I am familiar with. A traceroute visualizer provides a visual interface, information on the ASN, the geolocation, and the time measurements. This provides an intuitive view of where the server "could not be" located rather than could be located. We use several techniques to run basic diagnostic tests. This traceroute visualizer isn't an official test of IPinfo; it is something I vibe-coded together. There are far better internal tools, such as running ping and traceroutes across all of our ~1,200 servers simultaneously.
It is a network diagnostic tool. I think based on your comment, it is not just a tool, it is more of an abstraction of a tool! But it is somewhat useful I thinkl.
I am happy to hear your thoughts. I manage these servers and we are trying our best to improve our data consistently. So, any ideas or even random thoughts you might have can help us improve.
We run ping and traceroute operations from 1,200 servers across the world to every (within reason) IP address out there. So, we have ping RTT and traceroute data history of these measurements.
Considering the physical network topology does not match the physical network topology, we would love to hear your thoughts.
Unrouted IP addresses do not appear on the internet traffic, and we have limited network data for them.
If you can share your traceroute output and obfuscate as much data as possible, we will be happy to investigate and share feedback.
The dataset the user is accessing (IPinfo Lite) is licensed under CC-BY-SA 4.0 and does not come with an End User License Agreement (EULA).
Most free IP geolocation databases in the market do include restrictive EULAs. These usually require each individual user to register, obtain their own copy of the database, and limit usage to themselves only. In many cases, sharing the data or keys outside of the original intended scope is considered a violation of license agreement. Because of these restrictions, some providers also offer paid redistribution licenses, which can be expensive and typically targeted at enterprises.
Our approach is different. By licensing IPinfo Lite under CC-BY-SA 4.0 WITHOUT an EULA, we allow anyone to share or redistribute the data freely, as long as proper attribution is given. The goal is twofold:
- Make it easier for projects to use the data without legal barriers.
- Ensure that any questions or issues about the data are directed to us rather than to project maintainers.
Some large open-source projects, major global enterprises, and governement institutes are already using IPinfo Lite right now.I appreciate you taking a look. Let me know if you have any feedback for us. I am the DevRel, so I hang out in the community. If you see any issues where you think you need immediate support, post that in the community, and I will respond.
> wouldn't it be possible to just let users download the IPinfo data and use it locally? Does IPinfo offer database downloads?
Of course, you can download our free IP database right now: IPinfo Lite
> Also asking for potential personal use: How does the quality of IPinfo data compare to MaxMind, DB-IP, etc?
We are miles ahead of everyone in terms of accuracy. Currently, we have 1,100+ PoPs across the world running active measurements. While traditional IP geolocation services are no much more than ASN/ISP reported data aggregation and parsing services. Our priority above all is accuracy and at this moment we are likely the industry leader for that.
If you have the time, go through some of our posts in our community and you will be surprised how good our data is right now. I will share my recent favorite one:
https://community.ipinfo.io/t/the-north-korean-gamers-on-ste...
In both cases, they could opt to download our database locally and use it through their own API system.
We sponsor the AlmaLinux Foundation through a data sponsorship for their mirroring system: https://almalinux.org/blog/2024-08-07-mirrors-1-to-400/
But since privacy is a major concern for them, they should just use our IP-to-country database and host an API themselves on top of it: https://ipinfo.io/lite
We are happy to support and be part of any software that want to use our data.
Is making a connection to our API a cause for concern? If that is the case, we welcome OSS projects to user our local IP databases, which includes our free IPinfo Lite database that we primarily designed for firewall and privacy applications.
The intention is to go through a creative motion and develop tools and games around IP data without doubling down on any single idea. A form of confined creativity, I suppose.
This was coded entirely through ChatGPT/Github Copilot, and the entire functionality is based on the front end.
The selection of IP addresses is randomly generated through a simple function. Valid IP (non-bogon) selection and excluding popular ASNs is done through making an API call to IPinfo Lite. Geolocation information is gathered using the free IPinfo API.
There is a number of things that should improved and PRs are welcome: https://github.com/abdullahdevrel/abdullahdevrel.github.io/t...
https://ipinfo.io/blog/how-many-ips-change-geolocation-over-...
Nice to meet you. I love the work you have done so far for IPlocate.io. Keep up the good work. The point you have raised is interesting. However, I am not sure about the "Marketing plays a big role" point you made.
We are a developer-first company.
- Developers need a good product, so we are always committed to providing one, whether it is free or not. The best-in-class product we offer is a result of our obsession with meeting the expectations of our developer community. Our relationship with our developers largely (but not entirely) relies on the quality of our product. And that is not just data (which is again just miles ahead of everyone else). It is integrations, infrastructure, site reliability, uptime, dashboards, tools, CLI, and much more.
- For developers, if they cannot get the best product, you have to tell them why you are not the best. So, we write long community posts and have deep technical discussions. We built the community forum, just for our developers to have conversations directly with us.
- For developers, they need someone to be present and respond, and we are always there. In our community someone from our technical team is available seven days a week.
- For developers, they enjoy a product outside of work, so we will hold huntathon, online games, technical articles and events for them.
Being developer-first comes with revenue perks, sure. Our founder started the company 12 years ago, and it has been slow and stable growth. We have built trust in the community, and today it is impossible for us to find any developers who haven't heard of us. That is not because we have a billboard or spent X amount on paid marketing. They heard of us because they used our product, they like our product, and they know our team.
It is not marketing; it is just general goodwill.
We are constantly expanding our probe network, and my colleagues are also working on stabilizing and improving the network.
From 2025 and beyond, we will be putting a lot of effort into our R&D program. Our probe network provides the data that helps our data and research team to create better models.
Currently, a good portion of the data is available to all with no compromise. When we do make mistakes, we hope our community of users will point them out to us. Each ASN or IP address mistake generates a new ticket, our CEO/Founder is tagged there, our data team investigates it, pushes fixes, and provides explanations.
We are super obsessed about any points of friction any IPinfo user has, but to be honest, that is just what any developer-first business should do. We are obsessed about developer sentiment and perception, but that is what the industry should be. Our users consist of the smartest people I know, and because we go above and beyond to be helpful to them, they do not point out mistakes we make, they actually try to help us.
Consider our probe network server finding. We genuinely hit a wall after 750 servers. Then we reached out to our community of users, who found 150 more servers. There are even developers who will talk to their local hosting providers in their local language just to get us a server.
Then writing code to integrate our data into different places. Our team is extremely small, so we cannot actively contribute engineering contributions to open source projects. Our users usually help us a lot in writing high-quality code for open source projects they already use.
We are humbled because of the help our user share, and that is why we are obsessed with them and trying our best to go above and beyond to help them.
Integrations and maintenance were major issues when it came to users using the IP database. Usage of our IP database in software and platforms where data download facilities, maintainability (updating the database at regular intervals), and using an MMDB reader library were issues that were stopping us from universal adoption. For example, search/SIEM/threat intel platforms, distributed systems, firewall applications etc.
So, we just decided to launch an API to complement our data downloads. It is easier to use, and the unlimited requests make it a strong candidate.
We are rebuilding our backend in Rust and also developing a bulk enrichment API endpoint. The intention of the API system is to replicate the performance and features you get from a local database, with ease of use and minimal friction. Of course, the API is competing against the local database and will never be perfect, but I have to admit that using the IP database, particularly the binary database, and maintaining it is not as easy as using an API service.
Using a dataset-based implementation would require me to have a backend, which is out of the scope of this project. Right now, I'm generating random IPv4 addresses, but if I were generating random IPv6 addresses, I would have to go the database route. For that, I would use our free IPinfo Lite dataset: https://ipinfo.io/lite
My colleagues actually developed an extremely fast algorithm to select truly random IPv6 IPs from a series of CIDRs, which is what you see reflected in our dataset.
Let me know if you have any feedback or suggestions for me, please.
I'm generating random IP addresses on the frontend, then making an call to our free API to validate the "realness" of the IP addresses — mainly to remove bogon IP addresses, non-routable IPs, and IPs from large ASNs (national ISPs, the DoD, car companies, etc.).
Our free API supports 1,000 requests per day from unique IP addresses, so there shouldn't be any issues for low usage. However, if we get more power users who enjoy the game, I’ll switch to our Lite API service (which is also free, https://ipinfo.io/lite) to validate IP addresses, as it supports unlimited requests.
Let me know if you have any feedback for me :) I made it mostly by "vibe coding", I will write a post about the whole process of it.