The result of pinging all the Internet IP addresses
securityartwork.es
securityartwork.es
He found a lot of cool stuff. For instance, there are apparently 7 Windows NT 3.5.1 boxes with SNMP open sitting on the internet. And about 300 IIS servers that give out the same session ID cookie when you log on.
The visualizations are also really nice. The talk is here: http://www.irongeek.com/i.php?page=videos/derbycon2/1-1-2-hd...
Also, I thought this quote on the page you linked was pretty humorous: "[HD Moore] is the Chief Architect of Metasploit, a popular video game designed to simulate network attacks against a fictitious globally connected network dubbed “the internet”.
It seems they just stored if they got a response or not.
4 billion IPs, and if all you need to store is if you received a response, that's only half a gig of RAM. mmap and call it a day?
[0]: http://en.wikipedia.org/wiki/Linear_feedback_shift_register
[1]: http://cs.acadiau.ca/~dbenoit/research/webcensus/Web_Census/...
[2]: http://cs.acadiau.ca/~dbenoit/research/webcensus/Web_Census/...
[1]: http://www.shodanhq.com/search?q=port%3A80
http://www.myspace.com/kruhft/photos/19626037#%7B%22ImageId%...
Myspace lowres'd it and I can't find the original, but if there's any interest I can keep looking. I did a ping 'scatter' scan of a random set of IP addresses and then mapped them into 2d space with the boxes representing the ping times. I thought it looked cool in the end :)
1) The Opte Project (http://www.opte.org/maps/)
2) A hoax torrent download "HACKER TOOL EVERY IP ADDRESS EVER" (http://imgur.com/7CMCceQ)
It works perfectly well as long as your results are not used for anything important I guess. But if you have customers who needs reliable results, this naïve approach simply don't cut it in my experience.
You can split the address space up across several different scanners on different physical links. You can estimate the RTT to a network segment you're scanning and base your timings on that. Probing with TCP packets can yield better results than ICMP packets for this type of activity. There's so many variables involved.
Build a tool that allows you to send ICMP packets at a fixed rate (preferably in the kernel, or even without an OS at all if you're into that. Getting precise timings in user land is hard) or just a tool that sleeps between packets with the possibility of not sleeping at all. It's an educational experience. Scan a relatively small range of addresses bound to hosts on the other side of the world at different speeds and see the diff in results. Maybe there's a good tool for that already.
Whenever I read about "We've scanned/product X can scan the internet in X hours" I'm very sceptical. Unless the results are verifiable in some way (which is hard to guess/estimate for such a large sample) or the approach they took seems like a sane one (very subjective I guess), I assume they don't know what they're doing. The reason I assume this is because I've been there myself.
The problem is not sending packets fast enough. It's not about bandwidth. The problem is sending them just fast enough, which is impossible if you're scanning statelessly with just ICMP echoes.
Let's say you're on a 100 mbit ethernet, your uplink is only 8 mbit. If you send packets at a rate of 10 mbits, packet loss will happen. And you're not the only one using the network either, so this can happen way earlier. And that's only the part of the network that you control. There might be a lot of hops between you and the host you're sending packets to. And with your approach (the way I understand it) you're not gonna notice packet loss.
I might make too many assumptions here, but ten hours is just too short of a time period for a network of that size for a reliable result. I'm very sceptical. But please prove me wrong, because it will def. make my job easier.
I guess you could publish the code, so I could test it myself.
I am also curious about this: "With the extracted data more interesting analysis can be done,...such as the issue with network and broadcast addresses (.0 and .255)." Why do responses from .0 or .255 have to be an issue? My cable modem sits on a /20. It seems that there are a number of valid ip addresses ending in 0 or 255 in this range:
$ ipcalc XX.XX.57.26/255.255.240.0
Address: XX.XX.57.26 XXXXXXXX.XXXXXXXX.XXXX 1001.00011010
Netmask: 255.255.240.0 = 20 11111111.11111111.1111 0000.00000000
Wildcard: 0.0.15.255 00000000.00000000.0000 1111.11111111
Network: XX.XX.48.0/20 XXXXXXXX.XXXXXXXX.XXXX 0000.00000000
HostMin: XX.XX.48.1 XXXXXXXX.XXXXXXXX.XXXX 0000.00000001
HostMax: XX.XX.63.254 XXXXXXXX.XXXXXXXX.XXXX 1111.11111110
Broadcast: XX.XX.63.255 XXXXXXXX.XXXXXXXX.XXXX 1111.11111111
Hosts/Net: 4094 Class ARe: responses from 10/8; they may have some connectivity to local 10/8 resources; or it's possible someone was sending them fake ping responses, and the network path they're on doesn't do proper ingress filtering (many don't).
Some consumer routers filter traffic to/from addresses ending in .0 or .255 in a naive effort to prevent SMURFing
This has been complained about for years and means the most obvious approach to diagnosis of problems fails 93% of the time...
I mean the internet address space for IPv4 is now so tiny relative to our computing resources that visualizing and interpreting the data is fairly easy.
Of course storing a response packet for every IPv6 address might cost slightly more on S3.
I'd be really curious to know how long the full scan took. Couldn't find the info in the article.
Shocked at how fast they were able to ping all the IPs
The site is down for me, but I assume they used multiple machines to do this. I SYN scanned about 70% of the globally routed prefixes last month and it took a little over 4 days from a single box (but I was doing some detailed packet captures that hurt disk IO).
Looks like 10 hours unless I'm mistaken (it is pretty late here)