Y Combinator Dataset Of Posts
The dataset is 100MB, so only download it if you need it. This dataset may be removed in the next week or so.
The dataset is 100MB, so only download it if you need it. This dataset may be removed in the next week or so.
the latter cancels the former.
For what it's worth, here are additional mirrors.
Posts: http://dl.getdropbox.com/u/315/programming/datasets/ycombina...
Profiles: http://dl.getdropbox.com/u/315/programming/datasets/ycombina...
You also want to mirror the update utility ( http://news.ycombinator.com/item?id=173354 ).
Dir['*.html'].each do |user_file|
user_data = File.read(user_file)
user_name = user_file.chomp('.html')
age = user_data[/created:<\/td><td>(\d+)/, 1]
karma = user_data[/karma:<\/td><td>([\d\-]+)/, 1]
puts "#{user_name},#{age},#{karma}"
endTop users by "karma earned per day of membership":
Username Age Karma K/A
------------------------------------
dhh 1 48 48
pg 563 17544 31
oldgregg 3 87 29
nickb 429 11672 27
donw 3 52 17
pius 210 2803 13
edw519 428 5316 12
rms 427 5017 11
drm237 271 2851 10
hhm 246 2475 10
keating 10 103 10
sah 42 411 9
freax 5 45 9
further08 2 19 9
The ones who got their karma very quickly are the interesting ones to check out.I'll update with the link as soon as it's done.
Anyway, the mirror is:
Anyhow, I was expecting at most 20 concurrent connections before traffic decayed. I wasn't expecting concurrent connections from 56 unique IP addresses. The server is in London and it is transferring 1.7MB/s, mostly to European users. However, latency and routing to US clients seems to drastically reduce throughput to those users.
Most of the HTTP 206 [Partial] requests seem to be from Internet Explorer, despite this being a relatively obscure choice on this forum. I can only suppose that IE is quite inclined to re-establish connections after a connection briefly stalls. The latter would be because the httpd state creeped 39MB into virtual memory on this 64MB RAM server. This would also be why TCP window scaling didn't occur.
Anyhow, I knew it was risky to use this server to serve relatively large files but it will be used in the future for posting smaller tidbits.
One suggestion: it would be even more useful (for my purposes at least) if you had another version that only included the full posts, rather than having the full posts in addition to having separate files for each comment subthread. The way it is now, there is a lot of data duplication, since a comment of depth n will appear in n separate files.