Hacker News BigQuery Dataset
console.cloud.google.com
console.cloud.google.com
We're aware the dataset hasn't been updated since a month ago, and we are working to fix it. You can track the issue here:
- https://issuetracker.google.com/issues/127132286
In the meantime you can still play with the dataset, and dig into the full history of Hacker News - less this last month. I left some interesting queries to get you started here:
- https://medium.com/@hoffa/hacker-news-on-bigquery-now-with-d...
Late last night I had a conversation with someone explaining that Hacker News is not your typical message board -- it's owned and operated by YC and sits atop algorithms developed by some of the pioneers in spam and anomaly detection [1] [2], and it's is also an open dataset -- analyzed and scrutinized -- used by hackers worldwide to train and test bespoke AI.
HN is a live MNIST [3] for anomaly detection.
[1] http://www.paulgraham.com/spam.html
[2] http://googlesystem.blogspot.com/2007/07/paul-buchheit-man-b...
It's a side project so may have some issues!
Top posts about bootstrapping (https://news.ycombinator.com/item?id=19258249):
#standardSQL
SELECT *
FROM `bigquery-public-data.hacker_news.full`
WHERE REGEXP_CONTAINS(title, '[Bb]ootstrap')
ORDER BY score DESC
LIMIT 100
Count of YC startup posts over time by month (https://news.ycombinator.com/item?id=19185946): #standardSQL
SELECT TIMESTAMP_TRUNC(timestamp, MONTH) as month_posted,
COUNT(*) as num_posts_gte_5
FROM `bigquery-public-data.hacker_news.full`
WHERE REGEXP_CONTAINS(title, 'YC [S|W][0-9]{2}')
AND score >= 5
AND timestamp >= '2015-01-01'
GROUP BY 1
ORDER BY 1I manage the BigQuery Public Datasets Program here at Google. You're right, the dataset last updated February 2nd, but we intend to continuing updating it. We had an issue on our end that disrupted our update feed, but we're working to repair it now and get the latest data uploaded to BigQuery.
Hacker news is 12 years old. That's an average of 7 comments per day since inception. Wow
#standardSQL
SELECT
author,
count(DISTINCT id) as `num_comments`
FROM `bigquery-public-data.hacker_news.comments`
WHERE id IS NOT NULL
GROUP BY author
ORDER BY num_comments DESC
LIMIT 100;On the full table:
#standardSQL
SELECT
`by`,
COUNT(DISTINCT id) as `num_comments`
FROM `bigquery-public-data.hacker_news.full`
WHERE id IS NOT NULL AND `by` != ''
AND type='comment'
GROUP BY 1
ORDER BY num_comments DESC
LIMIT 100
tptacek is in first place with 47283 comments.I’m excited about this alternative
NB: That's easy to downvote without commenting...
The HN API [1] has been around in various forms for years and includes the same public data that's used to generate the public pages on the HN site, but rather than returning HTML pages designed for human consumption, the API returns the data in a JSON serialized form [2] designed for machine consumption [3].
When the HN API went live, it reduced the overhead and redundant work from all the programmers having to independently crawl and parse site. The HN BigQuery dataset is the same data returned by the HN API, Google just took the next step and did the work of loading it into BigQuery.
[1] https://github.com/HackerNews/API
[2] https://en.wikipedia.org/wiki/Category:Data_serialization_fo...
By uploading any User Content you hereby grant [..] a nonexclusive, worldwide [..] irrevocable license to [..] distribute [..] your User Content for any Y Combinator-related purpose in any form [..]
Agreeing to the T&Cs and deliberately sharing information publicly covers the GDPR's "consent" lawful base.
Even under GDPR this is not a situation where someone signed up for something else and then happen to have their personal data shared as a byproduct. They signed up to a site, agreed to T&Cs, and then explicitly and deliberately shared their personal data.
Also importantly, the GDPR requires that a controller not make a service conditional upon consent. Hacker News is likely not in compliance unless they make such data processing optional and require anyone interested to explicitly opt in.
But, then again, I'm not a lawyer, and even if I were, actual lawyers don't seen to know what the hell the GDPR actually requires either.
Again, I'm not accusing you of anything here, I'm just pointing out who benefits from framing the conversation this way. So far there is a lot of precedent for small operators shutting down their sites out of fear of GDPR, but there is actually no precedent for regulators having actually gone after small operators for anything resembling reasonable practices. The day may come where EU regulators try to crack down on forums for who are unwilling or unable to redact users messages post-facto, but we're nowhere close to that today and I don't see strong reason to believe that's where we're headed either.
All of us here are users of this forum, so this concerns the legal rights to our personal information. It’s not FUD for us to discuss how those rights are affected by things like this.
Now this position is certainly debatable, but I think it's at least a reasonable argument that you could take to regulators. Contrast that with the bullshit that Facebook, Google and a zillion ad-tech companies are doing with our data every day. You're free to object to the syndication of HN data, but personally I feel that is a distraction from the issues GDPR is meant to address, and I am hoping regulators feel the same way.
Correct. You can certainly attempt to assert your right of erasure with YC to erase your PII from their data (i.e. Hacker News).
But..! Because we give YC the right to distribute our content freely, we simultaneously realize that there may be many duplications and reproductions of this data. The consequence of this is that we must contact any/every user of that data yourself on a one-by-one basis to assert your right of erasure - there is no legal obligation for HN to track everyone who might have downloaded a legal archive of their data.
What we truly need is common crawl data then we can check specific site on our own.
Or wait, BigQuery simply can't handle common crawl size dataset in their public service!
Otherwise there is no reason to not add it, maybe it puts their search engine/ad business in geoparady.
Is there any other Google public dataset BigQuery like platform? Where their direct search engine/ad platform interests don't get in way of Common Crawl like data searching/indexing?
This is not true. Source: ex-googler.