Making open source data more available
github.com
github.com
SELECT repo.id, repo.name, COUNT(*) as num_stars
FROM TABLE_DATE_RANGE([githubarchive:day.], TIMESTAMP('2015-01-01'), TIMESTAMP('2016-12-31'))
WHERE type = "WatchEvent"
GROUP BY repo.id, repo.name
ORDER BY num_stars DESC
LIMIT 1000
Which results in this output: https://docs.google.com/spreadsheets/d/16yDS2wDdDOTxjVsjGvWm...Since the query only hits 3 columns, it only uses 15.4GB of data (out of a 1TB allowance)
More information on the GitHub Archive changes: https://medium.com/@hoffa/github-archive-fully-updated-notic...
SELECT
CONCAT("https://github.com/",repo_name,"/blob/master/",path) AS file_url,
FROM
[bigquery-public-data:github_repos.files]
WHERE
id IN (SELECT id FROM [bigquery-public-data:github_repos.contents]
WHERE NOT binary AND LOWER(content) CONTAINS 'easter egg')
and path not like "%.csv"
GROUP BY 1
LIMIT 1000
36s elapsed, 1.79 TB (so not free). Using github_repos.sample_files and github_repos.sample_contents only costs 31 GB (free) but not as many easter eggs :)https://raw.githubusercontent.com/bduerst/GithubEasterEgg/ma...
- https://medium.com/@hoffa/github-on-bigquery-analyze-all-the...
The Changelog also invited us to record podcast with Arfon Smith (GitHub), Will Curran (Google), and me (Google) - https://changelog.com/209/
Happy to answer any questions!
http://google-opensource.blogspot.com/2016/06/github-on-bigq...
Some big query tricks to make it work:
- TOP/COUNT is faster and more memory efficient than GROUP BY/ORDER
- Filtering data prior to join in sub-query reduces memory usage.
- Regexps and globs are expensive. Use LEFT/RIGHT as a faster version.
- Avoid reading all files to get around 1TB freebie scan limit. Only access file contents after filtering some paths.Hope you will find it useful
SELECT id, title
FROM [bigquery-public-data:hacker_news.full_201510]
WHERE title CONTAINS "Ask HN" AND url=""
LIMIT 1000
Output: https://docs.google.com/spreadsheets/d/12HZ2DqkR_nl380bpxM0B... SELECT title FROM [bigquery-public-data:hacker_news.stories]
where title like '%Ask HN%' LIMIT 1000People take their time "studying" everything these days, isn't it?
I can't imagine how would that be if the State wasn't paying them to do that.