1,539 karma · joined October 30, 2014
Just for fun:
- In Silicon Valley... Magic Johnson has more google searches than Larry Page.
- In India... Larry Page has more google searches than Magic Johnson.
Just checking Google Trends, as you suggested...
Guess: Who's the Alan with the most pageviews in Wikipedia? Who's the Steve?
I left the answers here:
- https://medium.com/towards-data-science/bigquery-without-a-c...
Basically a show-off for BigQuery, but to answer this specific question: The most viewed Alan and Steve in Wikipedia are the ones closer to HN.
SELECT title, SUM(views) views
FROM `fh-bigquery.wikipedia_v3.pageviews_2019`
WHERE DATE(datehour) BETWEEN '2019-01-01' AND '2019-01-10'
AND wiki='en'
AND title LIKE r'Alan\_%'
GROUP BY title
ORDER BY views DESC
LIMIT 10- https://www.youtube.com/watch?v=JvEvTcXF-4Q
(length 3:14)
We're aware the dataset hasn't been updated since a month ago, and we are working to fix it. You can track the issue here:
- https://issuetracker.google.com/issues/127132286
In the meantime you can still play with the dataset, and dig into the full history of Hacker News - less this last month. I left some interesting queries to get you started here:
- https://medium.com/@hoffa/hacker-news-on-bigquery-now-with-d...
- >500k cores
- >300PB storage
- >12,500 cluster size
- >1T messages per day
And there's also this other talk, "How Twitter Migrated its On-Prem Analytics to Google Cloud" - focused on their migration to BigQuery:
- https://www.youtube.com/watch?v=sitnQxyejUg
- 20 TB/ day of raw log data, >100k events/sec
- Loading ~1TB/hour into BigQuery.
- Serving 5,000+ complex queries / second. p99 ~300ms
Disclosure: I'm Felipe Hoffa and I work for Google Cloud https://twitter.com/felipehoffa.
How to measure number of pageviews instead? I have a solution!
- https://towardsdatascience.com/these-are-the-real-stack-over...
- https://medium.com/google-cloud/big-data-stories-in-seconds-...
Disclosure: I do work for Google Cloud, but all I'm here for is to see if that dog does look like a cat.
Disclosure: I'm https://twitter.com/felipehoffa and I work for Google Cloud. And I'm really excited to reprocess all public tables into clustered ones.
I quoted the reddit talk on the thread above/below :).
a) Do you look at the benchmarks that each company produces, and choose the one that publishes the numbers that make them look the best?
b) Do you ask "I have n data analysts with m different questions and I want to give them the most productive platform I can".
Well, this is what Twitter chose:
- How Twitter Migrated its On-Prem Analytics to Google Cloud https://www.youtube.com/watch?v=N3JAwCYGHU8
Or listen to Nick Caldwell, Reddit VP Engineering, moving away from AWS to BigQuery:
- "2017, which effectively brought us to the present system, we began forking all of our event data into BigQuery, after considering a lot of different alternatives" https://youtu.be/tKISLQ87GO8?t=426
I prefer method b) :)
Disclosure: I'm https://twitter.com/felipehoffa and I work for GCP
- https://seleniumhq.wordpress.com/2017/08/09/firefox-55-and-s...
"The bad news: from Firefox 55 onwards, Selenium IDE will no longer work."
Alternatives:
- https://www.katalon.com/resources-center/blog/selenium-ide-a...
My own metrics, comparing attention on Stack Overflow:
- Katalon Studio immediately started getting attention on Stack Overflow after Selenium IDE was discontinued.
- In Q2 2017 Robot Framework already had more pageviews on Stack Overflow than Selenium IDE. The gap has continued to grow since the deprecation notice.
- In any case, Protractor is the one with the most attention on Stack Overflow between these alternatives.
- https://news.ycombinator.com/item?id=14299731
What's new: # of questions is only half of the story - now you can also look at pageviews %!
- https://medium.com/@hoffa/hacker-news-on-bigquery-now-with-d...
Conspiracy fodder: how smart (or not) would a team need to be to plan a "well-executed marketing play" for an enterprise database warehouse -- in the morning of freakin Saturday December 23rd?
http://images5.fanpop.com/image/photos/31100000/Classic-Patr...
- https://twitter.com/felipehoffa/status/928681468024500224, https://twitter.com/felipehoffa/status/928705663060074496
Since that day, CoinHive has been removed from a lot of sites:
"In the 10/15 run, there were 1,040 mobile sites with the CoinHive Javascript embedded. In the 11/15 HTTPArchive run, this has dropped to 759 - a drop of 27%!" -- Rick Viscomi
- https://discuss.httparchive.org/t/the-performance-impact-of-...
To find all of these sites you can dig into HTTPArchive with BigQuery:
#standardSQL
SELECT
page,
req.url,
REGEXP_EXTRACT(LOWER(req.url), r'(cnhv.co|coin-hive.com|coinhive.com|gus.host|load.jsecoin.com|miner.pr0gramm.com|minemytraffic.com|ppoi.org|projectpoi.com|azvjudwr.info|jroqvbvw.info|jyhfuqoh.info|kdowqlpt.info|xbasfbno.info|crypto-loot.com|coinerra.com|coin-have.com|minero.pw|minero-proxy-01.now.sh|minero-proxy-02.now.sh|minero-proxy-03.now.sh|api.inwemo.com|jsecoin.com)') library
FROM
`httparchive.har.2017_10_15_chrome_requests` AS req
JOIN
`httparchive.runs.2017_10_15_pages` AS pages
ON
req.page = pages.url
WHERE
REGEXP_CONTAINS(req.url, '(cnhv.co|coin-hive.com|coinhive.com|gus.host|load.jsecoin.com|miner.pr0gramm.com|minemytraffic.com|ppoi.org|projectpoi.com|azvjudwr.info|jroqvbvw.info|jyhfuqoh.info|kdowqlpt.info|xbasfbno.info|crypto-loot.com|coinerra.com|coin-have.com|minero.pw|minero-proxy-01.now.sh|minero-proxy-02.now.sh|minero-proxy-03.now.sh|api.inwemo.com|jsecoin.com)')
GROUP BY 1,2,3"Patent information accessibility is critical for examining new patents, informing public policy decisions, managing corporate investment in intellectual property, and promoting future scientific innovation. The growing number of available patent data sources means researchers often spend more time downloading, parsing, loading, syncing and managing local databases than conducting analysis. With these new datasets, researchers and companies can access the data they need from multiple sources in one place, thus spending more time on analysis than data preparation."
- https://medium.com/google-cloud/showing-off-the-new-free-goo...
Feature wise Data Studio has improved a lot since that day, and will continue to do so.
(but just try it out, it's free https://datastudio.google.com/)
Data Studio has a different set of strengths: It's the quickest way I can get an interactive viz published with 0 infrastructure needed. Just build your dashboard, add some controls, and publish it to your closest connections privately, or publicly to the whole world. It will scale without any resource allocation on your side.
(disclosure: I'm Felipe Hoffa and I work for Google Cloud https://twitter.com/felipehoffa)
(in a parallel thread, someone else mentions how they use re:dash and Data Studio https://news.ycombinator.com/item?id=15446296)
- https://datastudio.google.com/org/aLzLLuH1QJC-2sBBmo7qdw/rep...
(the story behind: https://medium.com/@hoffa/the-most-famous-reddit-accounts-c9...)
| "premium operating system images including Windows Server, Red Hat Enterprise Linux (RHEL), and SUSE Enterprise Linux Server"
(disclosure: I work for GCP)
Not me or OP, but same team :)
(Of all respondents interested in Rust, only 3% are women. Compare with R 12%, and Ruby 11%)
Let's see. What if I take all mentions of each language on HN's who's hiring threads, vs % of women interested in each language?
There is correlation!
Chart:
- http://i.imgur.com/mcN6Ghz.png
If we take out the 2 outliers (r, go) - the correlation is 0.61.
SELECT *
FROM (
SELECT word, COUNT(*) c FROM (
SELECT SPLIT(REGEXP_REPLACE(LOWER(text), r'[^a-z]', ' '), ' ') words
FROM `bigquery-public-data.hacker_news.full`
WHERE parent IN (
SELECT id
FROM `bigquery-public-data.hacker_news.full`
WHERE title LIKE 'Ask HN: Who is hiring?%2017%'
)
), UNNEST(words) word
WHERE LENGTH(word)>1 OR (word='r')
GROUP BY 1
HAVING c>30
) a JOIN (
SELECT LOWER(WantWorkLanguage) language, COUNT(*) responses, ROUND(100*COUNTIF(v='Female')/COUNT(*), 2) perc_female
FROM (
SELECT SPLIT(WantWorkLanguage , '; ') WantWorkLanguage, Gender v
FROM `fh-bigquery.stackoverflow.survey_results_public_2017`
WHERE WantWorkLanguage!='NA' AND Gender!='NA'
), UNNEST(WantWorkLanguage) WantWorkLanguage
GROUP BY 1
HAVING responses>2000
) b
ON a.word=b.language
(caveat: "go" is an overloaded word)Any thoughts on why?
Row WantWorkLanguage responses perc_female
1 R 2477 11.87%
2 Ruby 3743 11.03%
3 Java 9409 8.42%
4 Python 11878 8.22%
5 SQL 10646 7.81%
6 Scala 2972 7.64%
7 JS 15451 7.37%
8 Swift 4282 7.36%
9 PHP 5039 7.22%
10 C# 9640 6.21%
11 C++ 7178 5.20%
12 Go 5500 4.87%
13 C 4536 4.83%
14 TypeScr 5435 4.51%
15 Haskell 2208 4.35%
16 Rust 2604 3.38%
(related thread: https://twitter.com/felipehoffa/status/879806078866776064) SELECT WantWorkLanguage, COUNT(*) responses, FORMAT('%.2f%%', 100*COUNTIF(v='Female')/COUNT(*)) perc_female
FROM (
SELECT SPLIT(WantWorkLanguage , '; ') WantWorkLanguage, Gender v
FROM `fh-bigquery.stackoverflow.survey_results_public_2017`
WHERE WantWorkLanguage!='NA' AND Gender!='NA'
), UNNEST(WantWorkLanguage) WantWorkLanguage
GROUP BY 1
HAVING responses>2000
ORDER BY COUNTIF(v='Female')/COUNT(*) DESC(From your browser to multiple Google zones)