HNHacker News
TopNewBestAskShowJobs

fhoffa

1,539 karma · joined October 30, 2014

https://twitter.com/felipehoffa
submissionscomments
fhoffa··on A map of the US where city names are replaced by most Wikipedia’ed resident
Oops, edited for clarity/intentions to "closer to HN".
fhoffa··on A map of the US where city names are replaced by most Wikipedia’ed resident
Ok, I see.

Just for fun:

- In Silicon Valley... Magic Johnson has more google searches than Larry Page.

- In India... Larry Page has more google searches than Magic Johnson.

- https://imgur.com/a/KwQ5wJt

Just checking Google Trends, as you suggested...

fhoffa··on A map of the US where city names are replaced by most Wikipedia’ed resident
Not really. This is people looking at Wikipedia.

Guess: Who's the Alan with the most pageviews in Wikipedia? Who's the Steve?

I left the answers here:

- https://medium.com/towards-data-science/bigquery-without-a-c...

Basically a show-off for BigQuery, but to answer this specific question: The most viewed Alan and Steve in Wikipedia are the ones closer to HN.

    SELECT title, SUM(views) views
    FROM `fh-bigquery.wikipedia_v3.pageviews_2019`
    WHERE DATE(datehour) BETWEEN '2019-01-01' AND '2019-01-10'
    AND wiki='en'
    AND title LIKE r'Alan\_%'
    GROUP BY title
    ORDER BY views DESC
    LIMIT 10
fhoffa··on Calculating a record-breaking 31.4T digits of Pi with GCP
Don't miss the video on how Emma did this:

- https://www.youtube.com/watch?v=JvEvTcXF-4Q

(length 3:14)

fhoffa··on Hacker News BigQuery Dataset
Hi, Felipe Hoffa at Google here.

We're aware the dataset hasn't been updated since a month ago, and we are working to fix it. You can track the issue here:

- https://issuetracker.google.com/issues/127132286

In the meantime you can still play with the dataset, and dig into the full history of Hacker News - less this last month. I left some interesting queries to get you started here:

- https://medium.com/@hoffa/hacker-news-on-bigquery-now-with-d...

fhoffa··on Twitter migrates data to Google Cloud
you missed this part "and our flexible compute Hadoop clusters"
fhoffa··on Twitter migrates data to Google Cloud
From the Twitter Hadoop to GCP talk that Boulos linked:

- >500k cores

- >300PB storage

- >12,500 cluster size

- >1T messages per day

And there's also this other talk, "How Twitter Migrated its On-Prem Analytics to Google Cloud" - focused on their migration to BigQuery:

- https://www.youtube.com/watch?v=sitnQxyejUg

- 20 TB/ day of raw log data, >100k events/sec

- Loading ~1TB/hour into BigQuery.

- Serving 5,000+ complex queries / second. p99 ~300ms

Disclosure: I'm Felipe Hoffa and I work for Google Cloud https://twitter.com/felipehoffa.

fhoffa··on What is going on with measures of programming language popularity?
With Stack Overflow most people measure number of new questions - but that's a partial measure that can lead to wrong conclusions.

How to measure number of pageviews instead? I have a solution!

- https://towardsdatascience.com/these-are-the-real-stack-over...

fhoffa··on What Being on the Front Page of Hacker News Does to Your Indie App
Related: I wrote this post showing the number of stars that GitHub projects get after being featured on the HN front page:

- https://medium.com/google-cloud/big-data-stories-in-seconds-...

fhoffa··on Google's AutoML: Cutting Through the Hype
Can you share a picture of your dog?

Disclosure: I do work for Google Cloud, but all I'm here for is to see if that dog does look like a cat.

fhoffa··on Machine Learning in Google Bigquery
And with the just announced clustering now you get even better costs and performance.

Disclosure: I'm https://twitter.com/felipehoffa and I work for Google Cloud. And I'm really excited to reprocess all public tables into clustered ones.

fhoffa··on Machine Learning in Google Bigquery
I really enjoyed this slide from the BuzzFeed talk "Hundreds of terabytes across 300+ tables, 100K+ queries a day" on BigQuery.

I quoted the reddit talk on the thread above/below :).

fhoffa··on Machine Learning in Google Bigquery
Can you post on https://stackoverflow.com/questions/tagged/google-bigquery? I haven't heard about this, but if you give the community good specifics, we'll work on it.
fhoffa··on Machine Learning in Google Bigquery
Let's say you are Twitter, Reddit, or the CTO of your current company. How do you choose your next analytical platform:

a) Do you look at the benchmarks that each company produces, and choose the one that publishes the numbers that make them look the best?

b) Do you ask "I have n data analysts with m different questions and I want to give them the most productive platform I can".

Well, this is what Twitter chose:

- How Twitter Migrated its On-Prem Analytics to Google Cloud https://www.youtube.com/watch?v=N3JAwCYGHU8

Or listen to Nick Caldwell, Reddit VP Engineering, moving away from AWS to BigQuery:

- "2017, which effectively brought us to the present system, we began forking all of our event data into BigQuery, after considering a lot of different alternatives" https://youtu.be/tKISLQ87GO8?t=426

I prefer method b) :)

Disclosure: I'm https://twitter.com/felipehoffa and I work for GCP

fhoffa··on Ask HN: Selenium is gone, How do you do SPA UI testing?
Selenium IDE was deprecated last year - maybe that?

- https://seleniumhq.wordpress.com/2017/08/09/firefox-55-and-s...

"The bad news: from Firefox 55 onwards, Selenium IDE will no longer work."

Alternatives:

- https://www.katalon.com/resources-center/blog/selenium-ide-a...

My own metrics, comparing attention on Stack Overflow:

- Katalon Studio immediately started getting attention on Stack Overflow after Selenium IDE was discontinued.

- In Q2 2017 Robot Framework already had more pageviews on Stack Overflow than Selenium IDE. The gap has continued to grow since the deprecation notice.

- In any case, Protractor is the one with the most attention on Stack Overflow between these alternatives.

- https://imgur.com/a/Jpe9KGl

fhoffa··on These are the real Stack Overflow trends: Use the pageviews
Previous discussion: When Stack Overflow introduced the trends tool (May 2017).

- https://news.ycombinator.com/item?id=14299731

What's new: # of questions is only half of the story - now you can also look at pageviews %!

fhoffa··on Ask HN: What has HN given you?
Even easier: All Hacker News posts are ready to be analyzed in BigQuery.

- https://medium.com/@hoffa/hacker-news-on-bigquery-now-with-d...

fhoffa··on Busting myths about BigQuery
Oh hi! I work for Google. I browse Hacker News in my free time - like now, from a plane.

Conspiracy fodder: how smart (or not) would a team need to be to plan a "well-executed marketing play" for an enterprise database warehouse -- in the morning of freakin Saturday December 23rd?

http://images5.fanpop.com/image/photos/31100000/Classic-Patr...

fhoffa··on Starbucks’ Wi-Fi Found Using People’s Laptops to Mine Monero
There are thousands of websites mining cryptocoins. I recently posted a tweet that went viral in Brazil - it showed that one of their government sites was mining coins - and draining 360% of my notebook's CPU power:

- https://twitter.com/felipehoffa/status/928681468024500224, https://twitter.com/felipehoffa/status/928705663060074496

Since that day, CoinHive has been removed from a lot of sites:

"In the 10/15 run, there were 1,040 mobile sites with the CoinHive Javascript embedded. In the 11/15 HTTPArchive run, this has dropped to 759 - a drop of 27%!" -- Rick Viscomi

- https://discuss.httparchive.org/t/the-performance-impact-of-...

To find all of these sites you can dig into HTTPArchive with BigQuery:

  #standardSQL
  SELECT
      page,
      req.url,
      REGEXP_EXTRACT(LOWER(req.url), r'(cnhv.co|coin-hive.com|coinhive.com|gus.host|load.jsecoin.com|miner.pr0gramm.com|minemytraffic.com|ppoi.org|projectpoi.com|azvjudwr.info|jroqvbvw.info|jyhfuqoh.info|kdowqlpt.info|xbasfbno.info|crypto-loot.com|coinerra.com|coin-have.com|minero.pw|minero-proxy-01.now.sh|minero-proxy-02.now.sh|minero-proxy-03.now.sh|api.inwemo.com|jsecoin.com)') library
    FROM
      `httparchive.har.2017_10_15_chrome_requests` AS req
    JOIN
      `httparchive.runs.2017_10_15_pages` AS pages
    ON
      req.page = pages.url
    WHERE
      REGEXP_CONTAINS(req.url, '(cnhv.co|coin-hive.com|coinhive.com|gus.host|load.jsecoin.com|miner.pr0gramm.com|minemytraffic.com|ppoi.org|projectpoi.com|azvjudwr.info|jroqvbvw.info|jyhfuqoh.info|kdowqlpt.info|xbasfbno.info|crypto-loot.com|coinerra.com|coin-have.com|minero.pw|minero-proxy-01.now.sh|minero-proxy-02.now.sh|minero-proxy-03.now.sh|api.inwemo.com|jsecoin.com)') 
    GROUP BY 1,2,3
fhoffa··on Google Patents Public Datasets: connecting public, paid, and private patent data
This is a huge deal for anyone working in this space:

"Patent information accessibility is critical for examining new patents, informing public policy decisions, managing corporate investment in intellectual property, and promoting future scientific innovation. The growing number of available patent data sources means researchers often spend more time downloading, parsing, loading, syncing and managing local databases than conducting analysis. With these new datasets, researchers and companies can access the data they need from multiple sources in one place, thus spending more time on analysis than data preparation."

fhoffa··on Google Data Studio
An authoring demo I wrote the day Data Studio went public:

- https://medium.com/google-cloud/showing-off-the-new-free-goo...

Feature wise Data Studio has improved a lot since that day, and will continue to do so.

(but just try it out, it's free https://datastudio.google.com/)

fhoffa··on Google Data Studio
Metabase is great - and another open source dashboard solution I'm a big fan of: https://redash.io/

Data Studio has a different set of strengths: It's the quickest way I can get an interactive viz published with 0 infrastructure needed. Just build your dashboard, add some controls, and publish it to your closest connections privately, or publicly to the whole world. It will scale without any resource allocation on your side.

(disclosure: I'm Felipe Hoffa and I work for Google Cloud https://twitter.com/felipehoffa)

(in a parallel thread, someone else mentions how they use re:dash and Data Studio https://news.ycombinator.com/item?id=15446296)

fhoffa··on Google Data Studio
Play with a live interactive Data Studio dashboard - the most famous reddit accounts:

- https://datastudio.google.com/org/aLzLLuH1QJC-2sBBmo7qdw/rep...

(the story behind: https://medium.com/@hoffa/the-most-famous-reddit-accounts-c9...)

fhoffa··on Extending per-second billing in Google Cloud
Specially interesting given other similar recent announcements, Google's per-second billing also covers:

| "premium operating system images including Windows Server, Red Hat Enterprise Linux (RHEL), and SUSE Enterprise Linux Server"

(disclosure: I work for GCP)

fhoffa··on Benchmarking TensorFlow on Cloud CPUs: Cheaper Deep Learning Than Cloud GPUs
https://medium.com/google-cloud/a-day-in-the-life-of-a-devel...

Not me or OP, but same team :)

fhoffa··on Increasing Rust’s Reach
The % is indeed a comparison to men.

(Of all respondents interested in Rust, only 3% are women. Compare with R 12%, and Ruby 11%)

fhoffa··on Increasing Rust’s Reach
True - but since I couldn't find that on time, I created my own ranking with the data I had on hand.
fhoffa··on Increasing Rust’s Reach
Interesting guess!

Let's see. What if I take all mentions of each language on HN's who's hiring threads, vs % of women interested in each language?

There is correlation!

Chart:

- http://i.imgur.com/mcN6Ghz.png

If we take out the 2 outliers (r, go) - the correlation is 0.61.

  SELECT *
  FROM (
    SELECT word, COUNT(*) c FROM (
      SELECT SPLIT(REGEXP_REPLACE(LOWER(text), r'[^a-z]', ' '), ' ') words
      FROM `bigquery-public-data.hacker_news.full` 
      WHERE parent IN (
        SELECT id
        FROM `bigquery-public-data.hacker_news.full` 
        WHERE title LIKE 'Ask HN: Who is hiring?%2017%'
      )
    ), UNNEST(words) word
    WHERE LENGTH(word)>1 OR (word='r')
    GROUP BY 1 
    HAVING c>30
  ) a JOIN (
    SELECT LOWER(WantWorkLanguage) language, COUNT(*) responses, ROUND(100*COUNTIF(v='Female')/COUNT(*), 2) perc_female
      FROM (
        SELECT SPLIT(WantWorkLanguage , '; ')  WantWorkLanguage, Gender v
        FROM `fh-bigquery.stackoverflow.survey_results_public_2017`
        WHERE WantWorkLanguage!='NA' AND Gender!='NA'
      ), UNNEST(WantWorkLanguage) WantWorkLanguage 
      GROUP BY 1
      HAVING responses>2000
  ) b
  ON a.word=b.language
(caveat: "go" is an overloaded word)
fhoffa··on Increasing Rust’s Reach
Data wise: According to the latest Stack Overflow survey, Rust is the language that women are the least interested in.

Any thoughts on why?

  Row	WantWorkLanguage	responses	perc_female	 
  1	R	2477	11.87%	 
  2	Ruby	3743	11.03%	 
  3	Java	9409	8.42%	 
  4	Python	11878	8.22%	 
  5	SQL	10646	7.81%	 
  6	Scala	2972	7.64%	 
  7	JS	15451	7.37%	 
  8	Swift	4282	7.36%	 
  9	PHP	5039	7.22%	 
  10	C#	9640	6.21%	 
  11	C++	7178	5.20%	 
  12	Go	5500	4.87%	 
  13	C	4536	4.83%	 
  14	TypeScr	5435	4.51%	 
  15	Haskell	2208	4.35%	 
  16	Rust	2604	3.38%
(related thread: https://twitter.com/felipehoffa/status/879806078866776064)

  SELECT WantWorkLanguage, COUNT(*) responses, FORMAT('%.2f%%', 100*COUNTIF(v='Female')/COUNT(*)) perc_female
  FROM (
    SELECT SPLIT(WantWorkLanguage , '; ')  WantWorkLanguage, Gender v
    FROM `fh-bigquery.stackoverflow.survey_results_public_2017`
    WHERE WantWorkLanguage!='NA' AND Gender!='NA'
  ), UNNEST(WantWorkLanguage) WantWorkLanguage 
  GROUP BY 1
  HAVING responses>2000
  ORDER BY COUNTIF(v='Female')/COUNT(*) DESC
fhoffa··on Comparing Bandwidth Costs of Amazon, Google and Microsoft Cloud Computing
You can measure it now:

- http://www.gcping.com/

(From your browser to multiple Google zones)

← PreviousPage 2 of 5Next →