HNHacker News
TopNewBestAskShowJobs

yaroslavvb

212 karma · joined August 30, 2009

https://medium.com/@yaroslavvb http://yaroslavvb.blogspot.com
submissionscomments
yaroslavvb··on Let's Remove the Global Interpreter Lock
Underlying implementations often have a way to disable parallelism, ie, OMP_NUM_THREADS=1 or MKL_NUM_THREADS=1
yaroslavvb··on Under the Hood of Google’s TPU2 Machine Learning Clusters
You could specify computation in high level API like TensorFlow, and then have the framework pick the best implementations available (ie, CuDNN for GPU, MKL for CPU, something custom for TPU)
yaroslavvb··on Google Is Battling a Russian Spammer Over the Use of the Letter 'G'
I found some fonts reuse exact same vector instructions for Cyrillic "А" and English "A" in the same file
yaroslavvb··on Without affordable housing, Vancouver risks becoming an economic ghost town
Vancouver is on the way to becoming the bar from the joke "nobody goes there anymore, it's too crowded"
yaroslavvb··on Mean Shift Clustering
With mean-shift you can get non-convex clusters
yaroslavvb··on Captchas are Becoming Ridiculous
We found that neural networks can solve CAPTCHAS much better than humans, 99.8% on the "hard" ReCAPTCHA instances: http://arxiv.org/pdf/1312.6082.pdf

This is why visual recognition is just one of the signals you need to use to tell humans and computers apart http://googleonlinesecurity.blogspot.com/2014/04/street-view...

yaroslavvb··on The Dyatlov Pass accident
The most plausible explanation is avalanche: http://en.wikipedia.org/wiki/Dyatlov_Pass_incident#Theories
yaroslavvb··on Google's Street View computer vision can beat reCAPTCHA with 99% accuracy
The 99.8% assertion comes from synthetically generated CAPTCHAs. We don't need crowd-sourced ground-truth -- we know what the true answer is because we generated it.
yaroslavvb··on How Google Cracked House Number Identification in Street View
"Off-the-shelf" approach didn't work, which is the reason we did this. And the simplifying assumptions are things we could get away with while solving the task at hand.

Traditional OCR pipeline would be to use some heuristics to find line of text in an image, use some other heuristics to break line of text into candidate characters. Some candidate blobs may need to be merged to make a single character, so you use a separate character classifier pre-trained on correctly segmented characters to score the candidates, and then Viterbi/A* search on those scores to find the most likely interpretation of input.

Many problems with this -- how do you tune the heuristics? How do you recover from error in an earlier stage of pipeline? How do you get character level ground truth from image/text pairs?

With enough engineering time, you can solve those problems, but it's a lot of coding and tweaking. The point of the paper is that you can skip those steps and read OCR output directly off top layers of the network.

yaroslavvb··on Protesters vandalize Google bus, block Apple shuttle
It's not really a strawman to say that shuttles take cars off the road because Googlers commuted from the city to Silicon Valley by car before the shuttle program existed. In fact, it started as a 20% project in response to the annoying commute.
yaroslavvb··on Protesters vandalize Google bus, block Apple shuttle
By that logic, we should also put some brakes on Bart/Caltrain because it makes it easier to work in Silicon Valley and live in the city.
yaroslavvb··on Protesters vandalize Google bus, block Apple shuttle
In SF they have permission yet people still protest. Some shuttle stops even have signs which look like they've been installed by the city, ie "No Parking/Shuttle Stop 6am-10am" on Van Ness
yaroslavvb··on Protesters vandalize Google bus, block Apple shuttle
There's a lot of irrational reactions to shuttles for some reason. I live on top of a "deluxe" cheese shop in Lower Polk, and the owner yesterday told me "Google shuttles, yeah, don't like them, they are causing all those techies to move in and drive up rents." Especially strange to hear from the owner since techies are probably the ones buying their $20-$120 bottles of vinegar
yaroslavvb··on Addressing Questions about the Salesforce $1 Million Hackathon
Not sure where the controversy is coming from, given that admin clarification from 11-14-2013 said, "You could modify an existing product to integrate with Salesforce and submit that, however you'd be judged on just that component, not the pre-existing product."
yaroslavvb··on Shepard tone
My dad made an animation with a visual effect in the same spirit: https://plus.google.com/109509141493915423605/posts/jokA4Ccu...
yaroslavvb··on Google’s Dremel Makes Big Data Look Small
It's like "Delta" that makes faucets vs "Delta" that flies planes
yaroslavvb··on Google to include people's Gmail in search results
If you are in Chrome, you can do Incognito Window
yaroslavvb··on Captchas Are Becoming Ridiculous
The tricky thing is that there are hordes of automated captcha-breakers that will recycle captchas until they get something that's easy to OCR
yaroslavvb··on Why the days are numbered for Hadoop as we know it
Oops, looks-like Flume is Google-only name, open-source implementation is called Crunch -- https://issues.apache.org/jira/browse/MAPREDUCE-1849
yaroslavvb··on Why the days are numbered for Hadoop as we know it
Other technologies to watch:

1. GraphLab2. Unlike Pregel's Bulk Synchronous Parallel Model, GraphLab2 it allows non-synchronous updates, which is more efficient for approximate quantities. For instance, for AltaVista's web graph, most nodes only need to be updated couple of times, while some nodes need more than 60 updates.

2. Flume: it's an abstraction on top of MapReduce, you program as if your data is contained in Java-like containers, and it turns your program into series of regular MapReduces

3. ScalOps (http://cs.markusweimer.com/pub/2012-DataEng.pdf): that's a higher level abstraction prototyped in Yahoo Research, might get resurrected in Microsoft.

4. AllReduce

yaroslavvb··on Google launches BigQuery - Analyze big data on the cloud
This is not MapReduce though, rather an execution engine specialized for data analysis: http://research.google.com/pubs/pub36632.html
yaroslavvb··on Building high-level features using large scale unsupervised learning
They are feeding 10M frames from random YouTube videos, 1 frame per video. Only 3% of 60x60 patches from those frames contained faces
yaroslavvb··on Tenzing: A SQL Implementation On The MapReduce Framework
Dremel aka BigQuery has a dedicated execution engine, roughly an order of magnitude faster than MapReduce for typical SQL queries
yaroslavvb··on How (not) to set a timeout on a computation in Python
Another problem with SIGALARM is that there's only one of them, so setting second alarm will reset the previous counter, and make the system try to run both handlers at the same time on trigger
yaroslavvb··on United States loses AAA credit rating from S&P
Just because the government has the printing press doesn't mean it won't default. When China downgraded US ratings last week they said it's because "neither the Democratic Party nor Republican Party has shown any consideration for the general interest in order to argue for their own partisan interest; they had a hard time making the correct choice in a timely manner"
yaroslavvb··on How I explained MapReduce to my Wife?
Suppose you want to count number of books for each author. Mappers take books from shelves and bring them to shufflers, shufflers make piles of books, one pile per author, reducers compute size of each pile
yaroslavvb··on Reinteract: a system for interactive experimentation with Python
Looks like Mathematica notebook
yaroslavvb··on Google reinstates account of thomasmonopoly
That's the trouble with child porn...when something is defined as "I know it when I see it", there's going to be a lot of cases incorrectly classified. This reminds me of Walmart child porn case http://trueslant.com/KashmirHill/2009/09/22/in-defense-of-wa...

To just highlight the difficulty of "what is child porn?" problem, it wasn't just Walmart's officials, but local police who took the complaint, prosecutor that initiated the case and probably a number of other officials in the pipeline who made incorrect determination

yaroslavvb··on A note to Google recruiters (and on Google hiring practices)
I also work at google, and also wonder what your bad allocation experiences are. A friend of mine started on Android team, didn't like it, and transferred to Google Books 5 months later. I think you are only supposed to transfer once every 1.5 years, but there's leeway to accomodate for bad allocations.
yaroslavvb··on A note to Google recruiters (and on Google hiring practices)
This approach can go both ways, I was hired despite just having a Bachelor's from a 3rd tier US university
← PreviousPage 2 of 3Next →