Applying BERT models to Search
blog.google
blog.google
The "frustration" is "increasing" "when" I "have" to "quote" nearly every "word" to get Google to actually return results with what I searched for instead of what it thinks I meant to search for.
And there's the frustration, computing today tries as hard as it can to figure out what it thinks I actually meant. I don't know if it is worse that a person who knows what they want can't get it when the computer disagrees or if the computer is actually mostly right and its algorithms start to push your desires in it's own direction and whatever motive.
Facebook already does this really, radicalizing people by engineering the most dopamine-driving content to the top either towards self-obsession or an us v. them bubble.
In other words, I just want a fucking regular expression instead of our new data-science overloads ruining our minds with artificial non-intelligence (for profit).
Instead of returning search results of relevant rap songs, you get a bunch of links on drug abuse, rehabilitation centers, etc. Thanks for assuming I'm a harrowed drug addict, Google, but really I was looking for music.
In my experience, the times google seems to have totally missed the point of what I'm looking for, it's usually the times that the answer I'm looking for isn't anywhere on the web. Things like "datasheet JK45690DFS" or "Types of asphalt available for local delivery today".
I wish Google had some way to understand your query and the results well enough to just be able to say "The answer isn't available on the internet".
$$$$$$'s I'd bet... One big tech company used to let employees do it on an internal internet mirror... It wasn't cheap.
Used to be you'd get the best matches from the meta data on a page.
Now there's linear algebra both trying to determine what the meta data means and what the question means, so it's going to have grouping biases.
And do things like exclude seemingly random strings of numbers, because in the training data, that's usually trash, but for you, it's a part or serial number that you're looking for
how to keep a cat interesting
looks like that pattern generally
How do Google's engineering teams get away with this obvious error? For how much they're paid, you'd expect them to be better about things that passive observers readily notice. Don't they use their own product, anyway?
If you have to have a clever read-my-mind search, then instead of blending it into the main search, have 2 search types of 'intelligent' and 'precise'.
In Chrome, you can set your search engine to "https://www.google.com/search?q=%s&num=100&tbs=li:1" to get this behavior by default.
when are bonuses usually paid
I had great trouble finding results which were not advice for business owners.
If you're looking for more search feedback, my email is in my profile.
This, so much. There are so many search queries nowadays that have been SEO'd to hell, and there are just pages upon pages of crappy 4 paragraph articles on B2B company blogs just parroting the same information over and over again.
On one hand, it seems that Google, by aiming for the least common denominator - greatly reduces that gap.
But not really. I still think there's a good advantage to having search knowledge. I hope it stays that way.
Just for fun, I thought I would try to figure out when the next solar eclipse would be: I'll use the simple phrase "Next solar eclipse"
I'd probably try Wolfram Alpha. On DDG I would type "Next solar eclipse !wa". It returns a date "Thursday, December 26, 2019 (2 months from now)" Nice.
Next I look at plain DDG: I knew it probably wouldn't be useful, but first result is a website that calculates the next solar eclipse. It requires another click, but it gets "Dec 26, 2019" right at the top of the page. Not bad.
Now for Google, who I knew wouldn't be sure (since their interpretation of the very exact phrase would be fuzzy): In big, prominant and confident letters it reads: "July 2, 2019".... thanks Google. Wrong.
The skill in searching is no longer in using the search engine to the best of it's ability, it's in picking the right place to look. This has sort of always the case, but has become more of a skill as Google has strayed away from improving for the knowledgeable searcher.
It would be great if the system could infer the level of specificity associated with the query. Some people are just exploring a topic while others want to get to a more detailed document sooner.
This update addresses the other side of the search spectrum which is meaning. Google has a very tough job of moving the slider between exact keywords and meaning every time someone makes a search. This is a step in the right direction, but the fundamental problem still remains - Google interface is optimized for ad conversion, not user experience.
For example a search for "can you get medicine for someone pharmacy" used to just show generic information about getting a prescription filled, skipping over the "for someone" bit.
The new results understand what the query is actually asking, which is pretty impressive.
I'm kinda with you, I grew up with a ctrl-f Google so I sort of prefer that behaviour, I think because I don't want to rely on an unreliable NLP AI.
...I was going to say "but" but.. no I think I just don't want to rely on an unreliable NLP AI. It's so frustrating when it doesn't work, which is often.
But my mom doesn’t type keyword searches like I do, she types out sentence/phrase questions. Maybe the average user benefits from this stuff?
They hired a thousand engineers to guess what I like when I already followed the people I want to see content from and I don’t (always) want it in some random order they think is best, mixed with liked tweets they never intended to publicly share, I just want a chronological list of actual Tweets from the people I followed again.
Not a dice roll the AI will get it right 51% of the time.
We did a "Semantic Similarity search" for some documents, where we represent a document as a vector using BERT, and had to look for documents close to a reference document.
The results where breathtaking. It really returned semantically similar documents. You can do it now using ElasticSearch(But you really should do it using Vespa.ai, it is much faster https://github.com/jobergum/dense-vector-ranking-performance )
If anyone is interested in hacking around with BERT, I work on an open-source project called Cortex that handles model deployment, and we have full tutorial for deploying a sentiment classifier using BERT quickly and easily: https://github.com/cortexlabs/cortex/tree/master/examples/se...
Note the quote "when it comes to ranking results, BERT will help Search better understand one in 10 searches". This is because of the "keywordese" point they noted earlier in the article. Most searches are 1 or 2 words - there isn't enough to grab onto for meaningful ranking with short queries and a similarity function for longer text documents.
Also, try keeping the systems afloat to handle search like this. BERT is not practical to use for search results by anyone without the scale of a company like Google. You need to have a server farm of GPUs to translate all your documents into tensors - and then keep them around somehow! A document of 10k text will balloon to ~1MB when converted to a multitoken vector representation. BERT uncased has 768 features - thats 768 floats per token you need to keep around. If you compress it using PCA or averaging across tokens, you lose all the juicy context that you need for the matching and ranking. Also, there currently isnt a good way to keep this stuff around yet (though there are active projects ongoing to get this into Lucene [1],[2])
I think this is definitely a great achievement in NLP - but it needs breakthroughs in other areas to be useable by product teams implementing search, with any reasonably large content size.
[1] https://arxiv.org/abs/1910.10208 & https://github.com/castorini/anserini/blob/master/docs/appro... [2] https://github.com/o19s/hangry
Inverted indices are very efficient. How much of that can you give up at what trade off? If I’m only going to be better for 10% of queries, is that a cost effective solution? What if I spend the same amount of time tuning a traditional engine a bit more and get better accuracy for 5% of queries? Tradoffs rule the world of practical search implementations.
The sniff test is if a person can’t do it, then a model can’t either. Lots of queries look fine for matching, but you really have no idea what the intent or information need of the searcher is.
Try for example to type "when was jfk born" in Google, you should see a factual answer fished from a kg.
Funny thing, I've seen people mention right here on HN that for DuckDuckGo you need to adjust the style of your queries. This notion puzzles me, probably because I kept the habit of ‘keywordese’ from the olden days. Most of the time, results are about the same for me in Google as they are in DDG and even in Yandex—with the exception that Google is better at grouping related or similar results, and also if there's one or two sources having the search phrase almost-verbatim then they're at the top in Google. Apparently, I already need to learn talking to the site like it's self-aware, to regret ditching it.
Now, if some wondertech helps me to home in on the answer to my exact software or programming troubles instead of hundreds of vaguely related SO posts—I could really dig that.
The Transformer architecture itself stays mostly unchanged in the 2 years after it has been proposed, and with BERT/variants, most (competitive) NLP models are now Transformer based, it makes sense to make custom chips to just run Transformers, the same as CNNs.
Furthermore, tpus are a moving target themselves: as ML needs change, the team build new operations and optimizations into the next generation of chips.
For each of the scenarios they described they are just like "here's potential hard search query, and BERT adds magic language understanding which makes it all better ". It's non-obvious how BERT is actually being used though, especially at the scale and latency they need.
(I get that that this is Google's "secret sauce" and they might not saying anything in this particular use of BERT. But I'm curious if anyone had seen anything related.)
For those who haven't noticed, this is the box that shows up under a search result when you return to google from the page you've gone to. 50% of the time I click the back button, this freaking box shows up, the whole page shifts, and I click the wrong link.
Oh and did I mention I never ever, not even once, actually use it?
Great, so the search results are going to get even worse?
"Google says it can now offer more relevant results for about one in 10 searches in the U.S. in English"
The method to combat that is to start wrapping every word in quotes which is annoying. I'm sure intermediate users will catch on at some point, start doing it too and Google will drop the quote modifier. Let's keep it a secret so that happens later rather than sooner.
There's no reason to think that they will enact this to the detriment of other queries. It _could_ happen, but I am skeptical - optimistic even. As they mention, this improves ~10% of queries. The other 90% likely represent different forms of query input and I would hope remain unaffected.
Could you perhaps expand with a few supporting details for your thesis?
But we'll see, maybe this will actually be better?
also, maybe we should stop googling in "kewordese"?
Top results:
* Managing Your Feelings Without Substances ...
* Why We Should Treat, Not Blame Addicts Struggling to Get ...
Google missed the ball here, the word substance in this case is not about drugs.
Best example is ads targeted to keywords on the page vs shows ads based on past purchases.
However for the opposite case, when you are trying to find something highly specific, even down to an exact substring match I find the results to be very poor.
I suppose you'll just have to make your own...
I was after a specific legacy driver last night which the vendor no longer has available. Google returns zero results for this, bing returned a few relevent results (but sadly still didn't help me get what I needed) in the end I went mooching through the way back machine.
It's this class of search that bugs me the most, I know for a fact something with that exact filename is out there on the web, I just can't find where easily.
If I run into more issues or can recall anything else, I'll forward it on.
https://news.ycombinator.com/newsguidelines.html
The comments below are much better.
On that website, they keep using their technology from 2008 and let me use it to search for what I want. I've had enough
What people don't seem to realize is, as much as you think Google has changed, the Web has changed even more. If you kept Google the same as it was ten years ago, your results would be far far more full of irrelevent, SEO'ed, spammed up content today.
$ host -t mx vuln.ninja
vuln.ninja mail is handled by 10 alt3.aspmx.l.google.com.
vuln.ninja mail is handled by 1 aspmx.l.google.com.
vuln.ninja mail is handled by 5 alt1.aspmx.l.google.com.
vuln.ninja mail is handled by 5 alt2.aspmx.l.google.com.
vuln.ninja mail is handled by 10 alt4.aspmx.l.google.com.
Maybe you should try running a bit faster.