Falsehoods Programmers Believe About Search (2019)
opensourceconnections.com
opensourceconnections.com
Anyway things are getting better now. More people got into search and info retrieval since I dropped this list. And there’s a great growing community out there of people who find and adore the problem space. For those who enjoy reading, I’m glad you do! For the people who don’t, … :)
Happy searchin’!
Of course the descriptivist in me would say that "reactionary" can also just mean "in reaction to" if enough people just use it that way and that people have picked up the political term through cultural osmosis and apply it to reactions in general validates this meaning, so make of that what you will.
Falsehoods Programmers Believe About Search - https://news.ycombinator.com/item?id=20039891 - May 2019 (179 comments)
Falsehood: Search can be added as a well performing feature to your existing product quickly
Falsehood: Search can be added as a well performing feature to your existing product with reasonable effort
I mean, I'm pretty sure most domain experts would tell you that about their domains.
They should have picked statements that are often true or true in certain situations so that they are "false" in the sense that they are not always true, drawing zero distinction between "mostly true/situationally true" and "completely false" in a field where the answer to most questions about system design is "it depends".
"Search can be added as a well performing feature to your existing product with reasonable effort"
if you append
", as long as your existing product is one that would benefit from having search added to it."
to the end of it
If product management think it's "just point lucene at the DB.. maybe a sprint or two", then you're in for an interesting conversation....
I don't know. If it was easy, you'd expect Atlassian make their wiki search usable.
On products, the article says it won't be easy, perform well, or give users a good impression every time. Again search quality issues.
》I would love a companion to these pieces that go into details.
They want you to give them a call, they do consultancy and the reason it's not complete.
I'm developing a next-generation search engine and the article is actually pretty light on topics, it's like search engine research got frozen around 2010. Google can't take user feedback for the recommendation without getting SEO-gamed way further, so we need a new kind of search to make it irrelevant.
For those two specific Qs you asked:
Search isn’t like a database because it isn’t ACID and shouldn’t be the point of record for data - even though lots of teams eschew this advice and use Elasticsearch or another engine as their record store…IMO content should be stored somewhere safer like a CMS, and added to the search engine for the search/discovery/recommendation use case.
Search can be plugged in to a product rather quickly, but the initial relevance is usually terrible - and it’s a beast that needs to be tamed and loved for awhile before it can make users happy.
When I use a search engine, I know I am often choosing my search terms based on suppositions about how the search engine works, but if I'm honest, I really have no idea. It usually devolves to trial and error until I find results that are close to what I wanted.
Some notes I found scribbled in his cage:
1. Never use the onsite search function. It's broken, undocumented, limited. Use google or ddg with 'site:...'
2. Learn all the search operators like intitle, inurl, etc.
3. Try different search engines. Sometimes one engine happens not to have indexed what you're looking for yet
4. Search for the text in UI elements of websites. E.g if you'm looking for a movie made in a particular year, go to IMDB and look at the part of a movie page that says the year, then search for that particular string like this: 'site:imdb.com "Made in: 1996"' you can turn almost any recurring element of a website into a tag this way.
5. Most of the above tips work best if you have a specific site to search use 'site:...' So, divide and conquer. Find the site(s) that will probably contain what you want and only then search for the thing.
Google seems to ignore these 90% of the time I try them. I have better luck with site:, but inurl is pretty much useless these days
> That hamster really learned how to use a search engine.
Unfortunately, the things that hamster learned are no longer relevant. Using + and - on keywords no longer works with Google - your luck on other search engines will vary. Sometimes adding quotes around the keyword / keyphrase functions as +, sometimes it does not. There is no reliable way to get - behaviour anymore. And even verbatim search sometimes includes synonyms.And allmighty help you if there is a product or popular media figure with the keyword you are searching. The search results will be flooded with links to some Marvel character that happens to share a name with the *nix daemon whose error messages you are searching.
You would think that really big companies would have search figured out on their own site. But Google usually works way better.
Site search gets tripped up on irrelevant false positives in some stack of fine print. Or even in a preview of a link to another page.
Or else it will totally ignore words and show partial matches high up, probably due to some terrible rank scheme trying hard to sell something.
Google's big issue for me is that it likes to show you those mirror sites that mirror reddit and github, sometimes in a way that makes it look like original content in the cache preview, but are actually boner pill ads if anyone but googlebot tries to view them.
Faceted search is a technique that involves augmenting traditional search techniques with a faceted navigation system, allowing users to narrow down search results by applying multiple filters based on faceted classification of the items