What Google search isn’t showing you
newyorker.com
newyorker.com
Teclis[0] and Wiby[1] are similar contenders in the non-commercial search space.
[0]: http://teclis.com/
[1]: https://wiby.me/
What the article mentioned about the influence of PageRank definitely rings true. An interesting variation is used by the Secret Search Engine Lab's CashRank algorithm[2].
[2]: http://www.secretsearchenginelabs.com/tech/cashrank.php
I listed a bunch of engines with their own indexes, FWIW:
https://seirdy.one/2021/03/10/search-engines-with-own-indexe...
run search startup, Breeze, so have a similar list and I was like damn :)
I think what you're doing with Breeze seems interesting, but the value-add of making it a commercial offering isn't clear; what does it offer that anyone else can't easily replicate with Google Custom Search? I'm not saying that there isn't a value add, I'm only saying that it isn't obvious as a user.
Something has to be a scarce resource; my guess is that the resource here is "labor" in building the CSE parameters and finding the sites to add to collections. Perhaps the effort that went into this should be emphasized.
Are there plans to move things server-side? Making client-side requests to Google has privacy implications.
tl;dr -- working on improving client-side privacy protection & there's couple of other options that may permit using Google with a slight bit of user config; premium version is all server-side since proxied via Bing, Gigablast &/or our index; there's an alt premium option that would basically proxy google via a cloud browser, say browserling or KASM or similar
1. The client-side will be set to no personalized ads -- just found out about a fairly hidden setting that permits that -- should be changed / live later today
2. Google's API precludes making server-side unless user configs an account which we can then drop in, since limited to 10K queries / day to do anything meaningful custom or full web; we're planning to do that as blog post as intermediate alt
3. We've started to instrument if client-side calls sidestep any privacy -- those results / protections are necessarily limited, however, we can perhaps uncover some things if anything in their client-side code that is privacy revealing & possibly mitigate some
4. premium is server-side with proxy to {ahrefs*, Bing, Gigablast, etc.} OR our Breeze index -- we scrape inventory-sensitive / time-sensitive sites, e.g., used car dealer pages
5. given enough ad or other revenue, we could alternatively give everyone a Bing proxy like DDG or other services -- bit too bootstrapped to do that out of the gate, plus Bing has more constraints on custom search, so that's a mixed bag of outcomes
6. another approach, also necessarily premium, would be to proxy through a service like browserling for Google searches, since the API doesn't permit it at scale
7. premium also includes alerts, especially for things like say car dealer pages that are more time sensitive and harder to config than what free google alerts do
tl;dr -- free version makes it easy for anyone to build, use & share custom searches; premium version is zero ads + better web alerts + easier access to alt web indexes
1. average user unlikely to go through config of custom search
2. even if they did, it's so poorly documented & finicky, they'd be unlikely to achieve comparable outcome
3. Bing CSE is even worse -- requires Azure account, limited to 400 up/down boosts, etc.
4. there's other technical reasons a user might not do more than a couple, along with the sheer scale issues that also mentioned
5. we're refactoring how the topic searches are experienced and making the topics, which we call branches, a more natural part of search experience
6. "blogs" is the first one to be done that way -- you can filter to just blogs after searching for X - any other branches / topics will be added in that way
7. e.g., podcasts & RSS feeds are coming out soon, along with more traditional filters such as shopping
8. that approach also makes it easy for us to expose what we're calling a low-code builder to let anyone build a custom search and either share / make public or keep private
9. that includes all possible filters - advanced keyword combos, site inclusion / exclusion, URL patterns, schema structure, etc.
10. premium for users includes zero ads, alerts, and some other features that are mix of TBD or too early to build
11. premium for teams includes similar things, along with the ability to config dashboards of searches, e.g., an HR dashboard of relevant custom searches, etc.
12. we're navigating that labor balance -- e.g., the blogs filter is fairly basic atm, whereas filtering college scholarships was a bit more nuanced to make it work
if added anything else, it's that we're positioning Breeze more as a search client + search engine, e.g., we use Google, Bing, etc. for webw-wide (search client) whereas we scrape car dealer pages for real-time alerts of inventory (search engine)
that same search client philosophy also why we're adding a low-code query builder so that anyone can build a really extended query, aka, custom search engine, and either keep private for their use or share with community, since we can't possibly build all the CSEs ourselves
in that sense, our long-term trajectory is about building what amounts to a deep reddit, where people can search sites / topics of interest in very direct ways that are way more substantive than wrestling Bing / Google
that's also why we're shifting away from /topics at the top of our site to integrating the branches / filters / custom queries directly into the search experience, e.g., blogs is first one we've done that with that's not easily available elsewhere, and podcasts is likely next.
I'm frustrated with Google search results lately, they don't reach far enough back and are diluted with all sorts of crud.
But these alternative search engines kinda just remind me of why people started using Google. When I use them, I don't get spam, but I don't really get things I'm interested in either. It's like they don't understand what I'm really interested in and I either get no hits, or get hits on things that are just completely unrelated. It's like old-school chat software or something, complete with all the computer misunderstandings and whatnot.
Don't get me wrong, I'd love to see a lot of competition in search. I use DDG a lot. But I'm surprised at how positive people's comments are about these alt search engines, because to me they have problems on the opposite end of the spectrum. To me, what Google returns is a pretty good understanding of what I want, but corrupted by manipulative spam; the other ones generally return no spam, but also not at all what I want.
So far, pretty happy with it. Before that I'd been using ddg but found myself so frequently !g that it was almost pointless.
DDG feels absolutely rubbish at local searches, but I may have just lost patience with it.
To be honest, I don't think DDG has a future purely because of the name. No way I'm telling non-techy friends that as I'll just get "what? Ducky go? Duck what? Are you being serious or is this a joke?"
Or that they don't have money to throw away on things that aren't things necessary to live. Your statement is only true if the person you respond to is quite wealthy, which is a leap.
Only as a side note/JFYI, naming something, particularly if aimed to be international/multilingual, is particularly tricky, duckduckgo may well sound funny to your friends, but - as an example - kagi (in Italy) would be pronounced the same as "cagi" which is a reknown historical maker of men underwear, to the point that "a cagi" is sometimes used as a synonym to "a tank top", and surely it would make some people think it as a joke.
Any sort of description or tagging or keywords or genre description needs third party vetting to be of any use what so ever. It's simply too profitable to misrepresent your websites for it to be any other way.
One could even let this vetting happen in decentralized fashion, by extending Web Annotation standards to allow for claims of the sort "this page/site includes accurate/inaccurate structured content."
In the past few months at least half of my Google searches have had spam in the top 3 results. Literally malware domains that community-made filter lists are aware of but somehow Google chooses to share anyway.
Extend this to online communities, and you can ask "what laptop would HN recommend for Linux?", etc.
Of course, there's a privacy issue to solve, but the functionality could be very useful compared to the crappy Google search results for commercial products.
To avoid sharing sensitive page / habit, maybe let the user review in batch and confirm before sharing out the list.
I'll be able to support the project starting a few weeks from now, and would love some non-Patreon options.
They do take a fairly steep cut, and that's not even considering that PayPal also wants their pound of flesh on top of that :-/
There are two menus for options ("Popular Sites", "Blogocentric Eigenvector", "Both Algorithms", "Experimental", "Allow JS", "Deny JS", "Require JS"), but does not explain them very well. (For example, I might want the search engine to not execute any scripts in web pages to determine their text, but if the text works when scripts are disabled that it can still be included in the search results (even if the web page has scripts, as long as those web pages work correctly even when scripts are disabled).)
Also, they have some documentation using Gemini format. I have a Gemini viewer in my computer, but it won't use it because of the "Content-disposition" response header.
The hotel industry is super heavily targeted with SEO spam. They're almost on par with online casinos in how much effort they put into manipulating search engines.
Places that are visibly there but not returned in search result.. It's the most bizarre thing.
Cherry-picked quotes aside, this part of the article really did make me roll my eyes. Users self-censoring searches to a _single website_ because the rest of Google’s results are so unusable is somehow a feature, not a bug.
The same way that Google improperly handling quoted searches and returning pages _without that exact string_ is somehow a feature, not a bug.
Couldn't find it now but that was refuted by Danny Sullivan (I think) in an HN thread about the Google search results quality. They gave a pretty convincing explanation of why that may happen (spoiler: it has something to do with the tokenization of words on the website) and I, for one, believed them.
*Edit:* Here it is: https://news.ycombinator.com/item?id=30356382
Looking at the source of the cached page,
https://webcache.googleusercontent.com/search?q=cache:sRJX_e...
It contains
<a href="/quotes/tag/don-t-give-up-quotes">don-t-give-up-quotes</a>,
<a href="/quotes/tag/don-t-give-up-the-fight">don-t-give-up-the-fight</a>,
After you eliminate HTML, that becomes "don-t-give-up-quotes, don-t-give-up-the-fight" and since punctuation is stripped, that matches.Full disclosure I work at Google, but not on Search.
If this is no longer possible, it's only because they're cheaping out on hardware and not actually indexing as much of the page as they could (no way to retrieve the matching token because it's not saved, but they can still match on it). This probably also leads to diluted results.
If they quote a search term they expect to find the exact term visible on the website. Not in alt text, not in invisible text, not any meta data. Quotes means this exact phrase visible on the website.
Of course, that's because my user agent renders alt text. I'm glad Google search results match to things I can see on the page.
Maybe it's just me, but I thought the middle section of the paragraph wasn't relevant to the first sentence. I speculate, with charitable interpretations, that the order of events was:
1) The reporter asked Sullivan whether the frequent use of "Reddit" in search terms indicated a problem.
2) Sullivan asserted that Google is useful for searching for results restricted to a domain, then shifted to say that Google still gives relevant results if you use Boolean search terms.
3) Potentially after a follow-up question, Sullivan conceded that much of the results that Google returns is commercial instead of community-made.
From a writing perspective, I found it unclear whether the middle section that starts with "Users on the whole have become passive, relying on Google to anticipate their desires" was analysis by the author or part of Sullivan's response. I interpreted the sentences as part of Sullivan's response from the context, though it would have been clearer if the writing more explicitly indicated whether this was part of Sullivan's reported speech.
No, Google made users passive to profit from controlling limited choices presented to them.
> I interpreted the sentences as part of Sullivan's response from the context
Looks like it to me, too. It's Google's propaganda for "we made you need us".
it's the layperson's version of recalling regex if only need it every few weeks -- lots of cognitive load for minimal return, easier to ask friends or reddit.
yes, search operators seem intuitive to perhaps the average HN reader or SO coder, etc. That's a small set of people compared to 7B+ people attempting to recall oh, right ~term adds synonyms -- easy for most people here, mostly useless or not worth remembering for most of the world.
Google's goal is to make them ultimately irrelevant and just give you what you're looking for, meeting the user halfway on the search query. That may not be possible, but it's the right mindset to have for solving the problem... Making the human talk like a computer is a UI failure.
we ran smack into that with a special topic we released this week, Ladypedia, where we ended up added a dial to filter out male results on a sliding scale for Wikipedia results. Doing it by default was too strong, and so we ended up selecting for female gender and then letting user dial in how many male-related pages to filter out.
1. Breeze -> https://breezethat.com/ 2. Ladypedia -> https://breezethat.com/p/ladypedia
No! Goodhart's Law. Karma in itself is pointless.
The reason that adding reddit to the end of a search query has better results is because you are querying a community for your results.
There are already bots on reddit that take submissions that hit the frontpage 366 days ago and re-submit, and then other bots that take the top comment and resubmit the comments to that submission. They get loads of karma doing this. Sometimes it's useful, sometimes not. But it does nothing to improve actual credibility.
The larger the subreddit, the SMALLER the community. r/AskReddit is not reliable. r/MechanicalKeyboards is. You want to find a small, tight-knit community of peers that know each-other. There are small communities of SEO experts just submitting Amazon affiliate links, so just 'small' isn't sufficient either.
You want to vet that the posters are acting in good faith. Find the top posters, find things you know(and it's better if you disagree with mainstream opinion here), and check their other posts and see if you agree with them there. If you do, it's more likely that you will agree on this subject you are referencing and don't know much about.
This process sucks - especially when you start - but over time you will accumulate a network of trusted people and communities that you can rely on. Anything else will be corrupted by greed.
We had an opensearchdescription.xml for literally most of the websites, and somehow Browsers (including Firefox and Safari) managed to mess that up.
I wish there was an easy browser extension that just persists searches correctly, and doesn't forget about them the next time I clear my browser cache.
Guess I'll have to implement it in my own one again :-/
Google is so full of content farms these days, it's ridiculous. And all of them are ranked higher than the sources because they use google ads on their pages. Just search for a quote that you read here on HN or on reddit or on SO...and you'll find hundreds of them ranked higher than the source.
Maybe someone should build a search engine that downranks all websites that have google analytics and google ads? Would certainly be an interesting experiment.
1. web, blog tab in search results is ours, and our first iteration of making it easier to dial in topics directly from web results -> https://breezethat.com/
2. early pre-web-wide topics before we added web search, https://breezethat.com/topics
3. ladypedia, an experiment on tuning out ugender bias for Women's History Month on Wikipedia pages, https://breezethat.com/p/ladypedia
4. we're currently sussing out the extent to which sites are human-ranked vs. using other machine-extracted information from a page / site -- it's a lot of rapid iteration at the moment
5. we have similar thoughts as Ahrefs on publisher profit sharing, although we haven't established if we'll hit their 90/10 mark or not. we also have option for users to go premium and skip all ads
There are a lot of useless, low-quality sites that I (so other people might find it useful) never want to see.
I want to avoid using uBlacklist for that.
1. why? - only way scales -- we can't possibly curate everything - some sort of shared gain -- token, rev share, that's WIP - core concept is what we call "deep reddit" - search all links in a subreddit
2. users will get same controls
- domains <- specific, subs, TLDs, etc. - keywords <- exact or fuzzy, akin to ~ - structure <- microformats, schema.org, etc. - format / content <- images, blogs, newsletters, video, etc. - user controls <- think form builder so don't have to remember query syntax - updates <- how often to snag updates - alerts, etc.
3. deployments
- private, unlisted, or public - branded or white labeled / embedded - basically your personalized equalizer for the web
Try Kagi.com. Also, you can manually rank or even block domains you choose.
I would suggest the correct framing here is "the rest of the internet is unusable".
How else is Google supposed to guess you want a personal opinion from Reddit other than... typing it in the search query? Google is a search engine, not God.
If you want an answer from a particular source you're going to have to specify that, Google can't know if you want Wikipedia, or Quora, or a news article or a TikTok video or one of the other thousand pages that plenty of users might be interested in.
The internet has exploded in size, so has the diversity of legitimate sources, the search engines are hardly at fault.
On a more serious side, I think, we may need to have another dimension to web search, regarding the type of origin, namely "conversational" versus "commercial" (to be applied to both text and image search). Something like this would provide a more generalized solution to the problem, instead of pushing one of the more commonly known origins to become yet another monopoly.
Well if you google "best toaster" then yes, Google will (and should) show you ads and sponsored posts and articles that have been seo-ed to death, because it's a meaningless search.
Google isn't in the business of assessing toasters' quality. Google is in the business of finding information relevant to your search.
But relevancy goes both ways: you need to search for something specific to find it.
"When will I die" will not point to anything interesting. "Pancreatic cancer symptoms" might.
When I search "best toaster" it's shorthand for "show me roundup reviews of toasters + forums where people discuss toasters + very highly reviewed toasters on shopping sites". It's not shorthand for "show me loads of ads and SEO bullshit".
There are many characteristics to any device and the "best" one depends on personal preferences. What's the "best car" for example? Does that even mean anything?
I don't think a generalist search engine should be used for shopping. You don't need Google to read reviews on Amazon and form your own opinion, again, depending on your own preferences.
People use Google as an oracle and then complain when that fails. But Google is a search engine. It's not the Pope.
I don't know. That's why I'm asking "what makes a toaster the best?". I'm trying to learn. If I get bombarded with ads and spam the search engine has failed me.
> But my point is that "best toaster" doesn't mean anything per se;
When I type "best toaster" into a search engine it does have a specific contextual meaning. It means "please help me educate myself on what makes a toaster 'best' so I can make an informed purchase".
A couple of observations since I last wrote my thoughts on this (last time i said this: https://news.ycombinator.com/item?id=30348492) :
- People who work in tech do not constitute the avg internet user now a days who uses Google to search. Naturally, google evolved to serve the lowest denominator by default, and that would probably not be satisfying for advanced users. (which opens up space for a new competitor I think?) Half the searches are when people are asking basic questions - hence they do not even result in a click. I have searched multiple times for a random calculation because google is faster than opening a calculator on mac.
- Adding site:reddit.com is a feature, and it's a utility of the engine that Google is allowing us to concentrate search to a subset. Same goes with many other operators, it's just another phrase. Ideally, there could be a wrapper which does these things for us, but I doubt it would work well. That does not mean Google search is dying, but more like we are now realizing the dependency we had on Google and the laziness when it comes to finding things on the internet.
Now people use "site:reddit.com" to limit the results to a space that is moderated by humans and not completely commercialized.
The loss of Old Google might be worse for humanity than the burning of the Library of Alexandria.
As if nothing good or the best in its breed ever came out from $PREVIOUS_YEAR or any previous years before that!
Tell me you're a garbage SEO website without telling me, just slap on "in $CURRENT_YEAR" or use some capital letters, I love how easy it is to tell the wheat from the chaff!
edit: someone with a more negative outlook/rant than my glee https://reddit.com/r/changemyview/comments/qowtws/cmv_seo_an...
It is, in fact, a 20+ year old design that's had some minor updates but is still essentially the same as the one it replaced.
But, wow! Who knew that the landscape of toastering moves this quickly and changes so much from year to year?!
Try to apply this to software (e.g., "I want the December 2018 version of Chrome, it was the best version of Chrome so far.") and observe the reaction you get.
That older one could still be the best, but this way you (or that’s how it’d be in a perfect world) know that they at least checked if anything better came out.
So fresh, so new! Except when you click, you realize the article is dated and was barely updated at all from when it was written!
I know most of these are blogspam sites and I hate them too, but the way I understand the rationale behind the year, and the way I would do it too is this:
Best $PRODUCT $CURRENT_YEAR can include products that were manufactured in $PREVIOUS_YEAR or $LONG_TIME_AGO. The $CURRENT_YEAR is saying that the list is up to date, not that all the items in the list were made in $CURRENT_YEAR.
(or alternatively, instead of "published" use "updated").
This has the added benefit that those searching for $product towards the end of a year can decide whether such lists published/updated near the beginning of the year are relevant or not.
Best toaster 2022: 12 toasters for every kitchen Published: Jan 12, 2022
For a laugh I tried "best toaster 2023" and there are already two worthy pages.
Google is the search engine, and if a business, or a thought, or a piece of information, doesn't show up there, then it's invisible to 95% of the (Western) human population that doesn't use duckduckgo or Bing or whatever failed alternatives there are.*
I do find it very amusing that I'm not the only one that uses reddit as a keyword in search results, because the results are almost spot on.
* - Yes, I use duckduckgo as my default, but use g! a lot because the results aren't superb for tech items, and it too falls prey to SEO type manipulation with dubious stuff at the top, in my experience.
Does Marginalia really prioritize sites that lack encryption?? The other things seem like good ideas, but don't we all want to encourage encryption?
But some sights just aren't... sensitive. They don't have logins, they don't have user-data stored. I don't get the point.
Does it matter if zombo.com really has an SSL cert?
I do like the ability to boost and downrank certain domains myself. Its TinyGem platform also looks like an interesting "social bookmarking" approach that's more minimal and focused than other mainstream solutions. I'm currently working on a way to POSSE[0] my public bookmarks to TinyGem.
The "lens" functionality which lets you gone searches to a subset of site is really excellent and genuinely useful. Plus you can remove sites from results, too.
I'm really optimistic about where it's going. And yes, they have a "pay for it" model which reassures me they're thinking about the future.
How long ago did you try DDG and do you have some examples of searches that didn't work on DDG but did on Google?
For starters, search is just critical to me (and I think pretty much anyone who does anything around development / design / content / ...actually any modern day desk-based work...) - so in time terms, no, I probably listen to Spotify more but in touchpoint terms I can't imagine how many searches I do in a week. I would imagine it's thousands.
Two things then follow: firstly, the quality of those results is totally critical. If (as is starting to happen with Google) I'm really struggling to get past the endless ad-based SEO'd b/s to get to the site or nugget of information I need, and each of those searches is taking me much longer and longer, then this really does become a genuine value / cost proposition. With a tool like Kagi I can say "just don't show me site X" or "boost Stack Overflow" - then this plus there are no ads taking up screen real estate becomes a viable value proposition.
Secondly: there's something to be said about supporting a viable non-ad based model for services on the web. I would argue that basing everything around clicks in an advertising way has degraded basically everything about the modern internet. It's lowered content quality, drowned out real voices in favour of bots and algorithms and causes major large scale issues with "truth". Where organisations / companies / individuals can afford to pay X and get a bigger platform with their made up b/s about vaccines / war / the opposition / guns / whatever), then this drowns out the viable, real, truthful alternatives from people who actually know what they're talking about.
Plus, of course preaching to the converted, but the tracking model being used by Google is pretty terrifying, too.
Plus... I mean, $120 a year is $0.3 a day... I think I can manage that ok...
Are SoL, as they say on the internet, because more or less any avenue where a genuine person can leave a review can also be used by a marketer to help steer you to the 'right' choice, their product.
I'm still with Bill Hicks on this one.
I’ve long thought that search should be provided by individual websites and cached/indexed by ISPs, like DNS etc.
You search for “foo”, your nearest Search Protocol Service (or called whatever) hosted on the ISP looks at all the content served by the ISP in the past which contained “foo”, then it asks those web servers for a detailed index for all occurrences of “foo” on their site.
If none found, query the next nearest search servers.
Search could and should definitely be decentralized.
While looking things up in an index partitions fairly well, search queries are typically extremely underspecified, so what you end up with is potentially millions of results.
This is why you need a search ranking, like this is not negotiable. If you don't have this, your results will be hot garbage. The crux being that this partitions in an orthogonal axis to document retreival.
If you don't solve this problem, you basically end up with YaCy. It's peer-to-peer, but it's slow as molasses and the results just aren't very good.
https://www.ted.com/talks/thomas_thwaites_how_i_built_a_toas...
To make good strudel you should heat it through consistently in a conventional over and then grill it.
How would he know whether they are/would be happy or not. The HN contigent is the only one that actually gives meaningful feedback. It seems like Google does not want that feedback.
Something about Danny Sullivan becoming a spokesperson for Google reminds me of when Microsoft bought out Mark Russinovich.
What would be good to have is an independent search engine that both works well, and which is not backed by anonymized Bing or Google results.
Google gets you nowhere near a link to that film at this moment because Google is not a search engine.
DDG has it as the 3rd result so you can decide for yourself.
It's critical to the functioning of the free market that consumers have accurate information.
The fact that Google is crap is less important. The fact that Google squashes competition and there is no unbiased search engine, is of national interest.
Current situation is because of a combination of Google's monopoly abuse and the fact that US govt does not recognise that search engines are(is) the entry point to a significant market. If you believe in the free market, then the state of Google search indicates huge inefficiency. A lot of people treat the term free market to mean "I, personally, am free to get rich, however I can, at the expense of everyone else." Which is pretty much the opposite of what it really means. It really means "you are free to join the market and compete on equal terms." Under those conditions best product/price should win.
The article still adds value as the reporter got a Google spokesperson to comment, and also brings the issue to a much wider audience outside of Hacker News.
I switched to DDG more than a year ago and I can honestly say that I've not had to revert to google search more than a handful of times since. A regular user would probably forget about Google in a breath if only they knew that such an alternative existed.
When I search for an opinion...