The google we grew up with was a tool that allowed you to precisely retrieve authoritative writings related to a subject, but today's google is a lot more like ChatGPT than that.
The google we grew up with was a tool that allowed you to precisely retrieve authoritative writings related to a subject, but today's google is a lot more like ChatGPT than that.
- I Googled "speed work running", wanting some workout suggestions. I visited four or five results that either didn't have concrete suggestions, were poorly written, or overly verbose. I ended my search with little faith that I was getting good advice. I typed the same thing into ChatGPT and got a suggestion of 6 different workouts that all seemed coherent.
- I asked Google Home "How do I make overnight oats". It replied, "I found a result for [...], should I read it?" then "The first ingredient is Nutella". That's it, that's all it said. I tried the same search on the web, and every result Google result was spam that was 20 paragraphs of inane blog content with a recipe tacked on the end. Again, ChatGPT then gave me a sane, baseline, recipe.
- I searched Google for "how much caffeine is in coffee". It gave me a calculator that said, "40 mg" and then a suggestion an alternative search for "Q: How much caffeine is in an average cup of coffee? A: between 80 to 100 milligrams". It turns out the calculator was normalizing the caffeine content to caffeine per 100mg of coffee.
I'm using Google exactly how their product is training me to use it, but the failure modes are all consistent. Google's AI features has not real understanding of the world, so their instant answers are frequently nonsensical. And Google clearly can't filter out SEO garbage.
Even if they could filter out SEO garbage, Google's early success killed the golden goose. 99% of web publishers are publishing content for Google, not for their readers and web publishing has become a cynically commercial affair. Individual publishers have by and large moved on from the web to other creator-focused platforms. So, the quality and experience of web content is absolute rubbish. Results are filled with cookie banners, ads, signup prompts, verbose SEO filler, poor writing, lack of authority, etc.
I wish I could search like I used to but we are where we are.
Uh, I do that and I’m 37. But I don’t scroll, I click the button that says "recipe" that’s on pretty much every recipe site.
IMO this is mainly about being used to the state of recipe sites on the internet, and has nothing to do with searching.
* Bio of the author - Google SEO rewards content written by "real humans", so this is there to appease the bot.
* Giant images of overnight oats - This helps the page rank highly on Google Image Search.
* Fluffy paragraphs - "The beauty of overnight oats is that you can make them as simple or creative as you’d like. The base recipe is delicious, and filling all on its own. But if you’d like to spruce it up, you can add a variety of toppings and mix-ins, including fresh fruit, nuts, seeds, spices, jams, and more". Is that supposed to be text written for other humans? Google rewards original content, which means recipe sites need to include massive paragraphs of filler to appear like an authoritative source.
* Ads - Gotta make the money.
* More fluff "Here’s a few reasons why you should to whip up this recipe today…" - Again, this is not text written for humans.
* Repeat the fluff, pictures and ads approximately 5 times.
* Common questions "Do you have to use yogurt?" - Trying to appeal to Google Instant Answer searches.
* Recipe - Finally!
* Recipe rating - This allows the site to put a star rating and review count on Google, again, for SEO
So, 95% of the content and functionality of that page is an attempt to convince Google into sending traffic, or ads generating revenue, rather than serving the end user.
Content on the web has dramatically changed from 20 years ago because publishers are creating content for a broken machine and not for other humans. This negative feedback loop is hurting the web, and not just Google. This leaves the door open for upstarts like ChatGPT or TikTok to gain mindshare.
> Recipes can be protected under copyright law if they are accompanied by “substantial literary expression."
https://copyrightalliance.org/are-recipes-cookbooks-protecte...
So the fluff is both an attempt at that substantial literary expression, and a way to differentiate this version of the recipe from 1000 other versions of the _exact_ recipe because the recipe itself is not unqiue or copyrightable! Without it, recipe's would be penalized by the duplicate content penalty Google applies.
This often works the opposite generational direction: when a young person has a flat tire they often call a road service and wait two hours, while the old guy would install the spare himself in less than 8 minutes. But either way, flat tires are inefficiencies that are best eliminated.
I'm over 40 and do that (and the fuzzy stuff / open questions in the top comment).
In fact, most 20-somethings are much worse than this - some studies even pin them as worse with most computing use that is not about social media apps.
For example, “speed work running”, were you looking to improve sprinting? (I’m not a runner, so it almost reads to me like “working fast while on a treadmill”). Reading the rest of your example, maybe “sprint workouts” would give better results?
(40+yr old guy here, Google had always worked for me, as many people had described it working for 20yr olds. Finding the right, specific enough set of keywords had always been my “trick”. Little did I realize I subconsciously generate embeddings like a LLM to Google search)
Speed work in running is a type of training that focuses on improving your running speed and anaerobic fitness. It involves various structured workouts and drills designed to increase your stride turnover and overall pace. Speed work is essential for athletes who want to enhance their performance in races or simply become faster runners.
Common forms of speed work in running include:
Interval Training: Interval workouts involve alternating between periods of high-intensity running (fast) and rest or recovery jogging. For example, you might run fast for 1 minute and then jog or walk for 1-2 minutes to recover. This process is repeated multiple times.
Fartlek Training: Fartlek, which means "speed play" in Swedish, is an unstructured form of speed work. During a fartlek run, you mix bursts of faster running with slower paces or jogging. You can do this based on how you feel or choose landmarks as your guide.
Hill Repeats: Running uphill at a high intensity is an excellent way to build strength and speed. Hill repeat workouts involve running up a hill at a fast pace and then jogging or walking back down to recover. This is repeated for a set number of repetitions.
Tempo Runs: Tempo runs are sustained efforts at a "comfortably hard" pace, typically just below your anaerobic threshold. This helps improve your lactate threshold and race pace.
Track Workouts: These are structured speed sessions done on a running track. Common track workouts include 400-meter repeats, 800-meter repeats, and 1,600-meter repeats, where you run at a fast, consistent pace and take rest intervals between each repetition.
Strides: Strides are short, fast runs of about 100 meters that help improve running form and leg turnover. They are typically done at the end of an easy run.
It's essential to incorporate speed work into your training regimen gradually to avoid overuse injuries and adapt to the increased intensity. Make sure to warm up properly and cool down after each speed workout. Consult with a coach or experienced runner to create a personalized training plan that suits your goals and fitness level. Additionally, listen to your body and allow for sufficient recovery between speed sessions to prevent overtraining.
Also, ChatGPT is usually really bad with long prompts when I use it, and is often kind of close when I give it a short prompt. I assume part of this effect is psychological though, where you expect the short prompt to have worse results.
(Generally just confused by how ppl always claim it’s trash or it got worse but there never is a solid benchmark)
I think the problem with Google is that 'garbage' is highly subjective, and what you think is awful is actually highly engaging for a lot of users. Google has massive amounts of data about what results users click on, how much time they spend on a website, how much scrolling they do, etc. To technical people we see that as SEO 'trapping' people to push engagement rankings up, but the reality is that it's just engagement. People could leave a site of they wanted to, but they don't. They scroll through those 20 paragraphs of recipe back story. That data means Google are finding accurate results.
If anything this shows that Google's results personalization isn't taking your engagement into account. A problem that will get worse as more people block analytics.
Man, you have strong coffee. :-(
Now, i was told this, repeatedly, over and over. It's probably false, but if it is false, cite the decision that precedented this.
"site:docs.python.org/3/ term" gets much better results. (You have to add the site: qualifier, due to SEO it's no longer just enough to say "Python docs". Hence "If I wanted the docs, I would get the docs." is not as trivial to do as you suggest. Things actually went beyond annoying to downright dangerous when searches for "Python 3.10 <topic>" were drowned out by older e.g. (3.6-3.8) version SEO stuff, instead of the latest Python docs.)
Queries like "awk by example" or "egrep by example" or "jq by example" seem to hit the sweetspot between completeness and working example: they give you third-party non-commercial unaffiliated expert bloggers not stuffed with upsell links to a bootcamp. Just pure information sites.
Having mentioned all this, now I've probably inadvertently cursed those sites...
Unfortunately even before anyone started talking about "semantic web" or whatever, substring-match was declining in utility due to spam, so a substring-matcher was only as useful as it's ranking algorithm. So some cleverness in ranking is necessary, but when ranking becomes too clever, or irrelevant due to the arms race with spam, it's an almost inevitable slide from necessary rankers to mediators.
Since all the search/spam arms race is pretty much played out with giants like twitter/google anyway at this point, it is kind of hard to understand why we can't get and keep a separate and truly basic internet search engine that just works. I don't have a handy example, but in my experience even DDG & friends seem to routinely fail the "exact substring match" test for unusual phrases that I know I read verbatim last week.
What I (and I think others) really want is just substring-match tool plus a very minimal amount of extra cleverness for basic stuff like spelling variations, some fuzzing for dates/digits, directly adjacent stuff according to a simple thesaurus. Not my area of expertise, but I would guess that there's a combination of problems that prevent this tool from existing.
a) Other search companies are all trying to do fancy-but-fuzzy "semantic distance" searches to emulate Google
b) It's still true in 2023 that just keeping an up to date index is a super hard problem no matter how many crawlers are in your army
c) Despite bots everywhere, the deeper web in social-media is just too login-walled for the rule-abiding crawlers to deal with
d) There's just no money in do-what-i-say-not-what-i-mean search tools
These days I use ChatGPT though.
I usually get non-working code snippets that pass the wrong parameters.
If you're a German living in Morocco trying to understand Quebec's immigration policies in the original text you'll be fighting the search engine at every turn.
I just give an example in another HN thread on a different context [1].
I am open for a 30' session showing a lot more examples.
Regarding less experienced people searching for information, try looking for health data and see what is in the top.
You could not have chose a more affirming counterpoint of an example.
At Google, they recommend to go to https://www.google.com/advanced_search so you can write all the keywords you want. In this case I searched for: kvt reverse exploit [1] you can quickly see that google search results are inaccurate and it shows first my HN post because it was very recent. Also the recent article where it is also mentioned and added after I post it here in HN [2] and the following results doesn't end in some correct previous post. I then try again [3] with a more precise search adding kconsole and only found three results without all the keywords while the advanced search says explicitly that all these keywords should be there.
Am I missing something?
[1] https://www.google.com/search?as_q=kvt+reverse+exploit&as_ep...
[2] https://dgl.cx/2023/09/ansi-terminal-security
[3] https://www.google.com/search?as_q=kvt+reverse+exploit+kcons...
So it's not that Google doesn't work, it's that you're using the wrong tool for your use case. If you want highly technical content or precise keyword search you shouldn't use Google. It's like you're going on Tinder to find a marriage partner.
You are saying one day you use your mobile phone for calling and the next day the device only work for calling 0800-*?
I hate the new Google as much as others, but if you don't adapt your search habits for 20 years when the whole ecosystem around you has been obviously changing, that's kind of on you. Just use another search engine that fits your use case. Personally I use Kagi and I haven't touched Google for the last year at all.
These are worth paying for.
That tells me that the main problem may simply be that the BugTraq post on seclists.org isn’t indexed by Google at all.
While that’d be annoying, I think it’d more likely be a specific configuration or robots.txt issue with that site than being a general issue with how searches are performed.
FWIW, the same query in Kagi found the original as the second result, just under a blog mirror of your 10/20 comment and above the comment itself. Since Kagi sources from Bing (along with Google) results, that reinforces the theory that it’s an issue specific to Google’s crawl.
Contrary to your other replies, I do think your style of query continues to work well on Google (where indexed of course)—-and so do full sentences.
I honestly think the issues when there are problems finding things usually come more down to 25-plus years of searchable internet history accreting a lot of clutter (especially in spaces like tech where old info ages out but never gets deleted), along with cynical SEO from low value sites deliberately skewing results for as long as the site remains indexed.
Neither is a Google problem, and the recency bias you observed is arguably the best way to combat both. Ads and site promotion are another story, and the reason I’m on Kagi, but I don’t think that hits tech as hard as consumer spaces.
sighs wistfully
https://chrome.google.com/webstore/detail/verbatim-search/oc...
So I tried verbatim. It still didn't work.
Verbatim + quotation marks. Still nothing.
I guess New Google Search simply doesn't recognize the "(-)-" part of the search term. But this is characteristic of its recent performance. I can't even count the number of times Google disregarded part of my search term and gave me an inane result.
Searching "BPAP chemical", "BPAP chemistry", etc. seems to work fine.
AFAIK, Google search has always ignored parentheses and most punctuation symbols (other than ones that are special to it, like +require_term -exclude_term "...")
https://support.google.com/websearch/thread/71287971/does-go...
Because as far as I know, I’m both and I find Google just as good as ever, if not much better.
The only place I see people complaining about Google is Hacker News and certain parts of Reddit and by no means are the opinions on these sites anywhere near universally shared by most people.
And by all means, I think calling yourself highly educated but not being able to adapt slightly when a tool isn’t working right a little rich…
No results: "vector plus phallophile" / "vector plus phallophile reviews" / "vector plus phallophilereviews" / "phallophilereviews.com vector plus"
vs.
Correct results: "site:phallophilereviews.com vector plus"
Yes, I have SafeSearch turned OFF. This is happening to pretty much anything that's remotely NSFW on both google.com and amazon.com -- it's getting completely impossible to search for perfectly legal things that are NSFW.
I don't think so. I've been using the "full sentence questions" and fuzzy questions since forever, and very seldom use "site:" and other such more formal constraints you mention.
Despite that Google results have been getting worse for the past 5 years at least.
This would be true if we could exclude commercial results. SEO is absolutely destroying search quality.
for some definition of better that is not "the results are more relevant than before"
They're perfectly capable of making things better, it's just more profitable if they don't.
uhh....
Measuring whether or not results are relevant to the query is the wrong metric to use. You don't want highly relevant misinformation. You want good information.
They should drop the charade and do a directory listing instead, like the "portal" pages of yore.