Internet search tips
gwern.net
gwern.net
The other day I searched for the Tom & Jerry full episodes on the web to no avail (streaming platforms and video platforms like youtube only have rubbish cuts).
The internet archive has every episode starting from the first one in 1940, in an easily accessible player without any ads or recommendations: https://archive.org/details/tom-and-jerry-all-114-episodes
Something I'd love to learn to do better is search WBM. I use WBM only a couple of times per month, but when I need it it's the only tool that can do the job and is therefore very valuable. Trouble is, I don't really know how to search it unless I have a record of the exact URL I want, which isn't always possible.
allintitle:Neil Diamond If You Go Away
on YouTube, and get exactly what you would think, results with all those words in the title. but now, you dont:https://youtube.com/results?search_query=allintitle:Neil+Dia...
now, I get crap like this:
Neil Diamond & Shirley Bassey - Play Me - "high quality"
Barbra Streisand - If You Go Away (Ne Me Quitte Pas)
how is that what I searched for? also, what is this:> A search for [site:nytimes.com] will work, but [site:nytimes.com] won't.
https://support.google.com/websearch/answer/2466433
did I just have a stroke? those two searches are exactly the same. I try to be understanding, but I am constantly tripping over big companies glaring software and/or documentation issues, it gets old.
Removed a few weeks ago.
Somebody posted in these pages the github diff showing the removal of the options.
I just had the exact same though while reading the page.
Makes me a firm believer companies should have their documentation on GitHub (or similar) so anyone can make a PR to tidy these things up.
> those two searches are exactly the same
Yes, that appears to be a recently-introduced typo -- the archived version from April does have a space: https://web.archive.org/web/20230412181331/https://support.g...
I submitted a feedback comment, hopefully they'll fix the typo.
(For future reference, here is a snapshot of the current version without a space: https://web.archive.org/web/20230722085337/https://support.g...)
In a few other languages I checked it has :)
Some libgen clone sites like z-lib have fulltext search on books with support for exact matches: https://zlibrary-asia.se/fulltext/?q=%22frank+sinatra%22&typ...
Even if you are going to purchase book on a subject, this finds so much stuff that is not in Google because of copyright delisting and is sometimes useful in knowing which books to consider.
Yandex results especially the non English ones can be good if you are willing to use a translator.
Too often I see in my students work evidence of lazy searching. It is as if they expect Google to be able to read their minds or even foresee their future intentions. Lack of variety of search terms is a key shortcoming. Also, lack of exploration of terms which are tangentially related to their search topic.
An internet search should be playful and exploratory. Above all, it should be understood that the internet is beyond simple linear indexing.
That's exactly what Google and other interests want --- gullible, uncritical sheeple that can be exploited to extract $$$ and worse.
I think there is / was an official standard or name, I can't remember it though. It never worked with Google really, but it worked on the search engines for our university and other academic sites.
You can use two numbers separated by two dots to represent all numbers in the range.
For example, a search for
taki 100..200
gives me results for taki 183.This is useful when you can't remember an exact year or number.
Internet Search Tips - https://news.ycombinator.com/item?id=26847596 - April 2021 (77 comments)
Internet Search Tips - https://news.ycombinator.com/item?id=18666574 - Dec 2018 (28 comments)
The forced R and L margins suck for people with oldster eyes who have to increase the text size.
Kids these days!
He did mention IA searches, but I didn't spot a mention of their https://scholar.archive.org/ with "over 25 million research articles and other scholarly documents preserved in the Internet Archive."
For book metadata, I find https://openlibrary.org/ search has a lot of 'MARC' type data. Esp useful for books with many editions.
Some services are getting more restrictive. e.g. WorldCat recently got harder to use, rejecting many searches. But if you can find a book's OCLC ###, then https://www.worldcat.org/oclc/### works every time. Useful for finding local hardcopy if you've got your location turned on. (With an ISBN , WPedia will also do this for you at: https://en.wikipedia.org/wiki/Special:BookSources/ )
Can anyone recommend a site that tracks those changes?
I find myself getting annoyed by serps of a given vendor and I might even be adopting to changes but ultimately jumping ship towards the next best thing if I can’t figure out a very low-effort way of influencing the results in my favor using flags that suddenly stop working - all I notice is a significant degradation of result quality.
But that is a great reposte for "prompt engineer"; my coworker and I joke about being prompt engineers, but really we are technical people putting new tools to use.
https://web.archive.org/web/20191201105759/http://search.lor...
Gwern's site somehow reminds me of this Fravia's site.
https://web.archive.org/web/20000620171446if_/http://www.sea...
Thanks very much for this; I'm very much interested in searching techniques, but wasn't aware of this site.
I have found a few obscure publications at the [Defense Technical Information Center](https://discover.dtic.mil).
Google should be quaking, and pissing themselves. It won't be their bardum, and they won't have the financial power to purchase said LLM.
Poof. Goodbye useless SEO spam, advertisement hell. We hated your guts.
You think LLMs won't also be used to generate more SEO spam? I think this will be an arms race, or a game of LLM-cat and LLM-mouse.
That sounds great, but are you sure that there won't be legions of LLMs generating SEO spam and advertisement hell in warfare with the LLMs that do search?
Just as an aside: Sci-fi seems to have one superintelligence (one Skynet) fighting humankind. If superintelligence emerges, I think it's likely that we'll have thousands, millions, or billions of superintelligences fighting each other, each with its own agenda.
And for superintelligence, let's clarify that this "Artificial Intelligence" moniker has almost nothing to do with actual intelligence, so we're probably a few years off from that.
LLMs are already being used to make the web even less useful, by shitting out vast ammounts of meaningless and even outright wrong text for SEO purposes. And don't forget the systems are being trained on the web in the first place… using a LLM that is able to utilize a web search already thinks listicles are useful information and not just a way to place affiliate links.
When the hype is over and the VC money dried out, companies will find ways to make the LLM interfaces and outputs an 'advertising friendly' affair.
On the contrary, my experience of the Internet since pre-web is those companies have gotten so powerful because now everyone 'creating content' wants to get paid.
Put another way, the gwerns of the world may start by posting ad-free content just because they feel like sharing. "Those" companies can't profit from individual pamphleteering.
But usually, as soon as a gwern has a measurable readership, they imagine money, and start supporting the ad industry. Surprisingly quickly, chasing more ad revenue becomes the point of their content instead of just sharing whatever they had to say.
Today, it seems most people start by wanting to get paid, and come up with a type of content to create.
Curiously, and as a result of the type of targeted searching gwern describes here, you'll generally find the best content is that which is still published free (self hosted or as open papers) thanks to motivations of the content's creator.
Also Telegram is like second internet, nobody talks about!
Do you have any groups/channels that you can recommend?
Both provide decent coverage of global happenings.
Most of the stuff that comes up on search is spam or garbage though, agreed. I’ve found the majority of news groups I follow through twitter.
It seems like 95-99% of the content you archive will be junk that you never look at again, but that 1-4% really might be worth it. Especially if it gets taken down.
What do you do for storage, though? I haven't invested in a NAS or personal storage over 1TB.
I think I understand now why some of these archive utilities offer you the option to upload to the internet archive in addition to storing locally. You can build up a local cache and then start removing the oldest page snapshots (but keep the link trail) once you start to run out of space.