Internet Search Tips
gwern.net
gwern.net
If you have a mouse with some additional side buttons (that you don't already have mapped) I can strongly recommend mapping them to ctrl-pgup and ctrl-pgdn, so you can go left/right with your tabs in the current browser without having to relocate either your hands or your mouse pointer. Because I tend to open many links to new tabs (middle-click) - a consequence of living in a country (Australia) with poor internet speeds - I'm now lost whenever I have to use a mouse without those buttons.
Watching people use the mouse to go to the hamburger menu to recover a just-closed-tab, or to the tab bar to switch between tabs, or moving the mouse to the scrollbar to move the content up/down, is the contemporary equivalent of watching someone play Solitaire intentionally poorly.
[1] https://addons.mozilla.org/en-US/firefox/addon/gesturefy/
These hotkeys are functionally the same as Ctrl+Tab and Ctrl+Shift+Tab, which - at least in my idle position - don't require moving your hands any meaningful amount.
If your hand is on your mouse, though, this means you don't have to move that hand back. (I'm assuming left-hand mouse uses won't find anything involving TAB to be hand-agnostic. : )
My usage pattern tends to be keyboard - search, tune search - then mouse - open lots of new tabs, then scroll up / down with mouse, and flick between those new tabs with the mouse, perhaps select some text with the mouse.
I think everyone agrees that frequent relocation between mouse / keyboard is sub-optimal.
Gwern’s tips for “quick searching” (which I implement using DDG !bangs and Firefox Saved Searches) are very useful too, when you know you’ll use a website regularly, I have about 30 custom saved searches now.
It came with a bunch of built-in web search keywords that sped up your search intents (all customisable of course). Usually two or three letter prefixes, that were terminated by a colon, so you could, f.e. 'ggl: khtml history' (takes you direct to first google 'lucky' result) or 'wp: charles eaton' (search and show wikipedia's top hit for charles eaton), etc.
Google is definitely far worse for obscure things than it was a few years or a decade ago. 2010 is roughly when I started noticing it.
(Google Web Search rate limits seemed to kick in around 2015 or so.)
You can still search Google (or metasearch) with bang queries, so, !S (for the Startpage Google proxy search) or !G (for google directly).
Google rate-limiting / CAPTCHA is vastly worse if you're on Tor or VPN. For the most part I simply avoid Google entirely.
It's absurd to me that this apparently independent blogger's website is the best argument I've ever seen for "JavaScript should be disabled by default."
make one of them appear and click on the gear icon top right.
I'm not sure what you mean by "can't actually click on them". Popups are positioned away from the link and do not cover up the original link, which can be clicked on like normal. (You can also click on the obvious thing in the popup, the title, which is a hyperlink to the original.)
The scroll thing is definitely bad, but fixing it is a little difficult. (GreaterWrong solves it by a hybrid of 2 listeners which tries to guess if you are mousing or scrolling instead of just on-hover.) I don't know if/when that will be fixed. In the mean time, I'm boosting the timeout since more than one person has complained about popups being too quick.
This was absolutely not the case, at least on my viewport size. FWIW, I also didn't realize about clicking on "the obvious thing in the popup", possibly because this is never a thing that you click on in websites.
> The scroll thing is definitely bad, but fixing it is a little difficult.
What's the motivation for this being a hover effect? Seems like clicking on the link to open it is a pretty well-tested paradigm.
[append]
One big difference between what I saw on your site and Wikipedia is that Wikipedia's are only previews and (AFAICT) feature no internal interactions.
Here are settings related to the feature:
To enable it from about:config, you want to set accessibility.typeaheadfind to true. The timeout after which the search bar disappears again is set as number of milliseconds in accessibility.typeaheadfind.timeout. The default of 5000 milliseconds might be excessive if you do not want the bar to be in the way during browsing. I'm very happy with 1500 for that which gives 1.5 seconds after the last keystroke to e.g. start editing the search string before the search bar disappears again.
Edit: it looks like you can enable typeaheadfind in the preferences nowadays. Tweaking the timeout still requires going to about:config, though.
1. Live in New York City. Yes, the New York Public Library seems to provide access to Illiad! I haven't actually tried making use of this, but it's on the website. Obviously you're not going to move to New York just to take advantage of this, but if you happen to already live there, you have this option!
I expect there are other cities and non-university organizations that provide access to Illiad; I mention New York just because it's one I know of. If other people know of others, I'd be glad to know of them!
2. Get a position as a "visiting scholar" at a nearby university. :) OK, this one will likely require knowing someone there, and maybe having a PhD since there may be some minimum requirements, but generally if you can get a professor to name you can one you can become a "visiting scholar" -- this isn't a job, they don't pay you anything and you're not required to do anything for them, and as such there's no hiring process, you just can get named one if you meet the minimum requirements (they may want to see some sort of CV also). They won't pay you any money but you will get library access, including ILL! So, y'know, that's useful. :)
Obviously that route isn't open to everyone either. But depending on your situation it can certainly be easier than enrolling as a student or getting a job as a professor!
Another option is that some university libraries have a membership program; you can borrow books and place interlibrary loans for an annual fee.
The website presents a graph of related works clustered by similarities.
The scite extension also works with connected papers so you can see that info there as well.
Disclaimer: I work on scite
I’m currently working on cleaning up the code and making installation as simple as pasting a GitHub release URL into the SurfingKeys settings. I hope to have this done within a week or two.
[0]: https://github.com/brookhong/Surfingkeys
Long ago I realized the the only reason I have a job is my ability to google stuff lol
Now they're terrified of not returning any results. Even when there aren't any results, they return a page full of ads that looks like results - at some point in the last dozen years, the empty "no results found" google page bit the dust.
As noted by many other people, Google's complete dominance of all web search for a decade makes the lack of any attempts of competition notable by their absence. If VCs are profit-maximizing, we should be seeing a new Cuil like, every month. Search advertising is a huge market, and is super profitable! Why isn't anyone trying to capture some of it? If Google is so bad now compared to some imagined heyday, then why is Bing also bad, despite the money Microsoft has spent on it?
On a side note, I wish the site had a simple, easy to read fonts option similar to the switchable light/dark mode.
I should remember to use the reader mode more often.
i think i would have found most of the examples gwern listed, maybe not as quickly. i go wild on google iteratively before jumping to another search engine
but, is there a tournament or contest along the lines of 'producing some result via searches' quicker than others? im thinking a form of this might exist at defcon/thotcon/similar
ironically, instead of searching, im asking here haha
Monster energy sponsoring the fastest ‘search parent directory’ surfers.
Ahhh if only.
Also while I’m at it I’d like ‘kicking stones against other stones at juuuuuust the right angle to ping against another stone while walking leashed dogs’ to become an Olympic sport.
If there is, it will definitely have deliberately horrible SEO.
Maybe that's because DDG ignores them.
"cats and dogs": Results for exact term "cats and dogs". If no results are found, we'll try to show related results.
The entire point of using quotes (or + back in the day) is to limit the results to the search term. Fluffling up the results with stuff we aren't asking for forces us to consider and disregard each one of those unasked-for results - until we get frustrated and go to Google.
Also google does the same thing all the time.
Also I just tried to get that behavior on duck duck go and it didn't do it. For my test quote, with three words that appear together on many pages but never adjacent, it just said no results found.
But if it works like google, it will tell you at the top of the page when there were no real results and it's serving fallback results. Look for that message instead of looking for a lack of results.
- There are numerous public-domain full-text archives, including Project Gutenberg, Internet Archive, and many small specialised library collections (usually focused on a given topic, e.g., Online Library of Liberty). Less useful for post-1925 materials, but often high-quality renderings (either scans or proofread re-typeset / typed-in documents) available. Google Books also allows full PDF downloads for public-domain works, generally.
- NYPL's Secretly Public Domain project has been reviewing copyright renewal records to find works published since 1923, and before 1964, whose copyright was never renenwed. Other projects (Internet Archive notably) have been flagging these works as being in the public domain, and hence freed of any download restrictions.
https://www.nypl.org/blog/2019/05/31/us-copyright-history-19...
https://www.nypl.org/blog/2018/03/30/unlocking-record-americ...
- OpenLibrary / Internet Archive increasingly have current under-copyright books available for at least 1hr and up to 14 day loan. The reader is less elegant than it had been in past, but is viable.
- HathiTrust is all but useless with its download restrictions. It's helpful to determine if records exist.
- Worldcat gets only a brief mention by Gwern. It's a union catalog (a combined library catalog of a vast number of libraries worldwide), of books, articles, and other document types, and is an excellent way of determining if a book exists, what an author's output is, and/or the documents within a given search space. !worldcat DDG bang search, "ti:" is title, "au:" is author, "kw:" is keyword. Space any colons (":") occurring within search terms, or omit them entirely. You'll still have to either find the digital record elsewhere, or track down a library, but quite useful.
- You can save online materials to the Internet Archive using the 'save' URL:
https://web.archive.org/save/<original_url>
So to save this particular HN discussion we'd specify: https://web.archive.org/save/https://news.ycombinator.com/item?id=26847596
You can submit that through any HTTP client (curl, wget, lynx, w3m, your GUI browser, etc.). Requests can be trivially scripted and batched.This is ... documented somewhere (I stumbled across it myself), though I'm not finding the specifics. Related "save page now" functionality is mentioned here: https://blog.archive.org/2019/10/23/the-wayback-machines-sav...
- Motorised paper cutters are available at some photocopy shops. Inquire as to whether or not you can have books debinded by them. (Generally anything resembling paper is fine, though the blades can be damaged by metal or other materials.)
The image search is especially impressive. Remember when Google Images used to give you actual results when trying to find the source of an obscure image? Yandex still does, and it does a bunch of other neat things too, like automatically trying to transcribe text from an image if it's text-heavy. My instinct is that a lot of this capacity exists in Google Images, but is either mostly hidden from the user or deliberately hobbled to stop the oh so evil content pirates.
Zero privacy of course. Assume the Russian government is watching in realtime as you hammer in another inane search. But for some use cases that's fine.
I thought it was a really dumb name at the time, but I remembered it some 20 years later.
I use a text-only browser to read HTML so what I am describing here is not designed with "modern" browsers in mind.
The approach I take to avoid search result limits is that I search from the command line and store the result URLs in standardised "search results files" in a "search directory", one file per unique query. When I reach the results limit for the particular search engine, I can repeat the search on another search engine. The search results file is created for two reasons: 1. it allows me to strip out all the cruft from SERPs and mix results from different search engines into one file that looks great in a text-only browser, and 2. it tells me where I left off for each search engine, so I can continue any search at a later time, i.e., get more results.
The script reads from the search directory and presents me with a menu of numbered searches. I continue a search at any time by selecting a number. For example:
1 this is an example
2 foo
3 foo bar
4 foo bar baz
I typically browse results by pointing the text-only browser at the search directory.The search result files are each named according to the URL-encoded search string. The file format is very simple. There are three type of lines: 1. a title tag for the search query, 2. a link for each result URL, and 3. an HTML comment for each HTTP request, indicating to the script where I left off, i.e., the last result number; this is more or less equivalent to a "Next page" or "More results" link. Each result URL and comment is prefixed with a search engine identifier, i.e., a prefix. Thus the script can read the search results file comments and I can easily see which results came from XYZ search engine versus ABC search engine.
<title>this is an example</title>
<!-- X q=this%20is%20an%20example -->
X <a href="https://example.com">https://example.com</a><br>
<!-- X q=this%20is%20and%20example&p=2 -->
<!-- A q=this%20is%20an%20example -->
A <a href="https://example.net/index.html">https://example.net/index.html</a><br>
<!-- A q=this%20is%20an%20example&start=2 -->
This example search results file above shows 1 result retrieved from XYZ search engine (prefix "X") and another result retrieved from ABC search engine (prefix "A"). It shows comments indicating where to continue the search for each search engine. Normally there would be around 50-100 result URLs per HTTP request.This approach assumes the search result limits are temporal, i.e., they are limits on how many results can be retrieved in a given period. That may not be true for every search engine.
Surely this is not suggesting that Google searches afford any privacy. ;)
While it might not be the Russian government who is watching, Google searches are certainly not ephemeral nor free from analysis in real-time. Aside from "things important to Russian politics", I would guess Google on behalf of its customers, who could be anyone, including governments, is far more interested in what someone is searching than the Russian government.
The point I am getting at here is that there is privacy from a government and there is privacy from a company. Each is a form of privacy, but only the company is in the business of commercialising the information it derives from violating privacy. (Not to mention that, at least in the US, the government is subject to a body of privacy law that does not apply to companies.)
Both the government and the company may violate privacy in the interests of staying in power. They could, e.g., suppress certain information when it is in their interests. However only the ad services company has the additional motivation to collect information to generate profits. The government may believe it has no choice but to monitor web search. The company OTOH freely chooses to monitor web search, as a business.
Anyway, here is a question I have about Yandex.
Google, Bing and other search engines such as DuckDuckGo now limit the number of results that can be retrieved in one session. Based on personal observation, with Google the ceiling is currently 300, with Bing and DDG, it's more like 250.^1 I wonder if Yandex is doing the same.
These limits by Google, Bing, DDG, etc. are eliminating "discoverability" via searching the web. "General" searches that would yield more than 300 results will not return more than 300 results. Users collecting large numbers of results from a general search is effectively prohibited. Google's idea of discovery is "I'm feeling lucky". Of all the silly changes Google has made to search, the most useless feature persists.
Perhaps "broad" searches do not benefit an online ad services business as much as more specific searches do. Limiting total results also puts more pressure on websites to try to be listed within the first 300. (Solution: Buy ads from Google.) A website who is at position 301 is undiscoverable thanks to Google's inexplicable truncation. Interestingly, on some, perhaps all, of their different "UI's", Google no longer numbers results.
1. To illustrate the truncation, try a search for a common string that would appear in the <title> tag of more than 300 pages on the web. https://www.google.com/search?q=title:[common string]&num=100&filter=0
Russia uses search data for things like applying pressure to political dissidents, blackmailing family members, harassing and attacking journalists etc. You can't easily opt out of this if you or your family is in Russia.
I consider the Google level of privacy completely acceptable, the Russian one not acceptable.
Depends on who you want privacy from.
Google's entire business model relies on them being the only company on the planet who knows as much about me (or you, or any other reader) as they do. I know that anything I do while using their services is going to be logged and analysed by machines like crazy for the rest of time, but I'm also reasonably confident that they're not going to let the interns page through this data whenever they get bored, and I'm very confident that they're going to do everything in their power to prevent my complete browser history from ending up for sale on some data breach forum.
[1] Those who are interested in a 'Search Engine Wall of Shame', URL for the discussion in my profile(#207).