The Vindication of Ask Jeeves
theatlantic.com
theatlantic.com
AI search can’t solve that problem. What it can do is give up on the idea of surfacing specific sources in favor of synthesizing answers from slurry of all the content, finding (or hallucinating) signal in the noise.
If you want strict keywords and boolean operators, you can’t expect to find that from the universal search engines that you’re used to. You want to look for and support curated engines that aggressively restrict what they index so that it contains high signal and low noise. Internal walled garden searches do this. Marginalia does this. Others probably do too. That’s what you’re looking for.
The reason I doubt this is because right now, I can find what I'm looking for through keyword searches on search engines that allow that, where I very often can't through the likes of Google.
The trouble is there's no way for me to know how true this noise theory is. I'm given no option by these search engines to "just show me all the crap." What search engines do today is so filtered and heuristical that there are apparently phrases and ideas that have never been expressed on the internet ever, or at least that's the impression I'd get if I was more naive.
Besides, it's not like there isn't noise anyway. Ever click on the top results only to find that nothing matches your query and probably never did?
Google tells us this is for the best. We are forced to take their word for it since their algorithms are not public. Google tells us this is also for the best.
Just figure out how not to index the endless content generation scripts, and places like Reddit that are 100s of millions of pages themselves. Or the content farms that are just copies of other sites.
Dead internet prophecy is coming true.
Go check search.marginalia.nu. Do some sample searches about history or linux or git or something and see how it often outclasses Google.
Then ponder the fact that it is run by one Swede and it runs on one (admittedly beefy but nothing unusual) workstation in his living room.
Personally at least I have seen enough to have made up my mind that it isn't impossible to do better than Google, they just don't want to.
That said I think my point still stands: If 1 Swede and 1 server can do a much better job than Google in significant niches such as history, Linux and Git, then maybe a hundred Swedes and 10 000 servers can outcompete Google across the board for traditional search, leaving Google to deal with plane tickets and local search? ; - )
And no, I'm not a Swede.
Other times, I want an intelligent system that figures out what I really meant and take care of the details. This is like having a chauffeur.
Which is better? It depends. But Google seems to think that only chauffeur style is important.
Natural Language is great for loosely defining what you want, whilst literal searches are great for finding the exact occurrence of something you know you are looking for.
I can imagine something like:
Search for the error code: "XY123", triggered by product "ABC" when I try and start it bound to port 123. I'm using version 1.0.0, and want results from the official documentation or sites of good repute. Only include solved problems, and give me references in your results.
Granted, a lot of that could be covered by special incantations in a more literal search, but that just feels far more flexible.
Note I'm trying to delineate between literal (in quotes) and Natural Language in this example if it's not clear.
When I'm searching the web, I'm very rarely doing that sort of search. I'm more usually looking for a variety of websites that talk about a thing, not for a specific answer.
I suspect it's very difficult to properly define a machine learning objective to specifically optimize for the long-tail. Without very careful crafting, I suspect learn-to-rank optimizes more evenly, resulting in poorer ranking performance in the difficult long-tail.
Back 20 years ago when I worked on web search infra at Google, a bunch of the "do what I mean, not what I say" features were disabled if you used any of the search operators in your query. If I recall correctly, it even used its "Kansas" database of per-user data to check if you'd recently used any search operators. Most of that would be very difficult and/or costly to implement with learn-to-rank, though, I guess you could pass a "power_user" boolean to learn-to-rank to try and regain some of the lost functionality.
In fact, as of November 1997, only one of the top four commercial search engines finds itself (returns its own search page in response to its name in the top ten results). [0]
Maybe I was a bit loose with my definition of "Early Google", but in the 2003-2010 time frame, the major search engines had pretty well figured out the most common 95% to 99% of queries.
IDEs optimize a part of my job that I honestly could care less about (writing massive amounts of green field code) at the expense of making occasional mistakes when I'm doing my actual job (examining and tweaking existing code). This trade off is not worth it.
I'll uh - I'll keep function autocompletion though. Trying to remember whether my coworker called it `getWidget()` or `WidgetRegistry::get()` is something I don't need to waste brain cells on.
But the lie was, in an app development schedule of 2000 hours, the tool made the 6-hour window definition effort into a 1-hour effort.
So what? It was good for demo-ware showing what an app might look like. But pointless for real work. Popular among the vaporware crowd, marketing and pitching etc. But not for software engineers.
In my day job I produce a lot of infrastructural/lib tools for other developers to make use of and good syntax sugaring is worth it's weight in gold. Being able to write concise and powerful expressions without rote overhead can drastically increase dev productivity while reducing ongoing maintenance cost. If you've got a noise ratio near 50% you could probably benefit from some refactoring to kill a big chunk of whatever you're repeating so often.
What is the difference between showing someone the winning advertisement versus showing them the winning "search result".1 The advertising company makes a decision according to secret criteria which ad/search result to show. A long, successful tradition of using alphabetical or chronological sorting for information, objective criteria, is jettisoned in favour of manipulation for commercial benefit of a middleman. The advertising company gets to decide what comes first and last, according to subjective criteria. Then make it infeasible to see what is last.
Google has admitted that in the majority of cases, an appropriate ad does not exist to pair with a given search query. That is what I call the noncommercial web. It has little direct value to "tech" companies and advertisers. Yet "tech" company intermediaries want to surveil and control our access to it nonetheless. As the parent states, "most of the internet is walled off."
In sum, the advertising company, and "AI", allow people to search for web content that exists, cf. allowing people to search for whether certain web content exists. The advertising company will make "suggestions". However not all content that exists is made searchable by the advertising company. If the content has no commercial value, e.g., it is "unpopular", it is considered garbage.3
1. When we search Wikipedia, typos do not result in "guessing" by an advertising company. If something does not exist, Wikipedia does not try to guess "what we want".
2. The term "winning" applies because the advertising company makes the process an actual auction or an effective contest.
3. If the web is mostly garbage, then its users should be entitled to know that through discovery. It seems as if the discoverable web, the one optimised for the advertising company, is mostly commercial garbage anyway. It is difficult to be serious about "AI" when it is trained on mostly garbage.
Kagi comparatively feels like a search engine that serves me as the user. It's a hard difference to spot, but the results _feel_ higher quality because I've curated them to suit what I actually want to see when I search the web.
If there's a site that's unreadable due to adblock-block nonsense or is generally very low quality, I can block it and move on. If there's something that generally has good content, I can raise it to prioritize it in results, or if it's something that I know I'll want to see a lot of (like AWS documentation when relevant), I can pin it to the top of the feed.
Not hard at all to articulate... Google is not a search engine, but an ad platform.
https://freakonomics.com/podcast/is-google-getting-worse/
"It used to feel like magic. Now it can feel like a set of cheap tricks. Is the problem with Google — or with us?"
I personally will happily pay $10 a month for a search engine that works and takes search quality seriously and has its incentives aligned with my best interest.
And no, Google, Bing and DDG are utterly broken for many of the searches I do every week and I have written about it a number of times here and elsewhere since 10 years ago.
"THE COLLECTIVE LIABILITY OF Kagi AND THE INDEMNIFIED PARTIES UNDER THIS AGREEMENT WILL NOT EXCEED $500 (FIVE HUNDRED DOLLARS)."
Why not limit liability to amounts received under the agreement. If the customer spends US$1000 on searches, then Kagi breaches the agreement, the customer can only recover "$500". WTF. Why is this so one-sided. That's left as a question for the reader.
There have been attempts in the past to insulate web searchers from Google's data collection but I do not recall anyone charging fees, openly engaging in data collection themsleves and asking paying customers to agree to a ridiclously low limitation of liability.
Two examples:
https://www.macworld.com/article/204924/googlesharing.html
https://en.wikipedia.org/wiki/Scroogle
In these cases, one was trusting a potentially random person running a proxy, for lack of a better term. However no account, no credit card or other payment was required. But that is not what Kagi offers. It requires disclosing one's identity via payment and account creation. It requires assenting to data collection. No way to opt-out. And to top it off, it requires agreeing to an pernicious limitation of liability.
In theory, having a paid agreement with a search engine website operator seems like it could be useful. The customer could, e.g., have an enforceable agreement with the website operator that prohibits the operator from certain activities, e.g., collecting data about the customer's use of the website and using it for financial gain or in a manner detrimental to the customer. If the operator breaches the agreement, then she can sue the operator.
What does Kagi actually promise to do or not do under this agreement. It promises:
"We will be good stewards of any personal information you share with us and we promise not to share your data with anyone else in any way, shape or form. Kagi's entire business is funded by its users and Kagi has no intention or interest in manipulating or monetizing user information in any way."
Why is "your data" not defined. For example, data that Kagi and its partners collect from your use of Kagi is arguably not "your data", it's Kagi's data. If it's Kagi's data then it can be shared with anyone "in any way, shape of form." Further, why is "user information" not defined. Is data collected about how people use Kagi considered "user information". If not, then it can be manipulated and monetised in every way.
Beyond the two statement about being good stewards, this entire "policy", like most "tech" company privacy policies, only states what the operator already does or does not do. At the same time it makes no warranties, so any statements about what it does or does not do could be grossly misleading or even false. There could be, and almost certainly is, other relevant information about what it does with the data it collects that is not disclosed.
A common tactic is to claim data is only used to improve the "product" or "service". Yet it is never disclosed exactly how this data is used to do that. What if the customer disagrees that such usage is actually improving the product/service. It leaves enormous discretion to the operator to use collected data however it sees fit. The customer cannot grant or deny consent for using data in some specific instance. She is not even informed about the instance, e.g., "We used would like to use your data to do X."
Other than the statement above about stewardship and keeping data for itself, does the operator does make any promises that it will do or refrain from doing anything in the future. "We do X" is not a promise. "We will do X" is a promise.
How does the operator satisfy its performance under this agreement. For example, what if the website stops working. The only thing it has agreed to do is "be good stewards" of data it collects, whatever that means. Performance is satisfied by being god stewards. Nevermind actually providing a product or service. It appears the goal of the operator is to collect data. The "product" is only a means by which to do so. This collection will be financed by user fees.
As the GoogleSharing and Scroogle examples demonstrate, it is possible to provide web search without collecting data.
The FAQ for Kagi says it's https://vladimir.prelovac.com/ which might be same as this guy https://cleverplugins.com/interview-vladimir-prelovac/ who works in SEO and was:
"developing real-time web analytics service with special attention to SEO. The project is called Cleveritics"
The thing I don't understand is... how come this is even a selling point? This is an obvious feature that was loudly demanded by people since time immemorial. It's something you realize you want five minutes into using your first search engine. Yet for some reason, neither Google nor Bing nor DDG nor literally anyone else I know of has ever implemented it. It took until ~2020 for Kagi to do it, and for me it's ~40% of reason I'm a paying user (the other ~40% is because they implemented a feature I proposed on their feedback board: the ability to rewrite result URLs via regular expressions - something I wanted to avoid having to use an extension to turn "reddit.com" results into "old.reddit.com" automatically).
There is an extension I used to use which did it for Google, though it's been long enough that I don't remember what it's called. I think that it is an okay workaround, but if I need to trim out this kind of stuff, I consider it a problem with the underlying system. It would be nice to be able to import or subscribe to someone else's list of domain-level overrides. But at a certain level, if you're subscribing to Kagi, you're probably looking to avoid that stuff by default, and you should just not have to deal with it in the first place.
2) is what makes browser extension route a no-go to me: too much hassle in getting an "unlocked" Firefox on Android.
Try it out: https://en.wikipedia.org/w/index.php?title=Special:Search&se...
> Showing results for typographical error. No results found for typograpihcal error.
Search Wikipedia for "exmaple.com", a common error when typing the string "example.com". By default, the results are the pages that contain this exact string.
https://en.wikipedia.org/w/index.php?limit=500&profile=all&s...
Now try the same query with Google. To get pages with this exact string, one needs to enclose the string in quotes. By default the results will not be pages that contain the string.
"Tech" companies want people to believe "search" and "suggestions" (non-objective ranking, insertion of "popular" items, insertion of the search provider's own web properties, etc.) are the same thing. I have no issue with suggestions. Sometimes, but not all times, suggestions are helpful. However I do have an issue with people believing search and suggestions are the same thing.
IIRC, there was a point in the past where Google was ignoring quotes and there were some HN threads about the subject.
I see what you did there ("typos").
Your reasons for search usefulness declining somehow ignore the elephant in the room: SEO. Because search as a business tends to be about advertising monetization, that's the system that gets gamed.
Want better search? Or at least a differently-skewed search? Find a way to make it not be about ad money.
It's similar to how music players don't use true random algorithms for shuffling music. The truth is that what people want from "random" is "I am not so likely to hear the same song too quickly" not "true random," which would have the possibility of playing the same song twice in a row.
By the way, in Google you can put exact queries in quotes, and there's also advanced search: https://www.google.com/advanced_search
Recent example. I've got a pine64 a64 board with an ES9023 audio "POT" daughter board that runs and boots linux using armbian fine. Doesn't know about the ES9023 audio DAC, spdif and i2s(?) rca output. Ok so first thing I think of is, what kernel module is the driver for it so I know if it is even available or do I need to compile a kernel.
Sure pine64 documentation is woeful unless you want to read schematics or something so it could have been there, but it isn't. This should not be a hard question to answer with a search engine.
pine64 a64 audio es9023 linux kernel driver
Try it. Watch google do the wrong thing no matter what you do. Quote. allintext: you name it. It's nuts. Google wants to show you what it wants, not what you want. That appears to be what it is. You want to find something, get stuffed. Google is no use to find things. Google is good to show you things it chooses in the general area of your interest. Ok but that is not a "search" engine. It's a recommendation system, fine if that is what you want but a flipping diaster of orwellian proportions if that recommendation engine is the best we have for search and we haven't noted it and responded.
I'm more likely to get the answer faster posting about how garbage google search his than repeatedly trying different google "tricks" just crazy.
So yeah the answer I really want is what else can we use that actually does search?
This is what I get: https://i.imgur.com/t0roTR2.png
Which of those results are you recommending as the answer?
edit: completely and totally prepared to accept I'm being a moron. The name of the driver, the link where you got it and how you got the link from google and I will decry it loudly here.
I tried using Startpage and DDG to see if I could do any better. I mostly found people complaining that their ES9023 POT wasn't working.
PINE64's forum search returned a few topics with half-answers. Apparently a distro called "Volumio" works? There's a report of someone having success after modifying asound.conf on Debian Stretch.
This leads me to believe that there's not much useful information about the POT, or that it's being posted somewhere that doesn't get indexed.
Although I do notice you wrote 9203 instead of 9023 in your suggested search. I get slightly more 9023-specific results with 9023 there: https://i.imgur.com/V5ty2In.png
I do think maybe you're expecting a little too much of any search engine for such an obscure query.
That's not obscure imho. What kernel driver do I use for this hardware is /exactly/ the kind of search google built its reputation on back in the old days when they did search.
The shock for me is I really /expect/ that to be easy becuase so very many similar searches have been, using google search, before. I've noticed it getting steadily worse along these lines in ways it really did not used to be. I've noticed it being because it wants to recommend not search. Seems others are noticing the same.
What alternative strategies do we have for search where we used to rely on google and now cannot?
Edit: also doing a search for “kernel module es9023” yields enough hints that to me indicates that there isn’t support for it even without it being a daughter board and thus there is no module. So the reason you’re not finding an answer is because there isn’t support for it and there’s no results confirming this 100% because it’s a very esoteric niche question. That you think it’s not is perhaps an inability to estimate how many people might be trying to run audio on this combination of hardware.
If you figure it out and update this thread, I'm sure Google will have an actually useful answer for the next person.
Ok so what is the name of the module that has the driver?
I'm sure that it ain't that hard to find out. I'm sure google with some google-fu used to give you an answer to exactly this kind of question in no more than 5 minutes of googling. 10 at most. I'm sure the driver exists.
Try it. Really. Watch google repeatedly drop and change your keywords, claim it hasn't then return pages with crucial keywords absent.
It's recommending. It's not searching. There is a difference and for this kind of thing it sucks very hard. You must have examples you've seen. This isn't my first it just came up in the last week so I picked on it here.
> What makes you think there is any support for this hardware?
Because I tracked it down reading forum posts, lkml, armbian docs, etc. Timeconsuming. Good test that google was no good for any of you either.
Google does recommendation, not search. It has happened and you don't have to like it.
I think a lot of people (with Google's encouragement) don't see them as two different sorts of tasks, but they are. In an ideal world, we'd have both.
If Google (or whoever) wants to be an answer machine, that's cool. My fear is that it will be done instead of, rather than along with, a search engine.
Jeeves is a recurring character in a series of comic short stories by P.G. Wodehouse. He is a valet for the hapless gentleman who narrates the stories, and he has a miraculous gift for keeping his boorish charge out of trouble.
I wonder if chatbots will end up serving a similar role. Sitting on our shoulder and demurely asking whether the esteemed customer is quite certain that pizza should be ordered for dinner again.
me> For the remainder of this conversation, pretend to be the butler named Jeeves from Ask Jeeves.
cgpt> Very well, sir/madam. As you wish, I shall assume the persona of Jeeves from Ask Jeeves for the remainder of our conversation. How may I assist you today?
me> What is the capital of Assyria?
cgpt> I apologize, sir/madam, but Assyria is an ancient civilization that existed in the Mesopotamian region of the world, and it does not have a modern-day capital. The ancient city of Assur was the original capital of Assyria, but it has long since ceased to exist as a functioning city. If you have any further questions about ancient civilizations, I would be happy to assist you to the best of my ability.
I’d bet that the things being searched for influenced which engines were best for a given individual. It’s too bad there’s no way to see what each excelled at.