What's weirder to me is that it seemed like they were going with my proposed route for a pretty long time and only recently starting providing dopey answers. Maybe its part of a grander experiment they're doing to vet question answering AI?
What's weirder to me is that it seemed like they were going with my proposed route for a pretty long time and only recently starting providing dopey answers. Maybe its part of a grander experiment they're doing to vet question answering AI?
Once upon a time it was possible to identify "good" pages by how often they were linked - effectively crowd-sourcing the problem.
Natural selection: sites have learned to manipulate the game and the crowd.
Under pressure to optimize metrics that lead to SEO and better valuations, the internet is getting less useful from a user perspective.
I don't want to watch a video/slideshow, download an app, register for a forum, or read through a 2000 word fluff piece with interspersed ads and links to more information in which site has buried the one-sentence answer to a 2-second question in order to maximize my time on the site.
This garbage is what makes it to the front page of google these days. I guess the poor sites that aren't user hostile just aren't SEO'd enough.
This is likely because marc.info isn't playing the SEO game, and spamsite #43254 is.
The original search engines ranked pages by their content. Naturally this led to gaming by including keywords (remember huge invisible sections of pages that just repeated keywords thousands of times?).
Google's original PageRank algorithm was a complete breakthrough for this, almost completely disregarding content and instead ranking results based on the text other pages had used to link to the page. This was so good, in fact, that most of the other search engines from the time didn't survive.
This once again led to large-scale gaming with techniques like link farms, and a small arms race as Google came up with ways to squash these new techniques. The quasi-legitimate "SEO" industry spun out of this.
I think now we're in the same cycle again, and at this moment scam sites are winning. What's to be seen is if there's a looming breakthrough, or the arms race will continue.
CAUTION: subjective experience ahead
I feel like in the last few years Google's utility for knowledge discovery has decreased. Instead it seems to be geard more and more to drive you towards purchase and shallow content sites. I think the growing popularity of "awesome" lists and other forms of curated content discovery is also driven in part by Google's lack of "good content" discovery.
You can tell there was a paradigm shift at some point and the search engine became a second (third? fourth?) class citizen.
But if you need to replicate google, you need bootloads of user data and in generally data to operate on ( i.e. to be able to extract the equivalence of certain terms and expressions, and so on ). And generally some kind of feedback system ( like user clicking a link on a SERP, is adding that question/answer combo a certain boost for the future ).
In short: it's complicated and after 10 years of experience I couldn't tell if a solution is so unversal as google.
For example, current solutions implies there is a "document" with some metadata. But what is a document in wikipedia or SO context? So is quite hard to get a an answer which satisfies both cases and you end up with cramming all kind of data together just to make a fit. So in the end, you still have to use the google model, of using the URL as a key for (data, metadata). So in this case, just let google/bing/yandex do it's job.
On the other side, I saw Google Search Appliance fail spectacularly in enterprise, because the google model couldn't fit and you end using real custom search solutions. There are quite a lot of companies filling this niche at this moment and so I think Mozilla wouldn't make a dent, if decides to enter this market.
My line of thinking was on working with "what is the user asking" vs. "what is the user looking for" in a way that could be applied regardless of the information being searched (by, for example, a non-profit with some existing insight - hence the Mozilla thought). But I know almost nil about the space so I suspect that is somewhat naive.
If your business pitch is "Like Google search, but Google doesn't control it," good luck getting investors.
It's one fewer engineering task for a busy content provider to take on.
This first is that you were probably a power user back when good search results required you to be good at search syntax. Back when Google wasn't the only search engine on the web, you could ask Google questions. Your results would be limited in relevance. If you tailored your searches by ensuring specific words or phrases were included, excluded irrelevant phrases, and added wildcards where necessary, you'd get pages and pages of great results. Google takes your search history and tries to predict what you're actually looking for.
The second is how Google filters your results to try to tailor what is relevant to you. Most people who hang out here on HN or StackOverflow would probably get results for the software framework filtered to the top when searching for "electron". Chemists, physicists, and other scientists of the like are probably going to get results for the particle filtered to the top. Adding this type of customization probably has ramifications beyond simple terms like electron. I don't have any evidence whatsoever, but I imagine that this customization would give biased results when searching political, economic, or social topics based on the sites that are visited.
Another theory I have is that my perception of good results changed over time. It could be that for the first time in my life I know enough about the domains I'm looking up that I can differentiate the good from the bad. On the other hand even with that, when I enter a new domain, e.g. AI it still feels "less discoverable", even though there are definately good introductions out there that you can find when going beyond Google.
If Google wants to make money by being the endpoint for every question ever asked, they need to accept the consequences when they get the answers wrong.
I feel like the article addresses this. Google has basically two different products that produce those top of the page, set-aside answers. One is "Knowledge Graph", which basically does what you are suggesting - grabs straightforward answers to simple questions from Wikipedia and similar. The second is "featured snippets", which is the one causing the outrages that the article highlights.
The problem with their algorithms.... is that all that statistics in the world can't help you when you're listening to a guy telling the truth vs an equally good liar.
I can tell you what Google doesn't have, a strong AI. It thinks it knows "facts" but these are merely patterns, and these can be gamed.
Because Google still lacks a strong, truly thinking AI, they rely extremely heavily on statistical models to rate content.
So how do you cheat google search?
Google's systems attempt to figure out the topic of your writing, the style, and quality. Is it scientific? An opinion piece? News? Is it a technical topic? A playful one? Fiction or nonfiction ?
The quality classifiers are much easier to game than topic and style analysis. They determine things like reading level of text but also things like the number of rare nouns, number of technical words, number of typos. Readability as far as font and formatting. Trustworthy signal of your domain and possibly the company and people they determine to be linked with it.
I also have a feeling google uses sneakier signals as well. These include your DNS registrar, phone number, email, and address listed on the site. Who you host with and what technologies you're using. Your mail servers and how trustworthy they are. Geo location, and visitor traffic info as soon as you put analytics on the site (or use amp)
Basically when Google says they have tons of signals, they do. They have a dataset that amounts to every site on the internet for the past 15 years, and they regularly run automated and manual "theory provers" much like quants do with historical stock market data. They find new signals constantly, and run tests to see if their new algorithms are better.
You know how sometimes google randomly takes a bit longer to load search results? My tin foily theory? They'll occasionally guinea pig you on prototype search results to see if they're better. I noticed their response time getting really bad a couple months before the public rollout of new AI powered search for example.
So gaming google? Do exactly everything that a large, legitimate, no-nonsense company would do. From where you host, what you host with, to who you link to. Bonus points if you have significant real looking mail and other traffic from your domain. Extra bonus points if you actually sell something real as cover and do it for at least a few years.
Once you've done enough convince google you're a big important thing IRL...write a ton of really subversive bullshit. Make it sound a real as possible, hell make 90% of it real, just with a single unverifiable fact. Keep pumping this shit out and make sure your garbage is never fake enough to get called out on. Or just make the fake part so hard to verify that nobody will waste the time, kinda like half the science world does when publishing papers.
Did anybody think the contrary?
I want a strong AI, because I'm more a fan of the depictions along the lines of Iain Banks' post-scarcity Culture Minds.
While that's true for sites in isolation, Google published a paper a few years ago[0] that describes how you could estimate the trustworthiness of a website. The basic idea is you assign a trustworthiness score to each website. Then, you determine how likely a fact is to be true based on the trustworthiness of the sites that state that fact. You can then recalculate the trustworthiness of each site based on whether it agreed with the fact or not.