Raise your hand if you want to go back to AltaVista/AskJeeves.
Raise your hand if you want to go back to AltaVista/AskJeeves.
With Google, I feel like I have to fight with it. If I'm searching for something obscure or perhaps a word that is misspelled on purpose it thinks it knows better what I'm looking for. It also often returns searches without the word that I searched for and often ignores when I prefix it with + or put in quotes.
However, 5 minutes looking through my Google Search History, and I don't see any examples where non-verbatim results would have been useful, so hmm.
e.g. "brakes NEAR ford NEAR (problem OR issue)" would bring back results with "brake issues with fords..." as opposed to AND where all words simply appear on the page.
no other search engine at the time offered anything near this power. the amount of crap eliminated by a properly constructed query was breathtaking
You can also click on the "Search tools" button and select "Verbatim" from the "All results" dropdown. This causes the search to only perform exact matches: https://support.google.com/websearch/answer/142143?hl=en.
And yes, I also spend a lot of time fighting Google's inferral rules. I remember doing a search for Biber at one point, and it asked if I meant Bieber instead.
PageRank was one part of the reason we have a search as powerful as Google, the use of links to assign an authority score. Another big part was the use of link context -- which isn't part of PageRank but part of the overall search algorithm at Google.
PageRank score is when Google revealed those scores to the public for any page. That's not something it had to do, in order to use PageRank as part of it algorithm. But in doing so, it fueled an explosion in link spam.
But I will actually claim that PageRank itself is a problem not just because it is so gamed but also because web page authors link to things they find via Google, creating a positive feedback loop that undermines the very premise of the PageRank algorithm.
Search engine research had stagnated because of Google's dominance, and Google itself is not motivated to change, much as Microsoft was unwilling to evolve its cash cow. Rather than innovate in search, it spends most of its resources on ways to shore up its dominance (Google+, Android), fighting the very thing its search engine feeds (Internet garbage), and strengthening its advertising business.
Do you remember when Yahoo was a website about wrestling ?
Or when having a good section in DMOZ was important ?
What? Really? I was around in 1995, and I don't recall that...
I can also remember the moment when I discovered Google. It was reading this article back in 2000 (I could have sworn it was 1998):
http://www.newyorker.com/magazine/2000/05/29/search-and-depl...
If you want to travel back in time to the internet as it was when Google appeared, give it a read.
For any gripes I may have about Google's search engine (like the fact I can never seem to easily relocate this article when I want to refer to it), it definitely solved more problems than it created.
I'd like to see a hybrid of the high-speed algorithmic scanning of the pages that we see now combined with an army of human reviewers, including super-reviewers who are recognized industry experts, who periodically rate indexed content based on quality. Backlinks and other derived consensus measurements should be given far less weight and a combined algorithmic and human quality rating should be of at least equal importance.
So I want to go back to the days of a manually curated web index, combined with the technology needed to make that span out over billions of web pages.
I don't think Google have too much faith in their algorithm - they know it's flawed. But it's the least worst algorithm anyone has come up with, and adding human tweaks leads it subject to subjective bias.
Yeah, I don't think it's a bad idea. I don't have the funding to start it, of course, and VCs crap their pants at the thought of anything that has overhead, so it's probably a non-starter.
>nor that recognized industry experts want to spend the majority of their time reading a huge amount of content and rating it for pennies an hour.
The recognized experts would be paid more than pennies an hour, and they wouldn't need to spend the majority of their time reviewing content. They'd be "super-reviewers", so their opinions would hold a lot of weight. It'd be a way for them to make some extra money without a lot of overhead, something they'd do occasionally for an hour here and there. Honestly the main thing we'd be looking for from these people is information about the cutting edge; things that are trending that we haven't picked up yet, things that are new and thus don't have a lot of consensus markers but are still worth attention, and information about the perception of the content within the industry. That classification can be used to inform on a variety of axes that could be good search parameters. We'd need to make sure we got opposing industry leaders so that the index didn't become solely representative of a single viewpoint.
Normal reviewers in a position analogous to a news reporter are more affordable, more consistent, and can classify the majority of content for a sector fine. Maybe Yahoo! could take their niche content mills and reassign the staff to rate pages in their index, then maybe their search will get somewhere.
Then you'd have MTurk style reviewers who provide the bulk of the content rankings and just give back a few basic pieces of info. These are the people that would be working for $2-$3/hr or less, at their convenience.
All of this is on top of a more traditional automated ranking algorithm that would use consensus markers and computer-perceptible quality markers to rank content. There's not necessarily an obligation that every page is sampled and reviewed by a human.
It'd be great if we could get good traffic data too; we'd be able to see where people are actually going instead of just what they put links back to.
>But it's the least worst algorithm anyone has come up with, and adding human tweaks leads it subject to subjective bias.
The bias is there regardless, it's just filtered through different parameters. This is inescapable. In general, not just in algorithm and computer design, we need less faith in cold systems and more faith in human intervention and judgment.
Yes, you have to be aware that any process is subject to gaming, manipulation, or bias, but I think affording sufficient room for human opinion and circumstantial judgment as very highly-weighted inputs prevents most of the egregious failures caused by runaway systems.