Algorithmic search is sinking
skrenta.com
skrenta.com
There are billions of webpages. Who is going to do this review?
Is someone honestly going to review http://stackoverflow.com/questions/4300234/how-might-union-f... and put it in the category of "How Union/Find data structures can be applied to Kruskal's algorithm?"?
No.
The closest thing to a editorialized web is www.dmoz.org, and that hasn't been properly updated in years (and never will be) because it failed.
Search has to be done with algorithms - there are just too many search queries to do it any other way. Udi Manber, Google’s VP of Engineering stated that 20-25% of all queries made each day have never been seen before: http://www.readwriteweb.com/archives/udi_manber_search_is_a_....
And noting circular irony one often sees, Rich Skrenta created a Yahoo knock-off in the bubble days called NewHoo and then sold it to Mozilla where it became the seed of dmoz...
Alternatively, for negative reviews, etc., use rel="nofollow".
To claim that algorithmic search is dead completely ignores the volume that Google is doing, or the fact that they are making $billions in algorithmic search and Ad placement. How much do curated places make?
Also, not to rain on anyone's parade or anything (just kidding, I'm going to rain it down) it would take decades of 10k people churning through pages to get even 1% of the new content that Google discovers daily.
You all saw the 24 hours of unique video uploaded to YouTube every minute of every day figure from a year or two ago, right? Imagine that, only text, and produced by 10x-1000x as many people at 10-1000x the volume posting to forums, newsgroups, social networking sites, blogs, etc., every minute of every day. Because of this, you can't just review a site, you have to review the content on each page of the site. That's going to kill any curated engine in the long term.
No, the solution to this problem is that GetSatisfaction et al use rel=nofollow. It's as simple as that. And arguably Google could improve its algorithm by taking negativity into account.
My case from this week was that I wanted plans for a bookcase. I searched, therefore, on "build a bookcase". There was exactly one useful link on Google's front page (a Popular Mechanics link), and the rest were regurgitated spam that I could improve on with a Markov chain algorithm.
I've read that as long as people click on ads, Google has no motivation to clean up spam, but surely this can't be the best even for Google?
[1] I haven't got a single spam in my Gmail inbox for ages, and other spam filters are pretty good as well.
Suppose you use bayesian filtering on the text surrounding the links to determine whether the connection is good or bad. With a reasonable amount of data it should be possible.
Note: I'm not an algorithms guy, I do business and strategy and a wee bit of programming, so maybe the example isn't good, but I thinkthe point is.
That's how I'd solve this particular problem though. As I said in the parent I only have cursory experience in programming, and almost none in algorithms.
http://oreilly.com/catalog/9780596529321
Naive Bayesian classifiers are just one of the more popular types; others include Support Vector Machines (SVMs), decision trees (and their relatives, random forests), and a bunch more. If you'd like to play around with some, Weka is good open source software for this:
Determining sentiment (the topic of the NYT piece) is considerably harder though, because it would allow for spammers to write negative articles about a site and link to it and negatively affect its rankings. Also, determining the tone/emotions of a piece of text is probably one of the hardest things to do with textual analysis
This could be solved by making sentiments act as a weight (i.e. a multiplier in [0, 1]). Positive sentiments would give a particular reference more weight, negative sentiments would give little to no weight. Then it would be impossible to negatively affect a site's rankings - only positively affect them. Just like now.
As for determining sentiment, it's not something I've ever tried to do, but is it really that hard ? Intuitively I would think positive and negative articles would have significantly different distributions of certain words.
As a net addict, I regularly find myself frustrated because I can't figure out how to get meaningful information out of Google instead of sites trying to sell me. And if I can't think off the top of my head of a website that will act as a relevant portal for that kind of info, then there isn't really any alternative to Google.
At least, not that I know of yet: can anyone suggest one?
Google has done amazing things for our ability to get what we want and fast, but it also is slowly eroding our independence from it and our ability to educate ourselves by other means.
Here's hoping they prove worthy stewards once they own all the information on the planet.
Now if you tell me that there is value in social search we could have a totally different discussion, but it's more about the persuasive power of personal recommendation than algorithms not working any more.
Giant swathes of Google searches are now overrun with datafog spammers. Ehow, squidoo, hubpages, wikihow, buzzle, how-wiki, ezinearticles, bukisa, wisegeek, articlesnatch, healthblurbs, associatedcontent - all thee and thousands more domains filled with spam semi-automatically generated by legions of Indians for a few cents per page.
There's not one word of useful information on any of those domains. But apparently they serve a lot of ads for Google, so they don't get delisted.
As an online discussion about PROGRAMMING grows longer, the probability of a comparison involving outsourcing or Indians approaches 1, if Godwin’s law has not already been satisfied
"As an online discussion about PROGRAMMING grows longer, the probability of a discussion of traditional medicine and spiritual beliefs surrounding childbirth in ancient sub Saharan African tribes approaches 1."
Godwin put forth the sarcastic observation that, given enough time, all discussions—regardless of topic or scope—inevitably end up being about Hitler and the Nazis.
Google is worth billions. They could easily afford to hire 10,000 people to do this full time and it would completely transform their results.
http://scienceblogs.com/goodmath/2009/05/dembski_responds.ph... http://scienceblogs.com/goodmath/2009/12/id_garbage_csi_as_n... http://scienceblogs.com/goodmath/2009/08/quick_critique_demb...
I think we need better Algorithm.
In any case, go to a shopping specific (sub)site if shopping. A google search is a terrible way of find either products or retailers.
A lot of his juice comes from every page (seems to be over 10,000 according to Yahoo Site Explorer) on his site linking with good anchor text to every other page. The fact that he ranks so low (on my Google he's number 6 or so) even with this on such an easy term shows something, doesn't it?