Cuil relaunches as 'cpedia.' The results are not pretty.
cpedia.com
cpedia.com
Even before the recent explosion of spablum (spam/pablum/made-for-adsense) sites, using the web to learn about something involved reading a lot of duplicate info. When learning about topic X, you might visit site #1 and learn A, B, C, D, E. Then, site #2 has B, C, D, F, G. Then, site #3 has A, C, E, G, H.
Even if H is the key thing you personally wanted to know, by the time you got to it, you'd encountered most other things 2 or more times, and even the 'gem' of a result, site #3, was 80% redundant.
A wonderful summary page would have A, B, C, D, E, F, G, and H, all in one place, properly contextualized. The best Wikipedia articles serve this role and that's why they're deservedly atop many search results. But the obvious utility of summary pages has prompted a race to create more of them, by automated and editorial means. When such attempts fall even a little bit short, they make the problem worse not better -- polluting the web with more permutations of the exact same information, hiding the rare gems in more muck.
I'd like to start a project to attack this problem head-on, and would love to hear from others who might want to collaborate.
This is probably a weakness, not a strength. If cuil were a bootstrapped one-person operation it would be very easy to cut the cord and move on to something else.
Or actually get it to work,like DuckDuckGo. One man operation. Impressive results. Contrasting DuckDuckGo and Cuil is enlightening in all kinds of ways.
Gigablast, duck-duck-go are outperforming cuil where it matters, and they're getting quite close to be being able to supplant google.
Cuils founders ought to be at least slightly embarrassed by that.
http://www.cuil.com/info/blog/2010/04/08/introducing-cpedia-...
"I was a stay at home dad, I dropped by Kleiner Perkins to try to get some money to write a search engine". This has to be an elaborate piece of satirical performance art.
Now, I am not going to try. Lesson learned.
Not really.
http://www.reddit.com/r/blog/comments/a5byc/interrobang_your...
Read literally, Paul Krill was the author of the entire article after the summary.
Because this content is pieced together by computers who have no idea of how the resulting piece reads as a whole, I wonder if there are hidden legal liabilities when words get stuffed into people's mouths or when the combinations of sections lead to false content (imagine one paragraph ending with "and found the following to be completely false:").
EDIT: Here's an article misquoting a lawyer. Not smart. ;) http://www.cpedia.com/wiki?q=palin#headline_31
Except, maybe, that mahalo's "content" is far better scraped.
I'm not sure what a crawler is supposed to conclude in that case.
I don't see a lot of cpedia content on Google yet. The only links I can see are clearly due to the threads here and on sites like Reddit.
User-agent: *
Disallow: /imgsrv
Disallow: /search
Disallow: /pref
Disallow: /suggestIt lets you in regardless of the username/password combination. I just logged in with this:
username: you morons
password: this defeats the whole damn purpose
I wonder if any of this gets logged at their end.
Here's one of the first grafs on "cuil" in cpedia:
Best Web Hosting Offer Cuil launches with an index of 120 billion web pages making it a most comprehensive magic! Affiliate Secrets search engine on the web and also a potential try to get the Google competitor. Clickbooth Cuil but not avail due to flooding traffics and making their servers 'too hot' to handle. After googling Cuil is down at the moment to 'cool' down.
The results for my real name are disturbing: http://www.cpedia.com/wiki?q=wyatt+greene
The page for us is interesting. Huge amounts of content, flashes of brilliance, and moments of sheer terror. http://cpedia.com/search?q=wikispaces
http://twitpic.com/1ek18g/full
The moods could also relink this thread of conversation to the screen shot.
One cpedia = one level of unintentional comedy.
For example, if the content that cpedia collected into an article ends up being unintentionally humorous, that's 1 cpedia. If the arrangement/ordering of the content adds to the humor, that's 2 cpedias.
http://cpedia.com/wiki?q=computational+fluid+dynamics
http://cpedia.com/search?q=London%20Underground
Can anyone find something useful? I'd like to see a good article.
http://cpedia.com/search?q=sam+ruby
I've been trying to find just one good article, but so far, no luck.
For a side experiment, I'm not going to fault or bash any company-- we've got a few of our own experiments which aren't the prettiest/best sites in the world.
> At other times it is weird — it does reflect the web after all.
I can see a lot of hackers wanting to try something like this and see if they can iteratively make it better and better and better.
Doesn't this sort of a site sound like something people here would love to toy with?
> For each query, Cpedia algorithmically summarizes and clusters the ideas on the web and uses this to generate a report. We do the heavy lifting of removing all the repetition, so that unique and novel content surfaces.
I personally think it's cool for people to try weird things like this-- even if the attempts fail.
the goal here isn't to find the one funny/stupid article...with this site the goal should be to find one article that actually makes sense
This is more nefarious than Mahalo. It's spam wrapped up in poorly summarised, chopped up content they have no right to be using.
And then the paragraph immediately after that doesn't follow at all. It starts:
>When he [Brian Mastenbrook] was able to reproduce the glitch at Basecamp, he began to suspect that the flaw was inherent to Ruby on Rails, the popular Web framework used by both websites.
Did they do QA before launching this thing?
>I find Cpedia best on topics that I thought I knew about. I find out things I should have known but didn’t. I’ve noticed productivity has slowed in the company since we have had it up for internal testing, as people ask each other about stranger and stranger trivia, or exclaim, “I didn’t know your middle name was Hector?”
And a search for Cheney does not show the former Vice President: http://www.cpedia.com/wiki?q=cheney
Some of the results are so funny that some groupthink has to be involved as well.
I'm surprised they even released it; the few search terms I attempted generated useless results.
I just hope this garbage gets blacklisted in real search engines ASAP, before it starts polluting results.
In Dutch it's called 'plaatsvervangende schaamte' which literally means "place exchanging shame". Shame felt on behalf of someone else, shame you feel someone else should feel. I'm embarrassed for them.
true that.
Watch your backs, Google, these guys are onto something.
http://cpedia.com/search?q=jeffrey+jenkins
What's really funny about this page is that a quick search reveals that the Milwaukee Brewers have a "Geoff Jenkins" and a "Jeffrey Hammond"
"Git doesn't include a source code editor, profiler, or web browser, either. That means your syntax-highlighted code-blocks look the same on your slides as in Emacs :)"
I'm not sure how it gets from here to there. I did write eslide, however...
http://cpedia.com/search?q=william+s+burroughs
Had they positioned this as an art project applying the cut-up technique to the Web as a whole, it would be seen as a real success.
cpedia itself speaks:
"Along with JG Ballard, Angela Carter and Anna Kavan, William S. Burroughs has probably done more to influence my writing, my reading, and my appreciation of literature than any other writer."
(No attribution in article. I like to think of this as the software's own appreciation showing through.)