IBM Watson Acquires Technology from Blekko
asmarterplanet.com
asmarterplanet.com
Freebase was acquired by Google in 2010 and powers their internal Google Knowledge Graph. Microsoft bought Powerset (company) in 2008.
Freebase has 2,751,614,700 facts, Wikidata has 13,788,746 facts. Wikidata may import some data of Freebase, but due its stricter guidelines (notability guideline...) many facts of minor will be lost/never migrated. A Freebase dump won't age well, in a lot of cases up-to-date facts from the real world are required.
Maybe some community project can rescue the Freebase community project before it is too late?
@downvotes: ?
https://gigaom.com/2013/09/10/diffbot-brings-big-time-search...
Others that I can think of are IXQuick (not sure if they use their own index or not) and Yioop (smallish index). There really is a lack of large players indexing the web.
Its also possible that we will see the rise of niche search engines, such as iconfinder.
DuckDuckGo makes a business not tracking users and I believe made over 1.2 million last year based on its donations.
A search engine doesn't need to track people to make money, they just make much more money if they target ads better by tracking.
What those neither of those give you is an up to date index, which is why small search engines still need Yahoo and Yandex's APIs. I'm not sure any free resource can match the speed at which big companies can index the web.
The real problem imho is building a distributed index and fighting spam. Both are incredibly hard problems to do well and very expensive. Hence so few are trying to do it.
I tried building a small search engine with a very basic algorithm and it worked very well for 90% of searches.
I agree with you in part but I also believe if you got the others 100% right you would have a respectable engine by itself.
IIRC google used to scan different pages at very different frequencies. Quite possibly because it has assigns pages into subsets every time it indexes.
Today's search engines are not just indexes of web pages. They give direct answers for "when did Lincoln died", they show detailed street and satellite maps with StreetViews in search results, they act as business directory, they act as people finder, they have elaborate image search (again super expensive to do), they recognize objects in images, they have freshest news search for thousands of sources, they can do video search, they have scans of 100s of thousands of books, they catalog millions of products and so on. Even getting some of this high quality data such as satellite maps, books, business listings, product catalogues etc would cost ~100 million dollars in licensing deals.
Even if you decided to get world's most productive programmers you will need at least 300 people in my most minimal estimation and more than 2-4 years before you can build anything that has non-negligible chance of competing. That's about $75M of cost per year right there at the ultra-minimal end. Of course, this is assuming you already had bigger breakthrough than that of PageRank and that you can beat state of the art machine learning and natural language processing techniques. Hopefully you can now see, for all intent and purposes, search business is closed to attack via startups. I can imagine Facebook and Apple would sooner or later get in to this business but it would be uphill battle for them primarily because of lack of talent that search engine business requires and more importantly lack of data that only Google has for billions of queries and users doing them. You can build AirBnB with smart college hires but building search engines needs truly the cream of the crop who are perfect rare blend of Computer Scientists and part time mathematicians while also being an exceptionally productive applied programmers. Google has been working for years to sweep away pretty much all talent in this area and it would take most competitors significant effort to build this kind of army.
PS: DuckDuckGo is not a "real" search engine. It has a very little of its own index and for most queries they just rearrange results fetched mostly from Bing while inserting links from their own little index. If Bing shuts them out for leaching off of them, they would be toast.
I can't tell you what the answer is, but I am pretty sure that it's outside the box you've built yourself into.
But perhaps the most interesting thing I've learned from my 5 years at Blekko doing search is that "phase 2 or 3" of the Internet is here. Crawling everything is a waste of resources because 95% of the "new" stuff coming online on the Internet is not information, it is just spam. This combined with advertiser burnout from getting scammed again and again on advertising that claims to generate leads or sales but leads only to click farms. And the whole eco system of the web is in being rocked, along with media distribution. The world is approaching some sort of climactic shift of orientation.
Blekko's key mission has always been to try to find the needles in this exponentially growing pile of hay. And it is something that the folks at Watson really liked about our technology when we first met at their outreach program to connect with startups. That is what lead to their asking us to join them, and no, they weren't particularly interested in the stuff we had done to provide more topical advertising signals. So as a technologist this has been both a validation of our work on finding the real information in the web and making it useful to people, and, when I'm honest with myself, a welcome step away from working on still more advertising technology.
Having the resources to pursue that, and an engine (Watson) that can put it to use, seems pretty exciting.
Best of luck! At least you're fighting the good fight.
IBM Watson absorbed AlchemyAPI a few weeks ago.
I helped a friend's company integrate IBM Watson into their product, and I have mixed feeling about IBM Watson: plenty of potential, but some rough edges.
People talking about needing to building a better search engine better get to work because I think that space is almost abandoned .
A recommendation , an opinion , and a result are a hell of a lot different .
My gut tells me IBM just needs some independent developers to develop some apps and then be acquired( as in that point they'd have no choice).
I may be wrong, but I'd rather not take that chance as I work on my projects.
Should the title just say "IBM Acquires Blekko"? The article seems to stop short of saying so.