Amazon CloudSearch - Start Searching in One Hour for Less Than $100 / Month
aws.typepad.com
aws.typepad.com
Full text SQL search
Apache Solr or something similar
Google Search Appliance
Custom search
Google free search on your site
Yay for search as a service.Are you saying it is too much? WebSolr etc have options that are cheaper.
http://www.midwesternmac.com/services/hosted-solr-search
http://www.netaphorsearch.com/products/solr-hosting
http://www.acquia.com/products-services/acquia-network/cloud...
(they're mainly targeted to drupal)
anything else?
Lucid Imagination is run by some of the most experienced Solr and Lucene devs. Specifically Yonik Seeley, who created Solr.
When I looked at WebSolr, the cost exceeded my entire VPS structure, even for their cheapest plans.
Not to mention, we've got transparent replicated redundancy on all our indexes—one of our better-kept secrets, I really need to update our marketing materials—so double your VPSs there.
For side projects I'd be okay with indexing being less aggressive, sizes being more restrictive, and response times being higher if that made the pricing more accessible.
I'd be happy to give you guys more money when I've got the traffic to justify it, but it'd be nice to be able to flip that on when I need it.
PS It looks like it's initially available only in US East Region
Searchify, Running the full open sourced IndexTank Search as a Service API
HoundSleuth, IndexTank Compatible API
IndexTanktoGO, IndexTank Compatible API
Bimaple, IndexTank Compatible API
IndexDen, IndexTank Compatible APIThere's also my own http://websolr.com/ running Apache Solr. Some other Solr services are mentioned elsewhere on the page.
I've also recently launched http://bonsai.io/ for a hosted ElasticSearch service. Because ElasticSearch is actually quite awesome (and I'm happy to answer questions about why).
For Sphinx, there's Flying Sphinx (by Pat Allen of Thinking Sphinx Ruby client fame, great guy), and IndexDen (which is Sphinx, not IndexTank).
Here's a good video on the subject from ElasticSearch's creator: http://vimeo.com/26710663
ElasticSearch has very little ceremony around creating a new index and getting started with using it. You will eventually need to do some configuration to tune its behavior for your specific application, but the learning curve is nice and gradual. This makes ES great for exploration.
The JSON document store aspect of ElasticSearch is indeed very nice. The RESTful API is simple enough that you don't really need a client, just grab your favorite HTTP client library and start integrating. Plus, coupled with solid distribution, you're looking at a pretty viable standalone data store, IMO.
Also, very good documentation. And its user/developer community is all full of the really smart, enthusiastic early-adopter types right now :)
Not least, Lucene itself is hands down the last word when it comes to search.
Elasticsearch
This tightly couples Sphinx to your application and your schema, and creates serious issues for your ops team since every app change potentially needs to modify the Sphinx config. It gets particularly hairy when you want to host multiple applications using a single Sphinx daemon.
We started out with Sphinx for our apps but quickly discarded it in favour of ElasticSearch, a much more elegant and orthogonal piece of software.
You send data to sphinx (when you update it), and its indexed right away.
The original disk-indexes (updated by a batch process is still available)
Also, I can't programmatically get a list of all "words" in the index with their frequency and the inverse dod freq, etc. With anything lucene based this kind of thing is really easy.
Yahoo! BOSS (http://developer.yahoo.com/search/boss/)
Not just customer service, too, but end to end developer experience can be a huge advantage. I'm personally very active in maintaining the popular Sunspot Ruby library for Solr, and have passing familiarity with the internals of half a dozen other clients as well. This makes for improvements at all levels of the stack.
While the CloudSearch API looks a bit more reasonable than the recently released DynamoDB, Amazon has traditionally been somewhat poor in terms of end-to-end developer experience.
And "near-real-time" amazon, is that all you got? :)
The answer is in two parts: 1) "fixed" cost to upload the data (say in one shot) and 2) the hourly/daily/monthly) cost to make these search instances running.
Also, other solutions allow indexing in different languages, I can't find this in the CloudSearch.
I don't see how that is possible unless that field wasn't indexed.
array[index] = updatedValue;
And then you can sort or filter by these numeric values, as well as the usual text searching. More info in the docs here: http://www.searchify.com/documentation/python-client#documen...
Note that if you update a text field, it does a normal reindex (although it's true real-time in IndexTank).
What i always tho is the ability to run search queries that also involve dynamic grouping (like grouping by random combinations of facets) and providing those aggregated results.
Only thing i've seen that can do this "on the fly" is SenseiDB. CloudSearch/Solr/etc seem to need preprocessing to get this right.
This is still going to be a painfull task for legacy data - it has to be massaged into shape. Should be interesting to see how this gets applied though.
I can see a lot of small businesses with S3 tools on the desktop getting excited about the ability to search their office document store, then discovering there's a whole lot of programming to do first.
Interesting to note that Pando Daily reported this as a rumour almost three months ago (though they got the announcement date drastically wrong: http://pandodaily.com/2012/01/17/good-news-for-ec2-customers...
I am very interested in the facet-search functionality, anybody know if there is sort options on the returned facet's ? Most search engines just sort facets by number of hits.
> run on Amazon’s proven computing environment
Yes, Amazon never explicitly stated, "We run our site on a fleet of EC2 instances, and you should too," but they certainly weaseled a connection with their main site that didn't exist.
Here's an excerpt from the Businessweek cover story[1] a few weeks after EC2's launch:
> Amazon is starting to rent out just about everything it uses to run its own business, from rack space in its 10 million square feet of warehouses worldwide to spare computing capacity on its thousands of servers, data storage on its disk drives, and even some of the millions of lines of software code it has written to coordinate all that.
Weeks later Bezos discussed the $2 billion investment in Amazon.com's infrastructure [2]. In effectively the same breath he mentioned 200,000 developers signed up for AWS. This is deliberately deceptive. Bezos linked AWS with Amazon.com's infrastructure, when the two are totally separate.
Signs point towards EC2 coming about through the intransigence of a couple people [3], not through a deliberate effort by top brass to rent out Amazon.com's infrastructure. Implicating the main site was a tactic to inspire confidence and provoke experimentation. It's unclear whether EC2 would have found the same success had Amazon not papered over the inchoate AWS architecture by invoking the Amazon brand.
I love AWS. Their blog is the only corporate blog I subscribe to, because everything they post is so friggin' cool. Sentences like these...
> If you have ever searched Amazon.com, you've already used the technology that underlies CloudSearch.
...are weasely and unnecessary. Maybe CloudSearch is functionally identical to Amazon.com search or maybe it isn't. Everyone understands there are tradeoffs to be made. I just wish Amazon were more transparent about its architecture.
[1]: http://www.businessweek.com/magazine/content/06_46/b4009001.... [2]: http://aws.typepad.com/aws/2006/09/we_build_muck_s.html [3]: http://itknowledgeexchange.techtarget.com/cloud-computing/am...
I don't see what is "weasely and unnecessary" about this. It's vague, sure, because "technology" is a vague word. But it seems a bit of an overreaction to believe this statement is actively trying to deceive. I read it as "we use this technology to run the amazon.com website", implying that it is up to the task of running your own web service. Time will tell if that actually proves to be true.