Riak 1.0
basho.com
basho.com
However, I still don't quite get it. What is Riak? It seems to be some sort of Dynamo implementation (like the ironically ill fated Cassandra), but apparently it has a workflow engine? What do people use it for? What is it best at?
Right now we're using PostgreSQL, Redis, and S3. PostgreSQL gives us ACID, Redis gives us fast in-memory access, and S3 gives us an infinite KV store. Is there some reason to use Riak? Would Riak just replace S3?
Given that, NoSQL adoption often appears driven by success stories. But Cassandra seems to have the opposite: numerous failure stories. Facebook, Digg, Reddit and a number of others have all tried Cassandra in production, and have either had serious complaints or moved off to either SQL or other solutions like HBase.
Of course, these failure stories are anecdotes, and numerous unrelated factors (like bad interactions between Cassandra and Amazon's EC2) could be at fault. But I'm not sure it matters.
Has anyone on HN had a really good experience with Cassandra? (This may be the wrong thread to ask for obvious reasons.)
There are videos and talks out there from several customers listed on Basho's site, just take a look.
I can't stand pre-announcements like this.
http://downloads.basho.com/riak/riak-1.0.0rc1/
I just did a rolling upgrade from 1.0b4, and it went smoothly. I love the LevelDB support, and it's holding up well under the considerable load that I'm throwing at it.
Riak 1.0 will be available later this month. To preview
some of the new features, download Riak, or to inquire
about a commercial deployment, please visit http://www.basho.com.
Their github page still just has 1.0-rc1 tagged. I'm excited though.There is nothing preventing you setting up a cluster that spans continents. What will deter you is the poor performance of the cluster due to the added latency between nodes.
In the Dynamo paper the ring spans DCs but they also have a very different network than most that allows them to do that. In Riak it is recommended that each ring is contained in a single DC. If you want ring-to-ring replication from Basho then you can pay for Riak EDS. You could also build it yourself as others have mentioned (Kresten Krab Thorup has done something like this in Riak Mobile [1]).
Nothing is stopping you from running a single ring (cluster) across DCs, and it might even be okay for certain apps, but it's not a choice that should be taken lightly. In general, if you don't understand the tradeoffs you're making in that regard then it's best to stick to one ring, one DC.
[1]: http://www.erlang-factory.com/upload/presentations/413/Erlan...
When you go to the Riak project on github, what you find is actually sort of a skeleton, that has as dependancies all those projects I mentioned above, such as riak_kv, riak_pipe, etc.
Riak ES, the commercial offering, is a superset of Riak. It has Riak as a dependency, and adds the feature of cross datacenter replication. I think the real reason you buy Riak ES is because you're wanting to buy support.
Riak ES being a commercial product doesn't make Riak any less open source, than Oracle Server being a commercial product makes Linux less open source.
Also, Basho is keen to develop users of Riak ES, and customers of Riak (who don't spend any money) still get some support from Basho. Basho has a "Riak ES for startups" program, which gives you a huge discount.
I'm building my business on Riak because Riak is open source. IF Basho goes away, I'll still have Riak. There's nothing missing from Riak that I need.
I figure if I get big enough where I want to be running out of multiple data centers, I'll be big enough to afford Riak ES, and if I can't afford Riak ES at that point, then I'll be able to build my own solution. (I don't think it would be that hard, actually.)
I was thinking along those exact same lines, but a big unknown was pricing on their enterprise offering. That information is unavailable on the web, and despite my skepticism in contact-us-for-the-price situations, I filled out their online form, which is a request to be contacted by a representative.
I haven't heard from them, but they did put me on a mailing list—I got an email about this 'milestone release' today! Not quite what I wanted to know, though :)
Nirvana, or someone using their Enterprise offering, perhaps you could fill us all in on the price?
I cant speak to Riak, but generally this model can create a conflict of interest between the "enterprise features" on the one hand and open source commitments on the other. For example if someone submits code to the opensource version that duplicates/overlaps an "enterprise feature"
Technically, what eventually put me off, is that I couldn't figure out how to maintain a clean secondary index. If you have a: SiteId, UserId, Data, and you want data to be accessible by SiteId or SiteId+UserId, I couldn't figure out a nice atomic way to maintain the secondary index. This is pretty basic stuff. I'm glad to see 1.0 will support native secondary indexes, but I think my inability to figure it out shows that their documentation is poor (or it could be that I suck).
More info here http://blog.basho.com/2011/09/14/Secondary-Indexes-in-Riak/
LevelDB seems mostly well suited for data that becomes (in terms of key size and number of keys) bigger than your RAM...
I'd be very curious to know a bit about the character of your data, the size of your cluster, etc. (I've only run test clusters at this point, so hearing from someone doing production work would be informative.)
My biggest RAM consumer stores historical data for a goods trading platform. Each trade is a unique key, with all the trade data being the value. Access speed is important, but not as critical as the other goodies I get from Riak (replication and automated rebalancing). Metadata is stored separately, but I hope to change that with Riak 1.0 secondary indexes.
Level also has to look down the entire tree if a key is missing. This means inserts end up being more expensive than reads or updates (which are all just a hash lookup in Bitcask).
Yep, this is a standard tradeoff. When you want your data to be iterable, you have to take the hit. In practice (I oversee a large cassandra cluster), this hit happens about ~1% of the time, which is either a lot, or a little, depending on your constraints.
"Level also has to look down the entire tree if a key is missing."
This is why Cassandra has a bloom filter on top of a very similar data store.
I don't have a company. Yet.
What I am interested in is a go-away button on this obnoxious ad bar so I can read your webpages on my vertically-challenged 11" screen.