Querying Riak Just Got Easier: Introducing Secondary Indices
slideshare.net
slideshare.net
# Query for category_bin = "armor"
curl http://127.0.0.1:8098/buckets/loot/index/category_bin/eq/armor
{"keys":["gauntlet24"]}
# Query for price_int between 300 and 500
curl http://127.0.0.1:8098/buckets/loot/index/price_int/range/300/500
{"keys":["gauntlet24"]}
Why not use URI query parameters for the query parameters? /buckets/loot/index?category_bin=armor
/buckets/loot/index?price_int=[300..500] /buckets/loot/index/category_bin,armor/price_int,300,500
Looks a lot like link walk syntax, don't it?1) If you want to be pedantic about HTTP, RFC 2616 states with respect to responses from URIs containing query strings that "caches MUST NOT treat responses to such URIs as fresh unless the server provides an explicit expiration time." Although this clause is broadly ignored by the vast majority of middleware, it could be argued that slashes are more correct. http://www.w3.org/Protocols/rfc2616/rfc2616-sec13.html#sec13...
2) Middleware like Squid comes with strip_query_terms on by default so if you put Squid in front of a Riak cluster and you wanted to see what was actually being run you'd have to make changes to the config file. Otherwise the request uri in the logs just reads "/buckets/loot/index?"
If the key is individual armor type, there could (and usually will) be values in any price range at every partition.
Is the index partitioned separately on its own key? Is a copy of the value stored in the index or is it then retrieved separately from the index scan?
In the examples given (price, license plate) there is no locality between the partition keys (armor id, person) and the index key. A query for all armor priced between 200-400 would end up touching every partition that contains armor priced between 200-400. Unless the set of armor is small you will end up needing to scan every partition.
For example, in a 4-partition ring with N=2, keys mapping to p1 are replicated on p1,p2; p2 on p2,p3; p3 on p3,p4; and p4 on p4,p1. As such, you only need to query p1,p3 or p2,p4 to cover the entire keyspace.
In general, approximately RingSize / N partitions need to be queried. The new smart coverage code figures this out as well as deals with routing around failed nodes and other issues.
EDIT: Since the replicas value (N) is settable per bucket in Riak, there's some interesting extreme cases that you could envision here. For example, you could have a bucket where N = RingSize, in which case the index is replicated to every node and you only need to query a single partition to lookup values. Of course, then you lose the ability to perform multiple queries in parallel with a more partitioned/distributed index space (which would be more useful for large results sets). As with database systems in general, the best configuration here depends on data and use case.
Technically, when you perform a write, Riak will always dispatch to N replicas. W simply requires Riak to confirm W writes before responding to the client. So W=N allows you to know N index sets have been updated, but it's not strictly necessary. At the end of the day, indexes are eventually consistent like the rest of Riak.
I agree PR and lots of BUZZ makes it really difficult to choose the right DB.
My latest obsession is Riak => Basho is VERY honest, and quick in helping you out and telling where and how Riak can help you, and most importantly what would NOT be a good fit for Riak. ( I am not working for them :)
/Anatoly