Pinterest open-sources Terrapin, a tool for serving data from Hadoop
venturebeat.com
venturebeat.com
We did look at a few options before building this. ElephantDB seemed a bit heavy handed, such as having to modify ring configuration every time we added/removed servers and also, modifying domain spec yaml files for newly added data sets. It did not allow us to easily change # of shards across different versions of the data - something that our developers do often to make their jobs run faster etc. Also, it does not GC out older versions and since our workflows write new versions every day, this was a problem.
We did look at Cassandra but we also did not want to operate another data store. However, we definitely wanted to get the data loaded fast i.e. through simple file copy operations. We found that for this option Cassandra had similar issues as HBase i.e. having to do major compactions to get rid of older data versions. Tweaking the number of reduce shards was also harder.
With Terrapin, we essentially tried to build serving system on top of HDFS given the recent improvements in HDFS performance when there is data locality. We felt that HDFS was rock solid and the best storage system (in terms of scalability & ease of operation) for immutable data sets. On top of that, we built versioning, cheap garbage collection, extensible serving formats etc. as mentioned in the blog
As for Apache Drill, it is more suited to running analyst queries with latencies ranging upto seconds or 100s of milliseconds. This is not acceptable for webscale work loads where the latencies must be < 10ms for lower level serving systems like terrapin.
Can you also elaborate, how you read HFiles and serve it out from Terrapin servers? Are you using similar functionality as HBase? (With block cache like design if yes how do you keep both in sync).
Your blog is missing this interesting detail.
While building one out, we looked at VoldemortDB, SploutSQL, and ElephantDB to serve bulk data coming out of Hadoop in batches. Voldemort turned out to be much rougher around the edges than expected, ElephantDB looked very bleeding edge, and SploutSQL wasn't as general purpose. In the end we turned to Cassandra and this tool - https://github.com/spotify/hdfs2cass.
Good to see Pinterest open sourcing this.
Very typical use case for recommendation systems etc. We face similar problems with latencies on HBase (At Groupon).
So this solution seems interesting. Would be good to have comparison of other solutions Pinterest tried before building this. eg. loading data into Cassandra instead of HBase etc.
In nutshell - very specific use case - but the one which comes across very often