The Twitch Statistics Pipeline
ossareh.posthaven.com
ossareh.posthaven.com
I'm a designer that worked with Mike on a short project for the science team. The 4 of them are a huge pleasure to work with :) Very scrappy and analytical when it's time to get things done, but playful and generally funny fellows to be around. Interesting problems to work on when they're #4 in US's peak traffic count, after much bigger guys like Netflix, Google, Apple (and an underdog story beating out Hulu, FB, Valve) [1]
Twitch's office is very thoughtfully designed with fan art and game-themed interior and murals. It's a nice testament to the community that they're built around. The office manager, Ashley, has done an equally thoughtful job with amenities -- makes it easy to be productive.
[1] http://s.jtvnw.net/jtv_user_pictures/hosted_images/wallstree...
Encoded json in get parameters? That's not what GET requests are for. use POST requests for that, or you'll quickly be limited by the max size of a get request, somewhere around 8KB.
We're in the process of moving over to POSTs. The second part of this series will go into more detail as to the "whys". The primary reason we use GETs is backwards compatibility; we wanted our new team to be a success and in ensuring that we thought carefully about the battles worth fighting - the Mixpanel stats client is a good one, so we opted for being "Mixpanel Protocol"-compatible; they use base64 encoded json blobs shipped using an HTTP GET, and thusly so do we.
Kinesis looks really interesting, and is definitely something that we're going to look at once we work out our ETL process.
The order of priority for us has been:
1 - Get a pipeline up and running 2 - Make it robust 3 - Make it fast.
Pipeline v3, our current one, satisfies (2). We expect to be working on (3) in the near future. ETL is the latter part of (2). (3) results in powering dashboards, we expect those to contain a lot of joined data and having a robust ETL process is pretty key to that.
I think Kinesis could replace the first three boxes in their diagram, and do it in real-time. (I haven't used kinesis so I could be wrong.)
It's amazing how fast big data infrastructure is evolving. For anybody looking to build something it seems your chosen solution will be obsolete by the time you release.