1,737 karma · joined April 11, 2011
http://www.cs.berkeley.edu/~istoica/classes/cs294/15/class.h...
Still - this estimate is based on 3KB/tweet, which is probably derived from looking at the raw twitter XML feed - I'd expect this to compress easily down to 10-30x smaller than the author's claimed numbers - more with proper data modeling.
Nevertheless - huge problems to be dealt with in the sciences and video.
If you really like to gamble you can buy puts on some of the bigger ones.
If I were building a speed-freak analytics database I'd be focusing on making my CPU implementations as fast as possible, since that's what 99% of potential customers are already running. Assuming you get to 40GB/s on this type of query, that's only a factor of 5 slower than the GPU implementation. I'd imagine that for most workloads, a factor of 5 speedup that requires new hardware and lots of energy is kind of a non-starter.
This is a use case where JSON shouldn't really ever be used, because the schema is pretty much fixed and highly regular. JSONB records essentially carry the schema definition with them per-record and in this case most of that information is duplicated - hence the blowup in its representation on disk.
While column stores are great for answering analytical queries that require scans over the whole table (like the single query example they show), they aren't as good at transactional queries (like serving webpages).
If I were citus, I'd have written the blog post using a dataset of highly irregular JSON blobs - e.g. log messages from lots of different systems or a big collection of web pages (serialized as json representations of the DOM). Maybe we'll see these "in the coming weeks."
Also, vw is what I'd consider "industry standard."
If I were making such an investment, I'd think carefully about the rate of computational depreciation (which is predictable) and probably hedge on the price of power. The only unpredictable thing in this setup is the price of bitcoin, which I believe they're extremely bullish on long term. You could probably figure out a way to hedge that, too, if you could find someone to take the action.
The fact that these are sets is an important distinction and enables things like predicate pushdown and other important optimizations. Would be interesting to see a Selinger-style cost-based optimizer built in Haskell!
But seriously, I've got some battle scars from the perl days, and hope not to revisit them. Honestly, there's very little I find I can do with perl and not python, and it's just as easy to express (if not quite as concise) and much simpler to maintain.
But, use the tool that works for you!
1) Be careful with CSV files and UNIX tools - most big CSV files with text fields have some subset of fields that are text quoted and character-escaped. This means that you might have "," in the middle of a string. Anything (like cut or awk) that depends on comma as a delimiter will not handle this situation well.
2) "cut" has shorter, easier to remember syntax than awk for selecting fields from a delimited file.
3) Did you know that you can do a database-style join directly in UNIX with common command line tools? See "join" - assumes your input files are sorted by join key.
4) As others have said - you almost invevitably want to run sort before you run uniq, since uniq only works on adjacent records.
5) sed doesn't get enough love: sed '1d' to delete the first line of a file. Useful for removing those pesky headers that interfere with later steps. Not to mention regex replacing, etc.
6) By the time you're doing most of this, you should probably be using python or R.
I very rarely bet any money on my round, nor do the people I play with. True, there are people betting on their rounds, but I'd estimate it's 1 out of every 10 groups that goes out for a round that does that.
The article blows it out of proportion a little bit, too. At the courses I play (in the bay area), there are usually maybe 20 people in the clubhouse, and a round of drinks would probably come to $100. The "insurance" crowd is a very small subset of the golfing population.
True - it is an expensive sport and maybe traditionally a game for the wealthy, but my weekly golf habit doesn't cost much more than a gym membership. People from all walks of life play and enjoy the game, and it definitely doesn't have to be expensive unless you want it to be.