417 karma · joined April 12, 2013
If I think the level of process and granularity is overbearing I say "no thanks" and pull the "Individuals and interactions over process and tools". If there is one thing that I wish I would have learned earlier in my career is that it's really ok to say "no thanks".
Unfortunately at work we have SCRUM teams that practice this methodology religiously, but not on side projects. One team spent two months building a buggy functional wrapper around gRPC. When asked why- "We did it because it wasn't functional". The result is gobs of wasted cycles building overly complicated and abstracted garbage, and then further amortizing that over time with ongoing escalations.
I'm not sure how to help them gravitate away from shiny things to focus more on outcomes (simplicity, reliability, speed of delivery, meeting needs and creating value, etc..)
* As a datapoint Pinot/Druid/Clickhouse can do 1B timeseries on one server. AresDB sounds like it's in the same ballpark here
* Pinot/Druid don't do cross table joins where AresDB can. My understanding is these are at (or near?) sub-second which would be a very distinguishing feature. I'm not sure how this will translate to when distributed mode is built out, as shuffling would become the bottleneck. Maybe there would be some partitioning strategy that within a partition allows arbitrary joining or something?
* Clickhouse can do cross table joins, but aren't going to be sub-second
* AresDB supports event-deduping. I think this can easily be handled by the upstream systems (samza, spark, flink, ..) in lambda
* Reliance on fact/dimension tables. - This design/encoding is probably to help overcome transfer from memory to GPU, which in my limited experience with Thrust was always the bottleneck. - High cardinality columns would make dimension tables grow very large and could become un-unmanageable (unless they are somehow trimmable?)
Standard clickstream data, maybe 50-ish parameters per event.
> What sort of queries are you running? > How is the data stored?
Depends on the use-case. For sub-second adhoc queries we go against bitmap indexes. Other queries we uses RDD.cache() after a group/cogroup and answer queries directly from that. For other queries we go hit ORC files. Spark is very memory sensitive compared to hadoop, so using a columnar store and only pulling out the data that you absolutely need goes a very long way. Minimizing cross-communication and shuffling is key to achieving sub-second. It's impossible to achieve that if you're waiting for TB of data to shuffle around =)
> How many machines / cores are running across?
Depends on the use case. Clusters are 10-30 machines, some we run virtual on open stack. We will grow our 30 node cluster in 6mo.
> Maybe Spark doesn't like 100x growth in the size of an RDD using flatMap
You may actually just need to proportionally scale the number of partitions for that particular task by the same amount. Also when possible use mapPartitions, it is very memory efficient compared to map/flatMap.
> Maybe large-scale joins don't work well
Keep in mind that what ever happens per task happens all in memory. For large joins I created a "bloom join" implementation (not currently open source =( ) that does this efficiently. It takes two passes at the data, but minimizes what is shuffled.
I could only speculate as to what this users issues were. One difference between hadoop and spark is that it is more sensitive in that you sometimes need to tell it how many tasks to use. In practice it is no big deal at all.
Perhaps the user was running into this- the data for a task in spark runs all in memory, whereas hadoop will load and spill to disk within a task. So if you give a single hadoop reducer 1TB of data, it will complete after a very long time. In spark if you did this you would need to have 1TB of memory on the executor. I wouldn't give an executor/JVM anything over 10GB. So if you have lots of memory, just be sure to balance it with cores and executors.
I have seen spark use up all the inodes on systems before. A job with 1000 map and 1000 reduce tasks would create 1M spill files on disk. However that was on an earlier version of spark and I was using ext3. I think this has since been improved.
For me spark runs circles around hadoop.
I don't know what you would count as major deployment, but I've deployed a 30-node cluster on HW for running sub-second real-time adhoc queries. I've also run many smaller 10-20 node virtual clusters on open stack. It is a rock solid platform. Our hosted ops loves it because it just works.
The amazing thing about spark is how insanely expressive and hackable it is. The best way I can describe it is this:
* Hadoop: You spend all of your time telling it how to do what you want (it is the assembly language of bigdata)
* Spark: you spend your time telling it what you want, and it just does it
It would be difficult for an article to convey this. Having used both Hadoop and Spark, practicality and expressiveness are precisely what made me fall in love with Spark. You can do so much with so little code. Don't take anyones word for it, see for yourself. Download the code and run the interactive shell, takes 2 minutes. It was totally mind-blowing for me.
Spark make doing MR easy. I've used other frameworks on hadoop MR, but nothing compares with the ease with which you can express computations using it. And it does both batch and real-time/streaming. It is a very well thought-out project.
That said- it's not a superior language, just a different paradigm. With Scala you get very strong typing, functional constructs (so incredibly powerful), matching, options, very clean closure syntax, etc.. You'll eventually fall in-love with _, too.
I haven't used scala.js yet, but here are some downsides to Scala proper: * slow the compile * tooling isn't great (this true of anything jvm to some degree though)
Also, when you signup I pull in your historical data and begin back-crawling immediately. This way you don't have to wait 3 months for the charts, trends, intensity maps, etc.. to fully populate before you take action on the data.
With that said, I would love feedback! If you sign up now there is free trial, and you can use this nifty code for an extra 10% discount: analytics
Cheers
I was in a similar situation (not quite as bad), but I couldn't leave because of a business obligation. I survived by insulating myself from the dipshits. We formed an isolated 'blackops' team within engineering that was beholden to nobody but ourselves. We were able to move at lightspeed compared to the rest of engineering.
After a year our small team actually began a major pivot for the company (culture, technology, etc..) and breathed new life into the company. Most of the diphits were jettisoned and the others were moved into supporting roles.