Data Wrangling at Slack
slack.engineering
slack.engineering
I totally get where you are coming from. Right now I'm thinking about a web API that feeds data into Kafka, to be processed (in Python, maybe Go?), stored into Cassandra and later on be the target of large Spark jobs, by the way, I need to present this info through pretty graphs and tables - Pandas will come in handy!
Sometimes it's better to just use what someone else has built, let them think about the implementation, the storage, the traffic and the maths... Here is where a third party solution falls apart: a) Costs. Data Analysis is stupid expensive. b) ... and this is the important one: Your sales/consumer facing teams want some extra numbers, literally the sort of thing that only fits your business. The solution you decided on doesn't support that use case, you are now stuck with an inflexible solution.
New Relic Insights is OK for some use cases, completely useless for the majority of analytics I need to serve, though. If it fits your bill, great! Save yourself A LOT of time and life span... Just keep everyone else on the business away from it, or they will start asking for things you can't give :)
A lot of the times these systems are built not only to serve business insight and stats, one of our main systems needed to answer two requirements: a) better/faster analysis for us; b) serve as a machine learning platform to serve better content to our users.
a) complements b) perfectly, as we collect data for analysis, that same data feeds into other areas of the business that help our users, on the fly.
You could argue there are solutions out there that satisfy a) perfectly, but the learnings of doing a) is what made b) possible.
Even if you're happy with a solution like New Relic (and by all means, I'm sure it's a good product, we use New Relic a lot!), what happens when someone has an idea like... oh I don't know... "can we build something that looks at the past 7 days worth of data and flags up any metric that moves away from the standard deviation line? Also, can you then match that against historic data and identify patterns/catch false positives?"... Just an actual, factual, example that I'm working on as well.
With that being said, and to your point, I would not be surprised if these systems were often over engineered when a sql query could get the job done.
RedShift takes SQL queries.
Together with cloud SQL tools like BigQuery or Redshift, it gets rid of the need to build a "full analytics stack" on your own. You can license the data collection/enrichment from us (we've already scaled it to billions of monthly events), you can use our clean starting schema (over 100 enriched fields per event), and then you can pipe the data into a fully-managed analytics warehouse, or just analyze it in raw form. Then you can actually spend all your time focusing on insights, rather than fussing about data collection clusters, pipelines, ETLs, etc.
I would love to hear what you think of the idea; it was launched just a few months ago.
Anyone know of anything out there?
If you have a small project (what is "small"?) you just deal with Google analytics or direct SQL requests to the single database you have. Don't need fancy tools.
The two free stuff I can think of are piwik and snowplowanalytics. They clearly suffer from "free open source" when compared to the paid tools out there.
Good data analytics software can only come when these two areas learn to teach each other. Programmers need to learn maths to the point where they are comfortable enough to implement a valid solution, Mathematicians need to learn about building software that others can use.
It can easily spiral into, "Pearson's Correlation" or "Give me the Linear regression of the bastard".
That being said. I guess that having had maths classes in my engineering degree skews my point of view, combined with working with Quants at times, who do analysis way more advanced than that.
Like yourself, I had quite a bit of contact with maths during my engineering degree - whether I took most of it in is a different question :) (Financial Calculus nearly destroyed me).
Developers aren't typically aware of concepts outside basic statistics, and even though a lot of algorithms are readily available for everyone to implement and benefit from, how can you use what you don't know conceptually?
I guess everyone has a different experience, depends where you're working, really. I do know of quite a few shops where the push for analytics came from the tech people, mostly because companies don't employ people with the math knowledge to identify these business gains.
The typical reddit developer who got a job without a degree is unaware of many many things.
The typical developer who got a job with a hardcore interview at random financial company and is surrounded by other master's and PhD. Not so much.
The typical tech company doesn't need much advanced analysis. If they could figure out how many recurring users and revenues they have, that would be a good start :D
The big data space still feels like an overengineered, fractured, buggy mess to me. I was hoping spark would simplify the user experience but it's as much of a clusterf*ck as anything else.
How hard can fast, reliable distributed computation and storage for petabytes of data be? He said ironically.
In general, I've learned to be skeptical of any new big data solution. Hadoop and hive are clumsy but as someone on my team said "they've found and fixed the tens of thousands of bugs".
It seems to take five years before any significant new solution is stable and reliable enough to be used on large, complex workloads.
Which makes me really uncertain how we get out of this situation. Maybe something like arrow is a silver bullet that fixes everything with minimal complexity and thus few bugs. But I'm skeptical.
Any single other component you may try to add will increase the complexity factorially. Better stick to the basics ;)
Hoping someone on this thread could answer a related question - how do you store data in Parquet when the schema is not known ahead of time? Currently we create an RDD and use Spark to save as Parquet (which I believe has an encoder/decoder for Rows) but this is a problem because we can't stream each record as it comes and use a lot of memory to buffer before writing to disk.
There is a bit of more latency when using S3 compared to HDFS, but it's not bad and the benefits overcame that. We do have a couple of jobs that store some intermediate results in HDFS, but in the end everything lands in S3.
We encountered a few issues with S3 at the beginning mostly around the eventual consistency, but nothing that could not be fixed.
Generally speaking, HDFS is going to be a clusterfuck to support unless you give a load of cash to cloudera (actually, it will be regardless but slightly better with the bill) -- even then you'll get the typical db vendor line of 'not running -some patchset ver-, then upgrade. Which is really risky on a large cluster which pretty much works as you want.
Also, unless you've got a load of hardware you can dedicate to this environment, then you're going to be spending a lot of money on IAAS bills and your performance is probably not going to be very good. (Yeah sure you can virtualize HDFS but generally I passthrough local storage to the VM's, and only run demo on AWS etc).
There was been a push towards such mental complexity and folks convincing themselves they needed to solve their problems in this manner, and now a bit of an ebb backwards (at least, in the general space) now that your avg deployer found out how hard it is to do this stuff even with good support. Massive data ingestion and huge batch jobs might be a solution to a given problem you have, but it's probably not the only one whereas it's almost certainly going to be the most difficult and expensive.
Personally, I'd avoid hdfs, flume, hfs, zookeeper and all the rest of the nightmares until you're absolutely sure that you need them (and if you're not already, then you probably don't).
Also: Check out manta from joyent. :}
That should be the de-factor standard for TB scale. In fact, don't bother comparing other products if you're TB scale, just use S3.
Say you're going to ETL or Map/Reduce over all that data a lot of times, you're telling me that reading it all for processing over S3's rest api (which is the only method?) instead of, say, a local array of 15k sas's over pcie hba's is ideal?
It's pretty expensive and inefficient to my eyes, what am I missing? I
In what way would S3 be better than running this on your own gear if cost and perf are clearly not going to be better (which are really the big factors in this decision)?
They are pretty cheap, efficient and simple to use ;)
Using S3 with EMR in production was breeze for us. Even cost effective, since you can play with spot instances depending on your jobs. You also improve utilization of your resources.
With recent Athena it is possible also to do ad hoc queries directly :) Before it required starting "QA" cluster.
- alooma.io (SaaS queing and transformation pipeline that saves to S3)
- segment.io (Saas analytics platform that can save to S3)
- snowplowanalytics (clusterfuck open source self hosted analytics pipeline)
Does anyone have experience with both that can talk to their strengths / weaknesses?
So this is not about data engineering, but data management/analysis.
One thing that sounds very interesting and worked surprisingly well when I played around with it was Amazon's Athena (https://aws.amazon.com/athena/) which lets you query Parquet data directly without relying on Spark which can get expensive quickly. I wouldn't trust production use cases just yet and it ties you more and more into the AWS ecosystem but might be worth exploring as a simple way to do basic queries on top of Parquet data. I suspect it's simply a managed service on top of Apache Drill (https://drill.apache.org/).
Since s3 listing is so awful, and the huge number of partitions we needed, we had to write a custom connector that was aware of the file structure on s3, instead of the hive metastore which has lots of limitations, so im a little wary of athena. create table as select is amazing too, write sql to generate temporary parquet/orc files back to s3 to query later, i hope will support this if it doesn't already.
Some amount of S3 listing optimisation is done by Qubole's engineering team for: https://www.qubole.com/blog/product/optimizing-s3-bulk-listi...
They also have features that allow you to auto-provision for additional capacity in your compute clusters as your query processing times increase.
I like sampling for figuring out how something works, it allows me to iterate much, much quicker.
However, if you need individual level predictions, sampling probably isn't going to help.
Slack, Hive, Presto, Spark, Sqooper, Kafka, Secor, Thrift, Parquet.
I sometimes can't tell the difference between real Silicon Valley product names and parodies. I'm starting to miss the days when it was all just letters and numbers.
An excerpt...
Which raises the question: With such a good tool for team communication, why does Slack need an office? Why not do all your work virtually?
Slack CEO Stewart Butterfield gives product manager Mat Mullen advice, and a ukulele serenade. “There are some conversations that are much easier in person,” says Brady Archambo, Slack’s head of iOS engineering.