The end of big data
benn.substack.com
benn.substack.com
What Databricks offers that you can't find in essentially the exact same products from other cloud providers, is that they have a pretty UI and a much more integrated setup for more advanced operations like upserts on a data lake instead of having to manage your own Hudi or Iceberg setup. Their notebook is a better offering than others have, but the money for big data companies is in the long running daily or weekly jobs, not one off queries.
Historically, databricks had been offering major performance improvements to their customers while not updating the Spark open source code with all of the same query optimizations, and other various major spark performance improvements like Dynamic Pruning and Adaptive Execution, they've since done an about face for the most part and have started pushing these improvements into the open source. But as I noted before, for these longer running daily jobs, you'd spend 10x less per instance hour in some cases running on a cloud providers generic offering rather than paying for databricks, as they've made performance improvements of their own such that the price performance almost universally will always beat databricks.
The point is that databricks is absolutely a very innovative company, it's just that it's impossible to compete against the bigger players. The only real major lasting strength of databricks is that Azure fully integrated them into their service as opposed to trying to compete directly like AWS and Gcloud have.
Sparks success is that it came out of Berkeley, was better than traditional map reduce in a number of ways initially, but more than that it was able to spread because a lot of map reduce product providers were able to push it for their clients because of it's apache license explicitly.
Could you elaborate on this? To my knowledge (limited!) there isn't any specific integration. I thought they just offered a turn-key managed service, but all that seems to do is spin up a preconfigured VM scale set...
First class in my mind would be if Databricks could push down compute to the Storage Account nodes or something…
Databricks has a product called Delta Lake that covers the infinitely scalable storage part. Here's a talk from a Delta Lake user that's writing hundreds of terabyes / up to a petabyte of data a day to Delta Lake: https://youtu.be/vqMuECdmXG0
Databricks recently rewrote the Spark query engine in C++ (called Photon) and provides a new SQL query interface for data warehouse type developers.
Databricks recently set the TPC-DS benchmark record, see this blog post: https://databricks.com/blog/2021/11/02/databricks-sets-offic...
This article doesn't align with my view of reality.
I think it is just API layer on top of existing storage systems like s3 and hdfs: https://docs.delta.io/latest/delta-storage.html#amazon-s3
That’s all.
You too can have transactions if you implement your access layer appropriately.
It’s a useful format, but let’s note pretend it’s any more magic than it actually is. I’ve also not noticed any improved performance above what partitioning would give you.
Yes, we can all "just" make it ourself, still nice when others make it.
Personally I very much like that it's conseptually easy and that it builds on an open source format (parquet). I also expect there to be a dragons, both small and large, in actually getting the acid stuff right,so I am happy we can fight them together.
Regarding “that’s all” - can you give an estimate how much time it will take to reimplement it from scratch on multiple cloud platforms?
It's not magic, and you can understand what the _delta_log is doing fairly easily, but I can testify that the performance improvement over partitioning can be achieved.
So they’ll just have to suffer running their non-SQL workloads locally on corporate issued laptops? Or they’re not really going to do data science at all?
The data scientists were using Databricks, but data eng was instructed to try to replicate a few core pieces because it was $$$$ as far as "a way to run notebooks and such" went.
The only downside is their marketing team. Sadly it's full of passive/aggressive people that won't let your company's engineers work because they'll bury you with hundreds of business questions just to resell a cool story internally.
I've seen projects and deals going down for this. I really can't understand why they're trying so desperately to waste their customer's time.
Has Databricks recently open sourced additional Delta table features that were previously only available with a paid license? I can't seem to find a relevant announcement.
How do you describe a chocolate gâteau to someone who has only ever eaten rice?
> Databricks is a big, fast database that you can write SQL and Python against
It really isn’t. Spark is not a database, certainly not an RDBMS. This reminds me of the people in the office who used to call the computer on their desk “the hard drive”.
If you’ve never worked with tools like Pandas, or R, or SAS, you’re just not going to understand Spark and Databricks.
Also, for the record, that's largely how I think of databricks ideally. A big distributed database/storage layer into which I can write queries in my chosen language...
So i...admit equal confusion...
A traditional database is efficient (i.e. cost effective) at doing a lot of the same query repeatedly, e.g. looking up a customer's account balance, whereas these query engines are good at doing infrequent queries with lots of complicated, expensive logic on very large datasets.
You could run simple account lookup queries in a CRUD app with Spark, but you'd be setting a lot of compute/money on fire, and your latency would probably be terrible.
Wikipedia: In computing, a database is an organized collection of data *stored* and accessed electronically.
Apache Spark (the thing that powers Databricks) doesn't store data, it only processes it.
---
I'll give this a go... A lot of this involves some massive simplifications but hopefully this might be helpful. Let's say you have a file like this:
Customer_Name, Order_Amount
Sarah, 15
Billy, 10
James, 20
Mukesh, 18
Kate, 42
If you want to find out the total number of orders in that file you'd load it as a dataframe (pandas/R) then sum the Order_Amount column. How about if you wanted to join this file up with another file: Customer_Name, Company_Name
Sarah, Big Multinational Conglomerate
Billy, EZ Groceries
James, The Coffee Shop
Mukesh, EZ Groceries
Kate, The Coffee Shop
To find the Sum of Order_Amount by Company_Name you'd load both files, join the two dataframes together and sum by Company_Name.Where is the data storage happening in this example? It's the files, right? The data is in the files and you've just run a query against those files. Once you close your python interpretor your dataframes cease to exist -- but the files still live on.
---
What if you wanted to access this data from multiple computers very quickly becuase, let's say, you have some software that should show the user their company's current order amount? Enter the RDBMS (Relational Database Management System) a.k.a. "the database". Now multiple people/computers can run queries at the same time without needing to:
1. read the files from disk
2. load each file's data as a dtaframe
2. join the dataframes together
3. Sum the Order_Amount by user's company
Instead they run a SQL query which handles all of that for them, as if by *magic*.Where is the data storage happening here? It's the database tables, right? After your query runs the data is still stored in the database tables. It doesn't disappear.
The "MS" part of RDBMS is the query execution part. The "SQL magic" that lets you do left joins and so on. It's not the data storage part ==> it's the query engine part.
---
Now scale this up to PETABYTES of data. Billions upon billions of rows of data. The pokey little 8 core CPU server your database runs on might not be able to handle processing all that data in any sort of reasonable time for a web browser based application (any longer than 3 seconds = the web application is broken).
As mentioned by MrPowers, Apache Spark solves this problem as a distributed processing query engine. The short version is:
Stage 1. split all the data into thousands of "partitions" (literally split the files into thousands of pieces)
Stage 2. run thousands of mini "Sum of Order_Amount" queries on each partition
Stage 3. combine the results of the mini queries to get the final "Sum of Order_Amount"
The clever thing about distributed batch processing is Stage 2. You can run the thousands of mini queries on thousands of indpendent servers, meaning your total execution time for Stage 2 maxes out at the time to run a single mini query. The part Apache Spark deals with is Stages 2 + 3 ==> executing a query in a cleverly distributed way, meaning the query is processed much faster than the pokey little 8 core CPU database server.Where is the data storage happening here? Stage 1, right? The data lives in all those partitioned files.
But Apache Spark deals with Stages 2 + 3 ...?
>> Databricks is a big, fast database that you can write SQL and Python against
> It really isn’t. Spark is not a database, certainly not an RDBMS.
You could store all your partitioned data in thousands of different database instances. Or store each partition as individual files on the thousands of servers used to run the mini queries. You can store it in a partitioned parquet file format in an Amazon S3 bucket. In all of these cases Spark will read the data from those locations, process it, then give you your result.
Much like loading that dataframe from the original simple file, the "Sum" query calculation is the data processing and not the data storage part. People often conflate "Database" with "Things I can run SQL queries against". The two are not the same thing.
Make sense? Anything unclear?
Addendum:
I can see why people think of these tools as databases. It’s the dunning Kruger effect in action.
Because I know the ins and outs of what databricks are (likely) using behind the scenes it means I can’t see it as a “database”.
But for folks who have a less detailed understanding of the, pardon the pun, bricks and mortar behind it, it makes sense as “the do all the data storage stuff for me == database”.
Which is the right/wrong way to think about it? I have no idea. Depends what you want to do with it I would guess.
Just wanted to say thank you.
Do you have a blog? Your writing / teaching style is very effective.
Designing Data-Intensive Applications: The Big Ideas Behind Reliable, Scalable, and Maintainable Systems
https://www.amazon.com/Designing-Data-Intensive-Applications...
I agree that the traditional Spark notebook query interface was a lot easier for people with pandas (or DataFrame experience) to understand. But their new SQL query interface is easily understandable for DBAs.
Hype is part of what drives technology and I don't see it going away any time soon - because a lot of people make a buck purely by milking hype cycles.
I've come to see there are two modes of intelligence. The ability to understand something and say 'yes' to it is complemented by bullshit detectors, a more mature ability to see through other people muddying the waters trying to look smart, and reject their nonsense.
Every cycle has its real innovators and hype-pedlars. At the height of blockhain madness a decade ago they had worked themselves into a frenzied fetish of near-supernatural woowooism. Once all the smoke and hullabaloo has cleared, and all the grandiose promises have fallen to the wayside, what remains is a core of genuinely useful and relatively simple core concepts that enter the canon, along with standards and protocols.
This line in TFA sums it up:
> I needed a database.
The same happened with "data". The hype problem there has always been a lack of telos reminiscent of the underpants gnome's flawed "profit?" scheme. Somewhere is a crucial missing step, where we ask "Why?". And the answer is "don't worry, the data plus magical "AI" fairy dust will reveal why". There is a quite religious (faith-like) flavour to this.
I blame the NSA. The idea that "collect it all" is anything but a thugish brute force excuse to waste of billions dollars building data-centres, has led to commercial obsession with "data" as a panacea.
It isn't "big data" per se that's a problem. Some applications, like in medicine or environmental science, absolutely thrive on large sets and sophisticated analytics. But the fact is that in most applications it's mostly useless, burdensome, energy-consuming, and space-wasting. But "data-hoarding" and over-analysis is pushed by those with sledgehammers to sell for cracking nuts.
has this happened with blockchain yet? I'm not seeing it.
Maybe I haven't been around the block enough because blockchain is the first hype tech I've seen that keeps not doing anything that existing tech can't do better. And yes I've looked, a lot.
The underlying tech sounds super promising and I hope someone figures something out. Not much luck so far.
No one wants to admit it for some reason, but Bitcoin is the killer app for Blockchain.
> The underlying tech sounds super promising and I hope someone figures something out. Not much luck so far.
Plenty of luck for me at least. Bitcoin solved exactly the problem it told me it would when it was created - permissionless digital "cash" transactions that cannot be reversed. It's a very specific use case, but one I am quite happy exists when I need it.
Those who have never had their money stolen by the banking system or denied the ability to use their money as they see fit likely won't understand my take. Those that do are using the currency today.
All the other crud around blockchain is line noise hype. Cryptocurrency already solved exactly the problem it set out to, it's just not easy to get rich off it.
If the killer app of the Blockchain is Bitcoin, what does it say about the underlaying technology that the "killer app" is useful only as an extremely volatile speculation vehicle, and an ecologically disastrous one at that, with a tertiary use case of paying for illegal goods and services?
This is the use case mentioned. The OP is saying that everything else about it is useless.
Having the option to do something illegal should the need arise is a benefit to some people, and government control of the financial sector is not beloved by everyone.
It's not good for crimes because Chainalysis can track the coin movements afterward - though maybe it's good enough. And porn doesn't seem to use it, maybe because the transaction costs are too high or it's too unethical even for them.
And it's not good for speculation because it turns out it's just correlated to US tech stocks.
Neither I nor any of my friends or family members have ever had their money "stolen" by the banking system.
People I know have lost their wallet keys and lost significant value in their Bitcoin "investments" they made in just the last year.
It's a unique situation in that hype is easily and definitely measurable: market cap. You would arguably need to subtract any use for real-world applications, i. e. the drug business using it, which can easily be approximated (to the closest full percentage point of market cap) by the formula x = 0.
Only if you stare at a chart on a dashboard. If you do serious analytics and collaborate with other people you need the data to be stable and not changing under your hands as you work with it.
Daily immutable snapshots produced by a daily batch are not only cheaper to produce but easier to work with and that's why the entire industry prefers batch.
The only use case for streaming are 'standing queries' in which query result is continuously updated with new rows incoming. Eg. real time anomaly detection in access logs, ML inferrence, fraud detection, plotting numbers on a dashboard.
All analysis is done offline in a batch setup. There is no 'stream analytics'. Stream processing only makes sense if the database is capable of hosting serving layer of the model inferrence application.
batch is business as usual, while realtime is disruption.
compare to banking: payments historically were batch (ACH), but realtime payments have disrupted it (zelle, venmo, fed wire, etc.)
Apache Iceberg and Delta Lake take care of that. The fact that you're reading a particular snapshot of data does not prevent you from adding, updating or deleting data in the meantime.
Also, your view is super narrow and focuses only on traditional analytics.
I have multiple friends from Oracle, IBM, Salesforce, Adobe, etc that are now AEs at Databricks making 3 or 4x total comp they were making previously.
Fun times!
While government is an extreme example, I'm sure there's a lot of non-tech companies in a similar position. The people doing the procurement don't know anything about tech, but they do know that there's random databases all over the place that are a constant headache and they want to centralise everything in a single platform. Secondly they steadfastly refuse to hire competent tech people, if databricks ends up being a very expensive replacement for an OPs team then that's a big winner for any organisation that is unable to, or refuses to, hire an OPs team but has lots of money to burn.
Hadoop may be losing its foothold though from what I've seen, these days people are getting by just running beefy Linux machines and doing everything in Python using Dask. Most people are not doing petabyte scale data analysis.
Some evaluation Criteria:
- Ease of maintenance and operation is almost paramount.
- It's fine if the solution never lives anywhere but 1 single virtual server that scales vertically (data might grow to a couple TB, but not PETA BYTES)
- Similarly, 20 9's is not a criteria. If the machine fails and it takes an hour till someone goes and re-deploys, that's fine.
- Declarative, reproducible deployment with an easy upgrade story would be great
- Ideal if the deployment can be run locally for quick developmnet
Databricks is their hosted version just like Aurora is a hosted PostgreSQL.
The fact you don't seem to understand the difference makes me question if you are qualified to be making statements that it is garbage and we can simply write SQL. Because having worked on many large, big data projects SQL is often not the right tool for every use case.
Instead, it feels more like the end of “data science as the sexiest job of the 21st century”.
DS is mostly running reports (so BI) or stats (perennially important niche) or manufacturing hype via “latest research”.
Engineering has been, is, and will continue to be the foundation for success.
So, given that enough industry analysts believe development will be heavily automated to make top-10 lists, and recent developments provide even stronger evidence (copilot writes better comments on existing code than my coworkers,) can someone providence a strong a priori argument or maybe even empirical evidence that utilitarian software development is somehow uniquely immune to this?
"The end of big data" but then goes on to rant about how they bought "big data" products and ends with "Big Data is finally starting to live up to its potential" - this is below Medium quality.
I am not happy about spending 10 minutes on it but realized I need another 10 minutes, and in the end would still be confused, and would need post this comment anyway.
So I just stopped there and wrote this comment.
Rip writing clarity, and the effective communication...
If you want to sell people a tesla, first and foremost, you have to tell them it's a car and it behaves exactly like a car. If that is not a given, they won't buy it, no matter which bells and whistles it has additionally to getting you from A to B.
So if we go back to the topic of the article, it means: If you want to sell a big data platform, first you must make people believe it can replace their old traditional database, but make it faster.
The article was written April 8, so ~6 weeks ago. As of May 26, Snowflake is worth $39B. Crazy how fast things can change.
spark is distributed and fast runtime for applications. it can be a platform powering the entire company/startup and its backend data pipeline.
rdbms like postgres cannot compete, because sql is about storage and querying (you need application code to do complex processing) and sql is hard to scale (unless you want to shard your db into million pieces and have ways to manipulate millions of db shards). also - how good luck shuffling data between thousands of rdbms shards, if you believe rdbms can replace spark.
the problem with databricks is that their offering is not differentiated enough from opensource spark. So their only value add is "we will manage your managed spark cluster, pls pay us, instead of your Ops/IT team?"
And they now provide Spark engines as well with associated notebooks. What's the unique offering in Databricks seen in this light?
Why not? The data is still being stored in S3/Cloud storage and computed on EC2/GCE so why should Amazon and Google care that they've been relegated to selling the pick-axes into to the gold rush?
Azure seems to be at the ‘embrace’ stage with Databricks. I’m waiting for the ‘extend’ phase, and we know what comes after that.