Big data is dead
motherduck.com
motherduck.com
The real issue is that business people usually ignore what the data says. Wading through data takes a huge amount of thought, which is in short supply. Data Scientists are commonly disregarded by VPs in large corporations, despite the claims about being "data driven". Most corporate decision making is highly political, the needs of/whats best for the business is just one parameter in a complex equation.
I did several experiments, and noticed that whenever I produced analysis that was in line with what management expected - my analysis was praised and widely disseminated. Nobody would even question data completeness, quality, whatever. They would pick some flashy metric like a percentage and run around with it.
Whenever my analysis contradicted - there was so much scrutiny in numbers, data quality, etc, and even after answering all questions and concerns - analysis would be tossed away as non-actionable/useless/etc.
if you want to succeed as a Data Scientist and be praised by management - you got to provide data analysis that supports managements ideas (however wrong or ineffective they might be).
Data Scientist's job is to launder management's intuition using quantitative methods :)
> Data Scientist's job is to launder management's intuition using quantitative methods :)
It’s no different than the days when grey bearded wisemen would read the stars and weave a tale about the great glory that awaits the king if he proceeds with whatever he already wants to do.
The beards might be a bit shorter or nonexistent, but the story hasn’t changed.
Ouch. This is savage, but sadly correct in many cases.
HOWEVER, to play devil's advocate here, I've also seen corporate data scientists overstate the conclusions / generalizability of their analysis. I've also seen data scientists fall prey to believing that their analysis proves would should be done, rather than what is likely to happen.
The role of an executive or decision maker is to apply a normative lens to problems. The role of the data scientist / economist / whatever is to reduce the uncertainty that an action will have the desired effect.
So, a synonym for 'consultant?' :)
Nonetheless, the CTO went on a multi-year, 10s of millions of dollars, huge data tech stack & staffing reorg shake up... with really zero data points explaining the driver, or what we would measure to determine it was successful.
So it became a self referential decision that we are successful by doing what he decided, and we are doing it because he decided it.
The skepticism isn't a problem, the unequal application of it, the potential to harm careers, and the chilling effect as people wisen to how best meet their own personal goals is.
The SNAFU principle: communication is only possible between equals. When an hierarchical divide exists the subordinate will tell the superior what he wants to hear.
I would say it was a very engineering driven org however, so if you could present compelling data it could go a long way.
>> Whenever my analysis contradicted - there was so much scrutiny in numbers, data quality, etc, and even after answering all questions and concerns - analysis would be tossed away as non-actionable/useless/etc.
It's a good sign at the company that I run, anytime our analysts/data scientists come up with metrics that say we're killing it, or that our ideas should bear a ton of fruit, the kneejerk reaction is to be extremely skeptical of the results. Usually they're still right.
When the data scientists say we're fucking something up, we tend to pay a lot more attention.
Only the paranoid survive, after all.
Also someone to blame if it doesn't work out
That was a plot point in Dirk Gentley's Holistic Detective Agency (1989), though the observation much pre-dates this.
I found it incredibly stressful to discover and provide analyses (even experimental results) that wasn't expected, or contradicted prior beliefs. The findings were always very harshly scrutinized, and typically lead to tons of pointless extra work to 'understand what is going on'.
My friend runs a successful market research agency and she says she gets called in when management have decided they need to make a change but need evidence to sell it to the shareholders and staff.
> Most corporate decision making is highly political, the needs of/whats best for the business is just one parameter in a complex equation.
100% Individual humans are emotional creatures with their own wants and needs, and it's important to understand how organizational incentives drive decision making.
> Data Scientists are commonly disregarded by VPs in large corporations, despite the claims about being "data driven".
This has not been my experience, though. The more common thing I've seen is that, sometimes data is boring and doesn't really show much actionable insight, but as everyone wants to justify their job, I've seen data scientists come up with really questionable conclusions that fell apart on further inspection (call it "p-hacking the enterprise").
Plus, a lot of this data in these data wearhouses is messy. Often times data scientists are siloed at the end of the process, but then you get "garbage in/garbage out" results, where there is some bug in data tracking that isn't understood until it's too late. Much better in my opinion to have data engineers and data scientists work much more closely with product engineering teams up front so they can help ensure the data they collect is accurate.
Oh, and a bunch of data scientists with zero domain knowledge for whatever data they are analyzing, preferably with PhDs in maths, but some ML background will do. And agile, because of course all those Palantir dashboards can only be developed using agile.
Once all is said and done, zero insight was created but a whole lot of consultants, contractors and project managers have been paid handsomely, while some higher ups can now put "implemented agile and big data at X" on their LinkedIn profiles.
Most of my analyses provide very little value because they are sort of common sense to people with domain knowledge. When I ask people what could be more useful, one of two things usually happen: 1) it's impossible due to data and/or infrastructure limitations, 2) what they ask turns out to be nonsensical in further analysis (like asking for average of something that follows a very fat tailed distribution with a few observations dominating the phenomenon. Of course it's usually impossible to explain this to people).
The more I think about this, the more I think that in truly data powered companies, both the decision making and data analysis have to be carried out by more or less the same people. The organizational hierarchies have to be much flatter. Essentially the employees will have to be some kind of "secret agents" who have both the skills and the mandate to steer the company in the direction they see fit. I sort of see this already happening in the FAANG companies where, or so I hear, it's very difficult to get hired, the staff count is quite small compared to traditional companies and the senior engineers have a lot of power in the company.
Using math PhDs or Palantir or whatever as a sort of modularized black box for "insights" while giving them no real skin-in-the-game does not work.
Are you going to take the time / money to set up a warehouse, get all the data into with an ETL product, set up dbt or some other transformation layer, set up a BI tool and build the reports and dashboards, etc.
Regardless the size of your data, you still need to get it in one place and model it in a way it's actually usable.
I told my VP that the engineering foul ups in the current product are easily fixable. Standard tooling and patterns exist to re-architect and solve the bottlenecks. What is much harder is a data architect to make sense of the complex data and make sure there is good value for our customers.
Guess what position I don't have on the team, and won't have due to budget issues.
The expertise we need in the industry is people who understand applications in-and-out and make great decisions on what data is worth keeping for present and future applications. And what data is needed to be kept, but only in aggregates (or anonymized, which reduces costs of maintenance)
But I have thousands of data points, thousands, I tell ya!
doesn't help that most of this data goes through multiple layers of BS where each person is putting it through filters to make themselves look better. And a good chunk of people don't have enough understanding of stats to understand when they are being tricked
Though I think there’s a better way— that is executive data science, what I do at Zapier. The key is that I’ve built up a huge amount of econometric and economic / business skills that I apply to affect good change in collaboration with company leaders. It allows me to work with Execs using sensible analysis. I improve growth and output by helping us catch errors of assumption before they go into production and cost growth / bad surprises. I also help the executives gain alignment around good information. This multiplies their departments’ output by allowing them to work better together, more in concert. That helps avoid issues with data being bent to decisions.
They typically carry a lot of hard-won valuable domain knowledge (that I combine with my economic-statistic knowledge and skills for rigor).
It’s my job to ensure Execs start with good sensible information regarding objectives. They usually ask fantastic questions and share a lot of great analysis of their own.
There are times I learn about what I might call controversial implications. This is typical of innovation using technology. It’s in these moments that I feel I create the most value by highlighting the trade offs I believe we face / potential regret.
I worked in a "data-driven" company of 5,000+ employees, with 4 data scientists who were spilt into two separate non-collaborating teams. In effect, they were so under resourced they got nothing done.
There's an agenda by some level of management, and they use "data" to forward their agenda, or disregard it due to "unexplainability" (a legitimate concern for BD/ML) if it disagrees with the agenda.
Which itself both takes time and is wildly unpredictable - neither of play well with today's Taylorist managements schemes.
kthnxbye
Most of what we as an industry are able to tell growers is stuff they already know or suspect. There is the occasional suprise or "Aha" moment where some correlation becomes apparent, but the thing about these is that once they've been observed and understood, the value of ongoing observation drops rapidly.
A great example of this is soil moisture sensors. Every farmer that puts these in goes geek-crazy for the first year or so. It's so cool to see charts that illustrate the effect of their irrigation efforts. They may even learn a little and make some adjustments. But once those adjustments and knowledge have been applied, it's not like they really need this ongoing telementry as much anymore. They'll check periodically (maybe) to continue to validate their new assumptions, but 3 years later, the probes are often forgotten and left to rot, or reduced in count.
There's initial value from training yourself on what something looks/feels like … but diminishing returns after that. Whether there is more value to be found doesn't seem to matter.
Factories would sensor up, go nuts with data, find one or two major insight, tire of data, and then just continue operating how they were before … but with a few new operational tools in their quiver.
Same is true of fitness trackers: you excitedly get one, learn how much you really are sitting(!), adjust your patterns, time passes … then one day you realize you haven't put it on for a week. It stays in the drawer.
Not unless they're threatened with ruin will people make changes to the standard way of doing things. This is actually … not bad! Continuity is important, and this is kind of a subconscious gating function to prevent deviation from a proven way of working. So, the change has to be so compelling or so pressing that they're forced to. Not a bad thing.
While we think things change overnight in this world, they generally take awhile … stay patient … it's worth it.
I definitely did not need a multi-PB storage array to reach that goal. Or did I? I am sure someone out there knows roughly the human brain storage capacity in MB.
"Mate, we don't need a chip to tell us the soil's dry"
Having tons of data is a Good Thing, so long as you can afford the marginal cost of gathering and managing all that data so that it's ready at hand when you need it later.
It's how you use the data that makes all the difference. If you're facing an issue you don't understand at all, don't go digging for random correlations in your mountain of data to find an explanation.
Think like a scientist: you need a valid hypothesis first! Once you have a hypothesis about what your issue might plausibly be, then you make a prediction: "If I'm right, I suspect our Foobar data will show very low values of Xyzzy around 3AM every weekday night". Only then do you go look at that specific data to confirm or refute the hypothesis. If you don't get a confirmation, you need to go back to hypothesizing and predicting before you look again. You can't prove causation by merely correlating data.
Absolutely. But in my experience, there's this massive trend across the tech world that flat out rejects the value of domain/subject matter expertise. Instead, all you need is an engineer who can throw some ML at the uncurated mountain of data your organization has collected. Little to no value is placed on the resources that can frame an actionable hypothesis, even though the entire value proposition arises from this exercise!
Maybe I'm just jaded. I end up wasting a lot of time trying to re-direct data scientists and engineers down more appropriate pathways than if the problem they're solving was just brought to my attention earlier. Sorry, I understand you spent two weeks shoe-horning dataset X into our analysis system for your work, but it's invalid for the question you're asking - use dataset Y instead, and you'll have an answer in an hour or two.
Sounds like the data scientists need to get together with the MBAs and they can do companies where nobody needs to actually know what they're doing.
You don't need a chip to tell you that the soil is dry, but if you can use that chip to regulate drip irrigation that can apply substantially different flow to different plants, then you can get a not-too-much, not-too-little watering even if you have a big variation in conditions.
You don't need a big analysis to acknowledge that everybody knows that a particular competitor has lower or higher prices and adjust your pricing; but doing that continuously on a per-product basis does require data and analysis.
I've worked on many product-led-growth initiatives in the software industry. The software industry is probably the biggest 'believer in data' there is -- many scientific-forward minds who understand the value. However, even in the software industry, it's really hard to convince folks that if you make 5 improvements that net 1% conversion gain each, you can dramatically improve revenue.
1. Keep the raw full data for short period of time, at most 1 month.
2. Downsample what you need for longer period of time (5-10% of the full data).
3. Aggregate your metrics on a yearly basis to save money and compute costs.
Not many orgs keep their data that long, though. Or even think about the future that far.
So hyperspectral, like big data, is useful up front. But in the end, much simpler tools and algorithms will solve the problem on a continuing basis.
I spent a number of exciting year developing a high frequency soil impedance scanner and finally understood why I was doing it. To confirm the obvious :)
That would only be a million bits (1 Mb). You're counting potential states, not bits.
If you want Toyota style continuous improvement you would need to improve in new areas of the process / new metrics, most of the time?
The problem is that they don't stay geek-crazy?
We are drowning in data, it's all around us. Information overload is real. Data enables most of our daily digital experiences, from operational data to insights in the form of user facing analytics. Data systems are the backbone of the digital life.
It's is an ocean and it's all about the vessel you pick to navigate it. I don't believe that the vessel should dictates the size of the ocean, it's simply constrained by it's capabilities. The trick is to pick the right vessel for the job, whether you want to go fast, go far or fish for insights (ok, I need to stop pushing on this metaphor )
This visionary paper from Michael Stonebreaker (2005) predicted it quite accurately and I think is still relevant: https://cs.brown.edu/~ugur/fits_all.pdf
Databases come in various flavours and the "trends" are simply a reflection of what the current era needs
Disclaimer: I work at ClickHouse
To say big data is dead sounds to me like someone desperate for eyeballs.
I do think there is a huge opportunity for DuckDB - running analytics on 'not quite big data' is a market that has always existed and is arguably growing. I've seen way too many people trying to use Postgres for analyzing 10 Billion row tables and people booting up an EMR cluster to hit the same 10 Billion rows. There is a huge sweet spot for DuckDB here were you can grab a slice of the data you are interested in, take it home and slice and dice it as you please on your local computer. I did this just this weekend on DuckDB _and_ ClickHouse!
Disclaimer: I work at a company that is entirely based on ClickHouse.
Weirdly there's a similar thing that can happen to codebases, specifically unit tests and test fixtures that outlive any of their original programmers, nobody understands what's actually being tested and before each release lose days/weeks hammering to "fix the test". The only solution is to throw it away, but good luck getting most teams to ever do that, because of the false comfort they get -- even though that fixture is now just testing itself and not protecting you from any actual bugs.
I mean how often does Netflix need to look a viewing habits from 2015? Summarize and throw it away.
Throwing out unit tests? If you make a change and it fails a test, then you fix the bug or fix the test. I can't even imagine in what universe it's a good idea to throw away a test if it covers code in use. In what universe are unit tests "false comfort"? And if "nobody understands what's actually being tested" then you've got huge problems with your development practices.
Similarly, viewing habits from 2015 are tremendously important. There may be a show they're releasing soon that is most similar to a title released in 2015, and those stats will provide the best model. "Summarize" requires knowing how data will be used in the future, but will likely throw away what you need. Not to mention how useful and profitable vast quantities of data are for ML training.
Storing data is incredibly cheap. I'm actually curious where this desire to throw away old data comes from? I've literally never encountered it before, and it flies in the face of everything I've ever learned. The only context I know it from is data retention policies, but that's solely to limit legal liability.
After that, they become a liability that slows down builds, makes changes brittle and code based schlerotic.
A few good unit tests are a lot better than a bunch of bad ones. And even from your statement we can tell a much more pernicious risk — the false beliefs that code coverage measures whether code is tested and that a code coverage percentage is a mark of quality or safety in its own right.
The villians here were monstrous test fixtures instead of mocks, "testing the fixture" instead of testing the code. Both were agency trading systems so "platforms" of a sort that needed significant refactoring to mock properly, so instead tests had to inject essentially fake concrete services.
Somehow I joined teams twice in my career that were trapped under this (who both indeed had "huge problems with their development practices") as their only coverage. The only way out is to write all new unit tests.
> An alternate definition of Big Data is “when the cost of keeping data around is less than the cost of figuring out what to throw away.”
This is exactly it. It's way too hard to go through and make decisions about what to throw away. In many respects, companies are the ultimate hoarders and can't fathom throwing any data way, Just In Case.
Really appreciated the post overall. Very insightful.
As an anecdote to this article, when business folks have come up to me and asked about storing their data in a Big Data facility, I have never found the justification to recommend it. Like, if your data can fit into RAM, what exactly are we talking about Big Data for?
In a larger sense, it's a challenge to throw away stuff, just as it's difficult to trim big data.
As I reach retirement, our attic, bookshelves, and cabinets must be trimmed -- and each item requires attention and a decision.
Some things in the attic are obvious liabilities (what to do with a mercury barometer? A radium dial pocket watch? Old electronics?) Disposing of other stuff requires time, insight, and a sense of the future (should we keep those fingerpainted scribbles from when the kids were 3? How about those cheesy trophies from chess club? Computer books from the 1970's? Betamax home movies? Record albums?)That's a fantastic point, and I keep mentioning the COST paper to anyone who cares:
https://www.usenix.org/system/files/conference/hotos15/hotos...
Big data was never going to be useful to even medium size enterprises, unless anyone can get public access to PBs of data, but that doesn't mean big data is dead. ChatGPT is literally changing how school will test their students, for a start.
Maybe what the author is trying to say is 'small-scale big data is dead, but big data chugs on.'
Sure, instead of schools checking for plagiarism from other students' papers using turnitin.com, they'll check for plagiarism using ChatGPT tools that scan for known output from their industrial-scale amalgamation of plagiarized materials. Big whoop.
[1] https://www.nbcnews.com/tech/innovation/chatgpt-can-help-foo...
In the small institution I am currently working with, the English courses, in one week, integrated chatgpt as a tool for students to work with. It's part of the collaborative idea building and development process now for every student enrolled in creative writing and writing analysis classes, and that happened in one week. I cannot stress enough how unbelievably fast that is for higher ed. That's faster than light speed.
And we're not even that well resourced. I have to imagine there are other examples where it's more than just running through a bot to scan for known outputs.
More Big Data!
Here's a novel idea: test students using pen and paper?
Actual good information will always be useful, most of this "big data" seems to be the equivalent of recording background static.
I guess the take away however is still that regular businesses really just can't play in this game and should not be assuming they have big data until that fact asserts itself out of necessity rather than the other way around.
ChatGPT and the like are not going to get much use from that kind of data and instead are looking at a giant corpus of text and images scraped from a variety of public sources to infer what humans might think sounds smart. It's possible the two worlds will meet, but that's probably not what's going to be announced this week.
At that point, does it keep scaling or is there an S curve where 100x more data and compute only leads to a 2x improvement?
Careful with the scales. 2x improvement could be interpreted from 80% human performance to 160% human performance. Or going from 10% error rate to 5% (which again crosses into superhuman territory on some tasks). Those last few bits are the critical ones.
That's pretty much exactly what the author says in the article.
That being said, what's the product? The website says "Commercializing DuckDB", but that doesn't give much of an idea of what they're offering. DuckDB is already super easy to use out of the box, so what's their value-add? It's still a super young company, so I'm sure all that is being figured out as we speak, but if any MotherDuckers are on here, I'd love to hear more about the actual thing that you're building.
[1]: https://techcrunch.com/2022/11/15/motherduck-secures-investm...
For a preview of what we're doing, on the technical side, a couple of our engineers gave a talk at DuckCon last week in Brussels, it is on youtube here: https://www.youtube.com/watch?v=tNNaG7e8_n8
(for context I'm the author of this blog post and co-founder of MotherDuck)
Assuming the above it true: I'll bet the reason they aren't so loud about exactly what they are doing is they want to get a head start on it. In theory anyone can build this stuff around DuckDB. From a marketing perspective the clever thing to do would be drive up usage of DuckDB while they build out all this functionality and then the minute corporates start seeing problems with their people using it (compliance etc), they have the solutions.
I think this is analytics equivalent of edge computing. Instead of one big-cluster cruching numbers.
1. User requests bunch of analytics
2. Server assembles a duckdb file
3. Sends this down to users laptop
4. User runs local queries on the duckfile
5. Go to step 1 for more analytics
My cheap no-name old laptop SSD writes with 170MB/s.
A customer has a name, address, email and order. Let's say 200 bytes for each. That means I can write 844000 new customers per second, far outside my personal marketing reach.
My disk is 240GB, which means I can store data for 1.2 billion customers. It'll take a while until I become that successful.
It will grow larger still if you include web logs from your e-commerce site and event data from your mobile app so that you can correlate these orders with items that customers considered but ultimately didn't buy. How will your laptop and SSD perform when you then build a user-item matrix to generate product recommendations for each of those 1.2 billion customers?
While plenty of organizations unnecessarily use Big Data tools to store and analyze relatively small amounts of data, there are plenty of customers with enough data to require them. I've seen plenty of them firsthand.
Right now I work for one of them: a global investment bank.
Within that organisation we have at least 100+ Spark clusters across the organisation doing distributed compute. And at least in our teams we have tight SLAs where a simple Python script simply can't deliver the results quick enough. Those jobs underpins 10s of billions of dollars in revenue and so for us money is not important, performance is.
So 1000 x 100 = 100,000 teams, all of whom I speak for, disagree with you.
These types of questions are orders of magnitude faster with a distributed backend.
Hold your horses... the beefiest servers that are in production today, unless you count custom-made stuff go to somewhere between 128 and 256 cores per board. These are hugely expensive. Also, I don't know if you can rent those from Amazon.
Typical, affordable servers range between 4..16 cores. Doesn't matter if you buy them yourself, or you rent them from Amazon. It's much cheaper to command a fleet of affordable servers than to deal with a high-end few. This is both because the one-time price of buying is quite different and because with smaller individual servers you have a fighting chance to scale your application with demand. Especially this is true in case of Amazon as you could theoretically buy spot instances and by so doing you'd share the (financial) load with other Amazon's customers.
Now... storage. Well, you see, in Amazon you can get very expensive storage that's guaranteed to be "directly" attached to the CPU you rent, the so-called ephemeral storage. This is the storage that's included with the VM image you use. It's very hard to get a lot of it. I couldn't find the numbers for Amazon, instead, I know that Azure tops out at 2 TB. In principle, this kind of storage cannot exceed a single disk, so, think Amazon probably offers the same 2 TB, maybe 4. But, again, it's cheaper to have a bunch of EBS's attached... but then you'll have to have more of them as the latency will suffer, and in order to compensate for that you would try to increase throughput, perhaps.
Also, think that, in practice, you'd want to have a RAID, probably RAID5, and this means you need upwards from 3 disks. Also, if you are using something like a relational database, you'd most likely want to put the OS on a single device, the database data on a RAID and database journal on a yet another device, and, probably, you'd want that device to be something like persistent memory / optane / something from higher-tier disks with dedicated power supply. And all this is not due to size, but due to different contingencies you need to have in order to prevent huge data loss... Now, add to this backups and snapshots, perhaps replication in 2-3 different geographical areas if you are running an international business... and that's quite a bill to foot.
There are similar problems with memory, since there can only be so many legs on memory bus and only so many pieces of memory you can attach to a single CPU, and if you also want a lot of storage, then, similarly, there can be only so many individual storage devices attached and so on.
Bottom line... even to reproduce the performance of your laptop in the cloud you would probably end up with some distributed solution, and you would still struggle with latency.
That's 8.4 hojillion megabytes per second right there.
That is not random access speed. For random access my relatively high-performance SSD only does 42MB/s reading and 80MB/s writing.
At first I tried to puzzle out a good sampling strategy to make sure I didn't bias the output, then on a whim I tried 2^32 samples and went to lunch. It took something like a half an hour to do 4 billion samples. Took me a couple times to figure out how to squeeze 4k megapixels into a graph so I ran it a few more times, but the results showed a very distinct banding pattern that confirmed that the problem was every bit as bad as I suspected, which was a blocking issue for our release. A couple of hours well spent, running through an 'intractable problem' that really wasn't.
Using any of the traditional SQL databases takes away a lot of complications. You can do transactions, you can query whatever you want, …
And if the database may get up to 1TB, still no problem with SQL. If exceed that, you may need a professional OPs team for your database and a few giant servers, but they should easily be able to go up to 10 TB, offload some queries to secondary servers, …
Both are distributed compute. Except that Spark allows you to mix/match code with SQL.
But linqpad is not useful if you don’t get the pro version, only then you get code completion. So it’s not really the answer to the problem.
I'm just awestruck, I could tell anyone that "large data" isn't really a bottleneck, but making sense of it is the very difficult part. My mentors kept pushing me to mention the sheer size of the datasets I process in talks because it sounds impressive, and I do do so, but I always knew it didn't matter because the interpretation and analysis is the hard part, not just the "sheer size."
[0] not going to use my real name
Especially now that 1TB datasets fit in memory on off-the-shelf servers and 100GB fits in memory on consumer hardware. You need a lot of data to run into real technical challenges that can't be solved by throwing a couple hundreed dollars a month (amortized cost) at hardware. And often you can get by with much, much less than even that.
At the same time, though, there was a lot of DNA sequencing data, we were designing CRISPR probes etc. But Spark and Hadoop aren't really that helpful in this area, so the Big Data team wasn't involved in those.
This paper is 8 years old and it was somewhat obvious then.
Scalability! But at what COST? https://www.usenix.org/system/files/conference/hotos15/hotos...
A big single machine can handle 98% of peoples data reduction needs. This has always been true. Just because your laptop only has 16GB doesn't mean you need a Hadoop (or Spark, or Snowflake) cluster.
And it was always in the best interest of the BD vendors and Cloud vendors to say, "collect it all" and analyze on/or using our platform.
The future of data analysis is doing it at the point of use and incorporating it into your system directly. Your actionable insights should be ON your grafana dashboard seconds after the event occurred.
I got sucked into "weekly key metric takes over 14 hours to run on our multi-node kubernetes cluster" a while back. I'm not sure how many nodes it actually used, nor did I really care.
Digging into it, the python code ingested about ~50GB of various files, made well over a dozen copies of everything, leaving the whole thing extremely memory starved. I replaced almost all of the program with some "grep | sed | awk | sed | grep" abomination that stripped about 98% of the unnecessary info first and it ran in under 2 minutes on my laptop. I probably should have tightened it up more but I was more than happy to wash my hands of the whole thing by that point.
Instead of improving the code, they just kept tossing more compute at it. Still heard all kinds of grumbling about os.system('grep | sed | awk | sed | grep') not being "pythonic" and "bad practice"; but not enough that they actually bothered to fix it.
1) 10 years ago, having access to 300tb of data that could sustain 10gigabytes/s of throughput would require something like two racks of disks with some SSD cache and junk.
2) people thought hadoop was a good idea
3) People assumed that everything could be solved with map:reduce
3) machine learning was much less of a thing.
4) people realised that postgres does virtually everything that mongo claimed it could.
5) people realised that cassandra was a very expensive way to make a write only database.
I gave a talk about using big data, and basically at the time the best definition I could come up with was "anything that's too big to reasonably fit in one computer. so think 4, 60 disk direct attached SAS boxes".
Most of the time people were chasing the stuff for the CV, rather than actually stopping to think if it was a good idea. (think k8s two years ago, chatGPT now, chat bots in 2020). Most buisnesses just wanted metrics, and instead of building metrics into the app, they decided to boil the ocean by parsing unstructured logs.
Not surprisingly it turned to shit pretty quick. Nowadays people are much better at building metrics generation directly into apps, so its much easier to easily plot and correlate stuff.
I don't know much about MotherDuck's plans, but I hope they're focused on making it as easy to collaborate on "small data" as Snowflake/etc. have made it to collaborate on "big data".
mongo's got sharding out of the box - which is nice - but you have to get your key right or it will suck.
Also no one should want to host a mongo db - unless that's your business.
And if there is any levelling off it's going to be because of the move towards cloud managed options e.g. Snowflake, DocumentDB rather than because PostgreSQL decided to add JSONB support.
[1] https://www.macrotrends.net/stocks/charts/MDB/mongodb/revenu...
https://github.com/frankmcsherry/blog/blob/master/posts/2015...
Also, using BigQuery as a metric of how Big Data is used is, IMHO, wrong. Real analytics companies usually have custom solutions because BigQuery is too expensive for any serious usage unless you are Google.
If anyone wants to follow along, the series is here!
https://alexpetralia.com/2023/01/19/working-with-data-from-s...
Unfortunately, I've found that many data teams focus more on making the data clean and available. They never drive the conversation about what actions are being taken with the data. That leads to them being treated as cost centers. Wrote a similar post about my perspective on it - https://bytesdataaction.substack.com/p/transform-your-data-t...
I'd love to chat about the space more with you if you're interested! Email in bio.
Making data the centerpiece of your business business could mean that your effectiveness of business process could increase several order of magnitudes. Funny thing is, you will not use some else’s model, unless you are building a ChatBox to infer, but you will need to build your own model and be trained in your own business process to be successful.
Consider a bank, here is my prediction of expected outcomes:
Enhanced Customer Experience: The system can act as a virtual banking assistant, providing customers with instant access to their account information, real-time transactions, and balance updates. The system can also answer customer inquiries and provide relevant information, improving the overall customer experience. Improved Fraud Detection: The system can monitor the bank's financial transactions in real-time and identify any potential fraud, helping the bank reduce its exposure to financial losses.
Automated Loan Processing: The system can analyze loan applications, credit scores, and other relevant data to approve or reject loan applications in real-time, reducing the time and effort required for manual loan processing. Personalized Marketing: The system can analyze customer behavior, transaction history, and demographic information to provide personalized marketing and cross-selling opportunities, increasing the bank's revenue and customer loyalty.
Real-Time Insights: The system can provide real-time insights into the bank's financial performance, customer behavior, and market trends, enabling the bank to make informed decisions and respond to market changes quickly.
What is interesting to me is, this is just the beginning of what could be…
There are lots of interesting things that can happen with "big streaming" than necessarily "big data". Like, cybersecurity is evolving to monitoring and reacting what everyone's machine is doing in the last 15 minutes, instead of having a huge database of hashes you trust. But not a ton of things really utilize what happened, say, 10 years ago on people's machines.
There's definitely some things that can use massive archives of old data, but I have found far, far fewer things that would benefit from it, and often that comes with some very big maintenance hassles. Most of the time, you can just set data retention to 30 days and be done.
They've been working to implement your ideas for decades and none of it requires LLMs or any machine learning techniques. Basic old ETL is more than sufficient.
The issue is that (a) the calculations they need to perform are complex and take time to run (b) there are financial regulations that weave its way through those system and (c) there is a lot of legacy code especially in the core ledger system which "just works" and people are reluctant to touch.
That said depending on your bank you can get real-time account activity, loan approvals in < 5 minutes etc.
But in this process, you don’t need ETL, nor all the process and development to accomplish these ideas. Conceptually the idea builds its self (it learns) how to threat the data, quite revealing and near real time. Considering you account for security and privacy, then you basically shift your input into the data stream and using a natural language get the data output you need, not clunky apps.
Imagine I just login, and say: me>how much do I have? bank>You have 100$ me>Please send 50$ to 1003 bank> Are you sure? Please add your security code to confirm
bla bla…
All this with little intervention.
Banks spend hundreds of man hours developing a lacking application while delivering a very poor customer experience. They spend millions on running decades old applications because it so expensive to change them… and thus the circle continues…
I’m really exited to see DataBases disappear conceptualy, data entry, mostly all that just disappear… I will ask my ChapBot for statement, give me a personal investment advice, and classify all my purchases and see where my wife has been spending all my money, all from the confort of my phone.
it’s a brave new wold we are wakening up to, that to me is exciting. And coming from having helped several major banks build their infrastructure, it’s just a boost to talk about something fresh, no more Hypervisor, core count, db licenses, ect. Ok, I’ll concede it’s pretty much the same old, just the nemonics will be different… How many GPUs, how quickly can you spin a container, how fast if your S3 datastore… oh wait, there is that circle again… >:D
In that case, chat-bots have existed for years and consumers largely don't like them.
In your scenario you can transfer money in a few clicks rather than having to write out an entire conversation.
Exactly, and I'd go further.
Are you in the perf/scale/data one percent?
So many people worry about scaling when in reality 99% of web apps will never reach above 100reqs/s.
I've been in web dev for 20+ years. Only once when working for a big international corporate client I had to worry about traffic spikes. And that was just for one of their multiple web apps.
Also, who thought their company would cease to function because surely they will hit google-scale dataset-sizes in the near future? Impossible for most except the biggest of the biggest
The current version of that article states: "There is no absolute amount of data that can be cited. For example, one cannot say that any database with more than 1 TB of data is considered a VLDB. This absolute amount of data has varied over time as computer processing, storage and backup methods have become better able to handle larger amounts of data.[5] That said, VLDB issues may start to appear when 1 TB is approached,[8][9] and are more than likely to have appeared as 30 TB or so is exceeded.[10]" https://en.wikipedia.org/wiki/Very_large_database
There is probably a rational, well thought out classification of different types of data bigness, as in CERN-big, Google-big, MegaBank-big, down to wordpress-log big and on the basis of that one would probably find that different designs are indispensable, address different pain points and cannot really "die". Hype has a more erratic lifecycle than real needs
The implicit problem is that even if the dataset fits in memory, the software processing that data often uses more RAM than the machine has. And unlike using too much CPU, which just slows you down, using too much memory means your process is either dead or so slow it may as well be. It's _really easy_ to use way too much memory with e.g. Pandas. And there's three ways to approach this:
* As mentioned in the article, throw more money at the problem with cloud VMs. This gets expensive at scale, and can be a pain, and (unless you pursue the next two solutions) is in some sense a workaround.
* Better data processing tools: Use a smart enough tool that it can use efficient query planning and streaming algorithms to limit data usage. There's DuckDB, obviously, and Polars; here's a writeup I did showing how Polars uses much less memory than Pandas for the same query: https://pythonspeed.com/articles/polars-memory-pandas/
* Better visibility/observability: Make it easier to actually see where memory usage is coming from, so that the problems can be fixed. It's often very difficult to get good visibility here, partially because the tooling for performance and memory is often biased towards web apps, that have different requirements than data processing. In particular, the bottleneck is _peak_ memory, which requires a particular kind of memory profiling.
In the Python world, relevant memory profilers are pretty new. The most popular open source one at this point is Memray (https://bloomberg.github.io/memray/), but I also maintain Fil (https://pythonspeed.com/fil/). Both can give you visibility into sources of memory usage that was previous painfully difficult to get. On the commercial side, I'm working on https://sciagraph.com, which does memory and also performance profiling for Python data processing applications, and is designed to support running in development but also in production.
This has been - to varying extends - my own experience working in large organizations that don't have tech as their core business.
Although there are some successful data analysis project, the potential of the collected data remains largely underutilized.
Right on point. In the past I have been obsessed with big data, looking for insights. Then I realized that a medium-sized specific data set is always better than a gargantuan general big data monster. There is so many applications in my field where only outliers matter anyways, and everything is very "centralized" to a few relevant observations. So the only thing about big data is that you maybe throw away 99.9% of the data right away and then you have some observations that you actually care about. There is soooo much data out there that is just noise, and so little that I actually care about. And that's why I still end up hand collecting stuff every now and then.
If you have no tangible capabilities to do above, asking customer "ARE YOU IN THE BIG DATA ONE PERCENT?" will be the quickest way out of the door.
I love use cases like the Rill Data (https://youtube.com/watch?v=XvP2-dJ4nVM), where you can suddenly run analytics with a single cmd line prompt and see your data just instantly visualized. Such use cases are only possible because of the "small" data approach that DuckDB tries.
Sighs...
Yes, thank goodness that part is dead. But meanwhile - we've still got more actual data than ever to store, and ever-tighter deadlines on finding and delivering it. If we can get back to that and let the PySpark bootcampers fade away, maybe things can get a little better for once.
In other words:
Even when querying giant tables, you rarely end up needing to process very much data. Modern analytical databases can do column projection to read only a subset of fields, and partition pruning to read only a narrow date range. They can often go even further with segment elimination to exploit locality in the data via clustering or automatic micro partitioning. Other tricks like computing over compressed data, projection, and predicate pushdown are ways that you can do less IO at query time. And less IO turns into less computation that needs to be done, which turns into lower costs and latency.
Big data is "dead" because data engineers (the programming ones, not the analysts-in-all-but-title) spent a ton of effort building DBs with new techniques that scale better than before, with other storage patterns than before. Someone still has to write and maintain those! And it would be even better if those tools and techniques could escape the half dozen major data cloud companies and be more directly accessible to the average small team.
I think there is a problem when someone with such proclaimed knowledge of the sector gets to this, and similar, pieces of data, and does not attribute it to pricing. Could it be queries are short because bigquery pricing for analysis, as confusing as this models are, is based on amount of data?[0]
Because the other line of reasoning is that a big chunk of that 90% of professionals being paid to do their jobs, do NOT take into account pricing of the tool and are using it for small data, instead of thinking that people are using the best tool with the lowest price, because there's plenty of options to process and analyse data right now in the cloud.
On the "business have low amount of data", that matches my experience as well. At first I thought I was simply dealing with smaller sized companies, but it's a trend of doing big data projects for data that'd fit a pendrive.
[0] https://cloud.google.com/bigquery/pricing#analysis_pricing_m...
It usually does not take many data points for an actionable insight and most actions then will invalidate small details in old data anyhow. Better to start every round with fresh eyes.
What do these blockchains do that have to keep data around forever, with high throughput, and need to expose it quickly do? Are you saying they should delete parts of data in the chain?
Seriously, I've spent my career working on big data systems, and while the answer is sometimes "yes you need to delete your data", I don't think that's going to always work.
"You probably don't actually have big data" is a very valid point, not that many organizations do - most businesses haven't generated enough actionable data in their lifetime to need more than a single beefy machine without ever deleting data.
It seems an odd pitch in general to say, hey my product specifically performs poorly on large datasets.
Most orgs collect the data that is easy to collect, and they are extremely lucky if that happens to be the data that enables the insights they desire. When the data they really need looks too hard to get, the org tries to compensate by collecting more of the easy stuff, and hoping that if blood can’t be squeezed out of a stone, maybe it can be squeezed out of 100bn stones.
Volume != Quality
I’m no statistician, but I’m like 99% sure that’s an exponential, not a power law
There’s a world of difference. The point of an exponential is that you can ignore big things. The point of a power law is that you can’t.
> I’m no statistician, but I’m like 99% sure that’s an exponential, not a power law
I'm no expert either, but it seems correct. The power law distribution has each X value in an X/Y series decreasing by a specific factor, like this: https://en.wikipedia.org/wiki/Power_law#/media/File:Long_tai...
The exponential has each X value increasing by a specific factor, like this: https://en.wikipedia.org/wiki/Exponential_function#/media/Fi...
This post doesn't go far enough. It challenges the assumption that everyone's data is "big data" or that every company's data will eventually grow to be big data. I agree that "big data" was the wrong model. We also need to challenge that all data should be stored in one place (warehouse, lake, lakehouse). We need to challenge that one tool can be used for every data need. We need to challenge how we build systems both from a technology and people standpoint. We need to embrace that the problems and needs of companies _are always changing_.
We are living with conceptual inertia. Many of our patterns are an evolution from the 70's and 80's and the first relational databases. It's time to rethink how we "do data" from first principles.
We've gotten to a point where the first and last step get skipped. Business leaders see other companies doing interesting things with data, so the answer must be "gather all the data"! Internal teams end up focused on gathering the data without the context of how it might be used.
We need to train data teams to not focus on the data as the product. Instead, they should be responsible for driving business actions. Gathering and cleaning the data should just a byproduct of that activity.
We don't use any "BigData" products yet, as there wasn't any need for them, even when we provide full search and relatively nice and rich set of analytics over all the data. Yet, based on the article, we're way above most of the companies relying heavily on such tools. Confusing.
A trick I saw is companies hiring experienced jack-of-all-trades back-end engineers into Data teams. A lot of things get migrated from Spark to Postgres, from Kafka to REST API calls, and keep working fine and become generally more responsive.
I'm on the same page as the author here: traditional BigData tech has its place and its uses, but before choosing it companies (CTOs, architects) should carefully consider if it is necessary, especially considering the cost of it and the risk of locking themselves down in a very specialized domain.
In the meantime SSD storage took off, so the IOPS from a stock drive have skyrocketed, business domains for large data sets have broadened beyond click/impression streams, and the challenge now is not "can I store all this data" it's "WTH do I do with it?"
Regardless of quantity of data, structuring and analysis and querying of said data remains paramount. The challenge for anybody working with data is to represent and extract knowledge. I remain convinced that logic -- first order logic and its offshoot in the relational model -- remains the best tool for reasoning about knowledge. Codd's prognostications on data from the 1970s are still profound.
I think we're in a space now where we can turn our attention to knowledge management, not just accumulating streams of unstructured data. The challenge in a business is to discover and capture the rules and relationship in data. SQL is an existing but poor tool for this, based on some of the concepts in the relational model but tossing them together in a relatively uncomposable and awkward way (though it remains better than the dogs breakfast of "NoSQL" alternatives that were tossed together for a while there.)
My employer is working in this space, I think they have a really good product: https://relational.ai/
It is about using modern tools (ClickHouse) for data engineering without the fluff - when you can take whatever dataset or data stream and make what you need without the need for complex infrastructure.
Nevertheless, the statement "big data is dead" is short-sighted, and I don't entirely follow this opinion.
For example, here is one of ClickHouse's use-case:
> Main cluster is 110PB nvme storage, 100k+ cpu cores, 800TB ram. The uncompressed data size on the main cluster is 1EB.
And when you have this sort of data for realtime processing, no other technology can help you.
Such a server could readily fit 384 threads with dual EPYC and would be available with enough RAM to keep more than 10% of that data in cache.
Your workload absolutely fits the definition of "fits on one machine".
I am not suggesting here that you should put it on one machine. You definitely could, though.
It doesn't help that we've shifted into a climate where hoarding data comes with a huge regulatory and compliance price tag, not to mention risk. But if the value was there we would do it, so this is not the primary driver.
Mission accomplished more than big data is dead IMHO.
Well, obviously, the realization that management is largely emotion driven and little data driven, is a prelude for the CEO-AI yet in the makings.
Of course this still got a face to it. A CEO who speaks and talks, as the voice commands, but does not do the part that even humans who think they are good at it, are bad at, decision making. The ground truth is there ("Cooperate history") going back to the merchants of sumeria. Lets learn that lesson, pack it into a decission tree, and wrap that bundle with Chat GPT smooth talking.
Although I don't think most organizations are blaming lack of actionable insights on the data size. It's the lack of prioritizing data usage over data accessibility. We need to be teaching data people business levers and teaching business people data levers.
Data should be a byproduct of an actionable idea that you want to execute. It shouldn't exist until you have that experiment in mind.
Logs are where data goes to die.
The article does allude to this definition when it states that "Most data is rarely queried". We have become data hoarders. Technology has made it easy (and relatively cheap) to store data, but the ideas of what to do with this data have not scaled in comparison.
Eye-opening. Especially when combined with a recent quote from Satya Nadella, "First, as we saw customers accelerate their digital spend during the pandemic, we’re now seeing them optimize their digital spend to do more with less."
Conclusion: SaaS is easy to drop off in downturns. Just as easy as it is to buy initially.
https://github.com/ClickHouse/ClickHouse/issues/45741#issuec... helps with that though.
Also, clickhouse-local exists https://clickhouse.com/blog/extracting-converting-querying-l... as a thing.
But, yes, when I think of DuckDB...I think embedded use cases...i'm also not a power user.
I also think of this very much as a 'horses for courses' or 'different strokes, different folks' sort of scenario. There is, naturally, overlap because 'analytical data.' But also, there is naturally overlap with R and this giant scary mess of data-munging PERL code I maintain for a side project.
The DuckDB team, the MotherDuck team, the ClickHouse team...we all want your experience interacting with data to be amazing. In some scenarios, ClickHouse is better. In some scenarios, DuckDB. I'm biased (as I work for ClickHouse in DevRel), but I <3 ClickHouse.
Try both. Pick the one that is best for you. Then...you know...tell the other(s) why so that we all can get better at what we do.
Meanwhile for example boilingdata.com seems to have already done that - by using AWS Lambda + DuckDB as distributed compute engine which I can't decide if its awesome, deranged or both.
[1] https://duckdblabs.com/news/2022/11/15/motherduck-partnershi...
SSDs we’re limited in capacity and still expensive.
Parallelizing work with MapReduce allowed using cheap fault-prone commodity hardware and disks.
If you’re dealing with terabytes rather than petabytes of data, you probably don’t need BigData
Can someone explain why this is the case? Is it due to more replications or maintaining more indices?
Hoarding is not a winning strategy.
wow, ya think? Must have been eye opening to see all those customers with a few million rows thinking they had "Big data" huh?
Don't get me wrong, I love it. It's about time people got off these stupid and shockingly expensive bandwagons.
Big data was vendor generated hype that convinced many engineers to confuse the size of their dataset with their, ahem, shoe size.
They didn’t do their employers any favors.
K, good luck!
Big data hype never felt to me like anything more than a hype campaign to help big tech research ML/AI.
Larry even rambled as much: https://arstechnica.com/information-technology/2013/05/larry...
It appears not all Googlers got the memo.
Everyone else is in the way of him solving big problems! Not like such work could not be distributed among technologists and researchers around the globe via the internet. Help Google do it!
I am leaning into “Deep Work” going forward; will slowly iterate on my own model creation and collaborate with like minded folks. I’m fucking done with intentionally empowering billionaire minority who convinced an ignorant political gerontocracy that minority is capable of magic.
Anyone prattling on with common tropes of “longtermism”; nation state nutters, religious, technocrats; are appealing to non-existent authority they see a a magical future for us! Give them your money to insure it arrives! They have zero ability to insure such outcomes and a lot of upside to making people believe such today.