Big data is dead (2023)
motherduck.com
motherduck.com
I patiently listened to all the big query hadoop habla-blabla, even asked questions about the financials (hardware/software/license BOM) and many of them came up with astonishing tens of thousands of dollars yearly.
The winner of course was the guy who understood that 6TiB is what 6 of us in the room could store on our smart phones, or a $199 enterprise HDD (or three of them for redundancy), and it could be loaded (multiple times) to memory as CSV and simply run awk scripts on it.
I am prone to the same fallacy: when I learn how to use a hammer, everything looks like a nail. Yet, not understanding the scale of "real" big data was a no-go in my eyes when hiring.
Problem is most students don’t listen to the first part “for the purpose of this course”. The professor does not elaborate because that is beyond the scope of the course.
But no, this particular person had been working professionally for decades (in fact, he was much older than me).
I would just back-track from a shipped product date, and try to guess who we needed to get there... given the scope of requirements.
Generally, process people from a commercially "institutionalized" role are useless for solving unknown challenges. They will leave something like an SAP, C#, or MatLab steaming pile right in the middle of the IT ecosystem.
One could check out Aerospike rather than try to write their own version (the dynamic scaling capabilities are very economical once setup right.)
Best of luck, =3
It’ll be slower, perhaps by a lot, but most “big data” stuff is already so god damned slow that mmap probably still beats it, while being immeasurably simpler and cheaper.
For CSV it flat out doesn't matter what you do since the format is so inefficient and needs to be read start to finish, but something like parquet probably benefits from explicit read syscalls, since it's block based and highly structured, where you can predict the read patterns much better than the kernel can.
But yeah, they might have that much RAM. At a rather small company I was at we had a third of it in the virtualisation cluster. I routinely put customer databases in the hundreds of gigabytes into RAM to do bug triage and fixing.
[0] https://aws.amazon.com/blogs/aws/now-available-amazon-ec2-hi...
[1] https://cloud.google.com/blog/products/sap-google-cloud/anno...
[2] https://azure.microsoft.com/en-us/updates/azure-mv2-series-v...
No business is passing 6tb data around on their laptops.
And how that price would compare to the equivalent big data solution in the cloud.
So yes, if need to be, I have 6 TiB of RAM.
To mitigate these bottlenecks you get fancy hardware (e.g oracle appliance) or you scale out (and get TCO/performance gains from separating storage and compute - which is how Snowflake sold 3x cheaper compared to appliances when they came out).
I believe that Trino on HDFS would be able to finish faster than awk on 6 enterprise disks for 6TB data.
In conclusion I would say that we should keep data local if possible but 6TB is getting into the realm where Big Data tech starts to be useful if you do it a lot.
The point of the article is 99.99% of businesses never pass even the 10 Gb point though.
I get that that post is only on 3.5GB, but, consumer SSDs are now much faster at 7.5GB/s vs 270MB/s HDD back when the article was written. Even with only mildly optimised solutions, people are churning through the 1 billion rows (±12GB) challenge in seconds as well. And, if you have the data in memory (not impossible) your bottlenecks won't even be reading speed.
[1]: https://adamdrake.com/command-line-tools-can-be-235x-faster-...
And I am not a specialized data scientist. But with time I am wondering if such a thing even exists... being a good backender / sysadmin and knowing a lot of CLI tools has always seemed to do the job for me just fine (though granted I never actually managed a data lake, so I am likely over-simplifying it).
...or put it into SQLite for extra blazing fastness! No kidding.
I had a heck of a time running the server locally before I discovered the CLI.
A good answer that strikes a balance between size of data, latency and frequency requirements is a candidate who is able to show that they can choose the right tool that the next person will be comfortable with.
Obviously organically grown Frankenstein programs are a huge liability, I think every reasonable techie agrees on that.
Check out "data science at the command line":
Lol. Data management is about safety, auditablity, access control, knowledge sharing and who bunch of other stuff. I would've immediately shown you the door as someone who i cannot trust data with.
Your entire post reeks of "I'm smarter than you" smugness while at the same time revealing no useful information or approaches. Near as I can tell no one should trust you with any data.
unlike "blows my mind" ?
> As stated the question didn't require any of what you outline here.
Right. OP mentioned it was "tricky question" . What makes it tricky is that all those attributes are implicitly assumed. I wouldn't interview at google and tell them my "stack" is "load it on your laptop". I would never say that in an interview even if I think that's the right "stack" .
You are assuming you know what the OP meant by tricky question. And your assumption contradicts the rest of the OP's post regarding what he considered good answers to the question and why.
I guess it wasn't but even if so, it would be legitimately baffling how people manage to project so much negativity in three words that are slightly tongue-in-cheek casual comment on the state of affairs in an area whose value is not always clear (in my observations, only after you start having 20+ data sources it starts to pay off to have dedicated data team; I've been in teams only 3-4 devs and we still managed to have 15-ish data dashboards for the executives without too much cursing).
An anecdote, surely, but what isn't?
I was never a data scientist, just a guy who helped whenever it was necessary.
No. You qualified it with "blows my mind" . Why would it 'blow your mind' if you don't have any data background.
This negative cherry picking does not do your image any favors.
buddy, you're just rolling off buzzwords and lording it over other people
No need to act smug and superior, especially since nothing about OP's plan here actually precludes having all the nice things you mentioned, or even having them inside $your_favorite_enterprise_environment.
You risk coming across as a person who feels threatened by simple solutions, perhaps someone who wants to spend $500k in vendor subscriptions every year for simple and/or imaginary problems... exactly the type of thing TFA talks about.
But I'll ask the question.. why do you think safety, auditablity, access control, and knowledge sharing are incompatible with CLI tools and a specific choice of file system? What's your preferred alternative? Are you sticking with that alternative regardless of how often the work load runs, how often it changes, and whether the data fits in memory or requires a cluster?
I responded with the same tone that gp responded with. "blows my mind" ( that people can be so stupid) .
> Consulting service: you bring your big data problems to me, I say "your data set fits in RAM", you pay me $10,000 for saving you $500,000.
A lot of industry work really does fall into this category, and it's not controversial to say that going the wrong way on this thing is mind-blowing. More than not being controversial, it's not confrontational, because his comment was essentially re: the industry, whereas your comment is directed at a person.
Drive by sniping where it's obvious you don't even care to debate the tech itself might get you a few "sick burn, bro" back-slaps from certain crowds, or the FUD approach might get traction with some in management, but overall it's not worth it. You don't sound smart or even professional, just nervous and afraid of every approach that you're not already intimately familiar with.
"not understanding the scale of "real" big data was a no-go in my eyes when hiring." , "real winner" ect.
But yea you are right. I shouldn't have directed it at commenter. I was miffed at interviewers who use "tricky questions" and expect people to read their minds and come up with their preconceived solution.
If you really must know: I said "blows my mind [that people don't try simpler and proven solutions FIRST]".
I don't know what do you have to gain to come here and pretend to be in my head. Now here's another thing that blows my mind.
Well why don't people do that according to you ?
Its not 'mind blowing' to me because you can never guess what angle interviewer is coming at you. Especially when they use the words like ' data stack'.
As for interviews, sure, they have all sorts of traps. It really depends on the format and the role. Since I already disclaimed that I am not actual data scientist and just a seasoned dev who can make some magic happen without a dedicated data team (if/when the need arises) then I wouldn't even be in a data scientist interview in the first place. ¯\_(ツ)_/¯
"not understanding the scale of "real" big data was a no-go in my eyes when hiring."
Why would you guess in that situation though?
It’s an interview, there’s at least 1 person talking to you — you should talk to them, ask them questions, share your thoughts. If you talking to them is a red flag, then high chances that you wouldn’t want to work there anyway.
My comment wasn't directed at parent. I was trying to be smart and write an inverse of original comment. Opposite scenario Where I as an interviewer was looking for a proper 'data stack' and interviewee responded with a bespoke solution.
"not understanding the scale of "real" big data was a no-go in my eyes when hiring."
i was trying to point out that you can never know where the interviewer is coming from. Unless i know interviewer personally i would bias towards playing it safe and go with 'enterpisey stack'
I've generally liked BigQuery for this type of stuff - the console interface is good enough for ad-hoc stuff, you can connect a plethora of other tooling to it (Metabase, Tableau, etc). And if partitioned correctly, it shouldn't be too expensive - add in rollup tables if that becomes a problem.
And most laptops have 4 CPU cores these days, and a multiprocess operating system, so you don’t have to wait for random access on a spinning plate to find every bit in order, you can simply have multiple awk commands running in parallel.
Awk is most certainly a better user interface than whatever custom BrandQL you have to use in a textarea in a browser served from localhost:randomport
If we're talking about 6 TB of data:
- You can upgrade to 8 TB of storage on a 16-inch MacBook Pro for $2,200, and the lowest spec has 12 CPU cores. With up to 400 GB/s of memory bandwidth, it's truly a case of "your big data problem easily fits on my laptop".
- Contemporary motherboards have 4 to 5 M.2 slots, so you could today build a 12 TB RAID 5 setup of 4 TB Samsung 990 PRO NVMe drives for ~ 4 x $326 = $1,304. Probably in a year or two there will be 8 TB NVMe's readily available.
Flash memory is cheap in 2024!
There are relatively cheap adapter boards which let you stick 4 M.2 drives in a single PCIe x16 slot; you can usually configure a x16 slot to be bifurcated (quadfurcated) as 4 x (x4).
To pick a motherboard at quasi-random:
Tyan HX S8050. Two M.2 on the motherboard.
20 M.2 drives in quadfurcated adapter cards in the 5 PCIe x16 slots
And you can connect another 6 NVMe x4 devices to the MCIO ports.
You might also be able to hook up another 2 to the SFF-8643 connectors.
This gives you a grand total of 28-30 x4 NVME devices on one not particularly exotic motherboard, using most of the 128 regular PCIe lanes available from the CPU socket.
Ideally if the data is laid out optimally on the spinning disk, a single process reading the data would result in a mostly-sequential read with much less time wasted on read head repositioning seeks.
In the odd case where the HDD throughput is greater than a single-threaded CPU processing for whatever reason (eg. you're using a slow language and complicated processing logic?), you can use one optimized process to just read the raw data, and distribute the CPU processing to some other worker pool.
You shouldn't have to set up a cluster for data jobs these days.
And it kind of points out the reason for going with a data scientist with the toolset he has in mind instead of optimizing for a commandline/embedded programmer.
The tools will evolve in the direction of the data scientist, while the embedded approach is a dead end in lots of ways.
You may have outsmarted some of your candidates, but you would have hired a person not suited for the job long term.
SQL should be fine for that.
Actually, I have a feeling that the awk solution will struggle if there are many unique keys.
For example if they in that dataset have a million customers and want to extract the top 10. Then there is an intermediate map stage that will be storage or memory consuming.
It is like matrix multiplication. Calculating the dot product is trivial, but when the matrix has n:m dimensions and n,m starts to grow, it becomes more and more resource heavy. And then the laptop will not be able to handle it.
(in the example, m is the number of rows, and n is the number of unique customers. The dot product is just a sum over one dimension, while the group by customer id is the tricky part)
It's a couple hundred bucks a month and $36 to query the entire dataset, after partitioning thats not terrible.
You can always save an intermediate data set partitioned and massaged into whatever format makes subsequent queries easy, but that's usually application-dependent, and so you want that control over how you actually store your intermediate results.
If you only needed this once, the BQ approach requires very little setup and many places already have a billing account. If this is recurring then you need to figure out what the ownership plan of the hard drive is (what's it connected to, who updates this computer, what happens when it goes down, etc.).
I later moved to a company with a fairly sophisticated setup with Databricks. While Databricks offered some QoL improvements, it didn't magically make all my queries run quickly, and it didn't allow me anything that I couldn't have done on the remote desktop setup.
Just dump it into Oracle, postgre, mssql, or mysql and be amazed by the kind of things you can do with 30year old data analysis technology on an modern computer.
The extend people goes to to not recognize how much the people creating the SQL language and the relational database engines we now take for granted actually knew what they were doing, are a bit of an mystery to me.
The right answer to any query that can be defined in SQL is pretty much always an SQL engine even if it's just sqlite running on an laptop. But somehow people seems to keep comming up with reasons not to use SQL.
enter duckdb with columnar vectorized execution and full SQL support. :-)
disclaimer: i work with the author at motherduck and we make a data warehouse powered by duckdb
> Consulting service: you bring your big data problems to me, I say "your data set fits in RAM", you pay me $10,000 for saving you $500,000.
This is from 2015...
[0] https://www.tomshardware.com/reviews/samsung-980-m2-nvme-ssd...
[1] https://www.tomshardware.com/reviews/ram-speed-tests,1807-3....
Once the index is too big to have in one place, thing get more complicated.
You want something like PM1735: PCIe 4.0 x8, up to 8000 MB/s sequential read.
And while DDR5 is surely faster the question is what the data access patterns are there.
In almost all cases (ie mix of random access, occasional sequential reads) just reading from the NVMe drive would be faster than loading to RAM and reading from there. In some cases you would spend more time processing the data than reading it.
PS all these RAM bandwidth rates are good for the sequential access, as you go random access the bandwidth drops.
https://semiconductor.samsung.com/ssd/enterprise-ssd/pm1733-...
The value of our data skills are getting eroded!
It's a specialist role, and most people with the skills you seek are generalists.
It seems like it would get a lot of swap thrashing if you had multiple processes operating on disorganized data.
I'm not really a data scientist and I've never worked on data that size so I'm probably wrong.
What device do you have in mind? I've seen places use 2TB RAM servers, and that was years ago, and it isn't even that expensive (can get those for about $5K or so).
Currently HP allows "up to 48 DIMM slots which support up to 6 TB for 2933 MT/s DDR4 HPE SmartMemory".
Close enough to fit the OS, the userland, and 6 TiB of data with some light compression.
>It seems like it would get a lot of swap thrashing if you had multiple processes operating on disorganized data.
Why would you have "disorganized data"? Or "multiple processes" for that matter? The OP mentions processing the data with something as simple as awk scripts.
A better question would be:
Why would anyone stream 6 terabytes of data over the internet?
In 2010 the answer was: because we can’t fit that much data in a single computer, and we can’t get accounting or security to approve a $10k purchase order to build a local cluster, so we need to pay Amazon the same amount every month to give our ever expanding DevOps team something to do with all their billable hours.
That may not be the case anymore, but our devops team is bigger than ever, and they still need something to do with their time.
1 TB of memory is like 5 grand from a quick Google search then you probably need specialized motherboards.
Not necessarily - I might not want it or need it. It's a few TB, it can be on a fast HD, on an even faster SSD, or even in memory. I can crunch them quite fast even with basic linear scripts/tools.
And organized could just mean some massaging or just having them in csv format.
This is already the same rushed notions about "needing this" and "must have that" that the OP describes people jumping to, that leads them to suggest huge setups, distributed processing, multi-machine infrastructure, for use cases and data sizes that could fit on a single server with redundancy and be done it.
DHH has often written about this for their Basecamp needs (scalling vertically where others scale horizontally having worked for them for most of their operation), there's also this classic post: https://adamdrake.com/command-line-tools-can-be-235x-faster-...
>1 TB of memory is like 5 grand from a quick Google search then you probably need specialized motherboards.
Not that specialized, I've work with server deployments (HP) with 1, 1.5 and 2TB RAM (and > 100 cores), it's trivial to get.
And 5 or even 30 grand would still be cheaper (and more effective and simpler) than the "big data" setups some of those candidates have in mind.
Im just trying to understand the parent to my original comment.
How would running awk for analysis on 6TB of data work quickly and efficiently?
They say it would go into memory but its not clear to me how that would work as would still have paging and thrashing issues if the data didnt have often used sections of the data.
am I overthinking it and they were they just referring to buying a big ass Ram machine?
In that 6TB is not that huge of an amount
That's their total dataset, and there's no "real time" requirement.
They can start a batch process, process the data, and be done with it.
Here's an example of someone using awk (read further down for the relevant section):
https://livefreeordichotomize.com/posts/2019-06-04-using-awk...
"I was now able to process a whole 5 terabyte batch in just a few hours."
>They say it would go into memory but its not clear to me how that would work as would still have paging and thrashing issues if the data didnt have often used sections of the data
There's no need to have paging and thrashing issues if you can fit all (or even most) of your data in memory. And you can always also split, process, and aggregate partial results.
>am I overthinking it and they were they just referring to buying a big ass Ram machine?
Yeah, they said one can buy a machine with several TB of memory.
I'm not advocating that this is generally a good or bad idea, or even economical, but it's possible.
is this what they were referring to just by a big ass Ram machine?
Why not take a few days to get familiar with AWK, a skill which will last a lifetime? Like SQL, it really isn't so bad.
This effect then forces the person to be reliant on the LLM for answering all questions, and they’ll be less capable of figuring out more complex issues in the topic.
$20/mth is a siren’s call to introduce such a dependency to critical systems.
The only time this becomes an issue is when the data needs to be processed as close to real-time as possible. In those instances, I still tend to log the raw data to disk in another thread.
Not only that, it provides a useful separation from the storage format so you can use it to query a flat file exposed as table using Apache Drill or a file on s3 exposed by Athena or data in an actual table stored in a database and so on. The flexibility is terrific.
Why would use use bigquery just to get SQL?
You tricked candidates with your nonsensical scenario. Hate smartass interviewers like this that are trying some gotcha to feel smug about themselves.
Most candidates don't feel comfortable telling ppl 'just load on your laptops' even if they think thats sensible. They want to present a 'professional solution', esp when you tricked them with the word 'stack'. which is how most of them prbly perceived your trick question.
This comment is so infuriating to me. Why be assholes to each other when world is already full of them.
In my view, it's an appropriate question.
OP even mentioned that it his favorite 'tricky question' . It would def trick me because they used the word 'stack' which has specific meaning in the industry. There are even websites dedicated to 'stack's https://stackshare.io/instacart/instacart
Penny wise and pound foolish, plus a dash of NIH syndrome. When you're the only company doing something a particular way (and you're not Amazon-scale), you're probably not as clever as you think.
Most business have static sets of data that people load on their PCs. (Why do you assume laptops?)
The only weird part of that question is that 6TiB is so big it's not realistic.
The largest dataset I worked with was about 60TB
While that didn't fit in ram most people would just load the sample data into the cluster when I told them it would be faster to load 5% locally and work off that.
You don't get extra credit points for mind-reading, if anything you get more esteem for requirements-gathering, which would lead you towards either a professional solution or a laptop solution: whichever fits the business needs.
It might be a totally unreasonable question if it's provided context-free as a form on a screen, but it is a perfectly reasonable conversation-starter in an interview.
In terms of maturity of solution and diversity of downstream applications you'll go much further with BigQuery/Athena (at comically low cost) for this amount of data than some cobbled together "local" solution.
I thoroughly agree with the author but the comments in this thread are an indication of people who haven't actually had to do meaningful ongoing work with modest amounts of data if they're suggesting just storing it as plain text on personal devices.
I'm not advocating for complicated or expensive solutions here, BigQuery and Athena are very low complexity compared to any of the Hadoop et-al tooling (yes Athena is Trino is in the family, but it is managed and dirt cheap).
(On the other hand, I wouldn't really want the average business analyst walking around with all our customer data on their laptops all the time. And by the time you have a proper ACL system with audit logs and some nice way to share analyses that updates in real time as new data is ingested, the Big Data Solution™ probably have a lower TCO...)
I doubt it. The common Big Data Solutions manage to have a very high TCO, where the least relevant share is spent on hardware and software. Most of its cost comes from reliability engineering and UI issues (because managing that "proper ACL" that doesn't fit your business is a hell of a problem that nobody will get right).
I'm not sure there is a way to get this right unless there is a programmatic integration into the org chart, and ability to describe and parse in a declarative language the organizational rules of who has access to what, when, under what auth, etc. It has otherwise been for me an exercise in watching massive amounts of toil manually interpreting between the SOT of the org chart and all the other applications mediated by many manual approval policies and procedures. And at every client I've posed this to, I've always been denied that programmatic access for integration.
A lot of sites try to avoid this by designing ACL's around certain activity or data domains because those are more stable than organizations, but this breaks down when you get to the fine-grained levels of the ACL's so we get capped benefits from this approach.
I'd love to hear how others solve this in large (10K+ staff) organizations that frequently change around teams.
How could anyone answer this without knowing how the data is to be used (query patterns, concurrent readers, writes/updates, latency, etc)?
Awk may be right for some scenarios, but without specifics it can't be a correct answer.
I would not conclude that they over-engineer everything they do from such an answer, but rather just that they got tricked in this very artificial situation where you are in a dominant position and ask trick questions.
I was recently in a technical interview with an interviewer roughly my age and my experience, and I messed up. That's the game, I get it. But the interviewer got judgemental towards my (admittedly bad) answers. I am absolutely certain that were the roles inverted, I could choose a topic I know better than him and get him in a similarly bad position. But in this case, he was in the dominant position and he chose to make me feel bad.
My point, I guess, is this: when you are the interviewer, be extra careful not to abuse your dominant position, because it is probably counter-productive for your company (and it is just not nice for the human being in front of you).
It may not be reasonable to suggest that in a role that traditionally uses big data tools
The person who says "I'm going drive back to the restaurant and take my professional equipment home to cook the steak" is probably offering the wrong answer.
I'm obviously not a professional cook, but presumably the ability to improvise with whatever tools you currently have is a desirable skill.
It is a sort of “interview hack” example that’s been used to emphasize the idea of a simple unspecialized skill-test that went around a while ago. I guess upcoming chefs probably practice egg scrambling nowadays, ruining the value of the test. But maybe they could ask to make a bit of steak now.
What would be the equivalent for a technical interview? Perhaps: Implement a generic linked list or dynamic array.
E.g., if a HN'er takes this as advice they're just as likely to be gated by some other interviewer who interprets hedging as a smell.
I believe the posters above are essentially saying: you, the interviewer, can take the 2.5 seconds to ask the follow up, "... and if we're not immediately optimizing for scalability?" Then take that data into account when doing your assessment instead of attempting to optimize based on a single gate.
Edit: clarification
The few times I've ignored red flags from a company in an interview I've been in a world of pain afterwards.
Give both answers, and explain why the obvious "hard" answer is wrong
If people in high stakes environments interpret hedging as a smell - run from that company as fast as you can.
Hedging is a natural adult reasoning process. Do you really want to work with someone who doesn't understand that?
Last I heard theyd promoted one unix guy on the inside to baby sit a bunch of chron jobs on the biggest server they could find.
As an interviewer, why not asking: "how would you do that in a setup that doesn't have much data and doesn't need to scale, and then how would you do it if it had a ton of data and a big need to scale?". There is no trick here, do you feel you lose information about the interviewee?
As a sibling puts it though, it's a matter of level. Senior/staff and above? Yeah, that's mostly what you do. Lower than that, then you should be able to mostly trust those upper folks to have seen through the trick.
I don't know about you, but in my work, I always have more than 3 seconds to find a solution. I can slowly think about the problem, sleep on it, read about it, try stuff, think about it while running, etc. I usually do at least some of those for new problems.
Then of course there is a bunch of stuff that is not challenging and for which I can start coding right away.
In an interview, those trick questions will just show you who already has experience with the problem you mentioned and who doesn't. It doesn't say at all (IMO) how good the interviewee is at tackling challenging problem. The question then is: do you want to hire someone who is good at solving challenging problems, or someone who already knows how to solve the one problem you are hiring them for?
You may say "yeah but you just have to think out loud, that's what the interviewer wants". But again that's not how I work. If the interviewer wants to see me design a system, they should watch me read documentation for hours, then think about it while running, and read again, draw a quick thing, etc.
Turns out he was laid off and my suggestion was used.
(Okay, I'm being silly, the layoff was a coincidence)
On the other hand, "hey I have 6TiB data, please prepare to analyze it, feel free to ask any questions for clarification but I may not know the answers" is much more representative of a real-life task.
Or, if you can guess what the interviewer is aiming for, state the assumption and go from there "If we assume it's gonna stay at <10TB for the next couple of years or even longer, then..."
Then the interviewer can interrupt and change the assumptions to his needs.
Even if it would say "it will stay at 6TiB" I would probably prefer a senior candidate to briefly question it, such as "It is surprising that we know it will stay at 6TiB and if this were a real project I'd try to at least sanitycheck that this requirement is really correct, but for now I'll assume this is a given..."
At least if, as the interviewer, I told them to treat the challenge similar to a real request/project and not to not-question the given numbers etc.
How much do they pay? How long has it been since my last proper meal? How long until my rent is due?
I had an interview "design an app store". I tried asking, ok an app store has a ton of components, which part of the app store are you asking exactly? The response I got was "Have you ever used an app store? Design an app store". Umm ok.
But now as a senior, I have the same questions and answers regardless if I’m being interviewed or not.
DING DING DING!
About 30% of all interviews I've had in my career the person doing the technical interview was using it as a means to stroke their ego. I got the impression they don't get the power they want there and this was a brief reprieve where they are "the expert" and get to thumbs up or thumbs down in their kingly intellectual facade.
TBF, I personally cut them some slack because they may actually be right about a lot and not given the authority to do what they know is right.
Still, if we're subjected to AI resume filtering, then how about we replace the technical interview process with AI too and eliminate the BS ego trips?
Tell me this isn't big data, and then, if you must, tell me about hadoop (or whatever big data is)
.parquet preserves data types (unlike CSV)
They are 10x smaller than CSV. So 600GB instead of 6TB.
They are 50x faster to read than CSV
They are an "open standard" from Apache Foundation
Of course, you can't peek inside them as easily as you can a CSV. But, the tradeoffs are worth it!
Please promote the use of .parquet files! Make .parquet files available for download everywhere .csv is available!
But it is because it makes a lot of things much simpler, and that a lot of people have not realized that. Tooling is moving fast in this space, it is not 2004 anymore.
His arguments are still valid and 86 days is a pretty long time.
3 times across 3 months is hardly astroturfing for big parquet territory.
So, of you want to query a giant collection of protobufs, you end up reading and deserializing every record. For parquet, you get much closer to only reading what you need.
Dremel was pretty revolutionary when it came out in 2006 - you could run ad-hoc analyses in seconds that previously would've taken a couple days of coding & execution time. Parquet is awesome for the same reasons.
I believe that Parquet files have rather monolithic metadata at the end and it has 4G max size limit. 600 columns (it is realistic, believe me), and we are at slightly less than 7.2 millions row groups. Give each row group 8K rows and we are limited to 60 billion rows total. It is not much.
The flatness of the file metadata require external data structures to handle it more or less well. You cannot just mmap it and be good. This external data structure most probably will take as much memory as file metadata, or even more. So, 4G+ of your RAM will be, well, used slightly inefficiently.
(block-run-mapped log structured merge tree in one file can be as compact as parquet file and allow for very efficient memory mapped operations without additional data structures)
Thus, while parqet is a step, I am not sure it is a step in definitely right direction. Some aspects of it are good, some are not that good.
But nobody tells me that I can hit a hard limit and then I need a second Parquet file and should have some code for that.
The situation looks to me as if my "Favorite DB server" supports, say, only 1.9 billions records per table and if I hit that limit I need a second instance of my "Favorite DB server" just for that unfortunate table. And it is not documented anywhere.
"Dictionary Encoding is effective across data types (even for floating-point values) because most real-world data have low NDV ratios. Future formats should continue to apply the technique aggressively, as in Parquet."
So this is not critique, but assessment. And Parquet has some interesting design decisions I did not know about.
So, let me thank you again. ;)
A former colleague of mine is now working on a memory-mapped log-structured merge tree implementation and it can be a good alternative. LSM provides elasticity, one can store as much data as one needs, it is static, thus it can be compressed as well as Parquet-stored data, memory mapping and implicit indexing of data do not require additional data structures.
Something like LevelDB and/or RocksDB can provide most of that, especially when used in covering index [1] mode.
[1] https://www.sqlite.org/queryplanner.html#_covering_indexes
Most tools can run queries across parquet files.
Like everything, it has its strengths and weaknesses, but in most cases, it has better trade-offs over CSV if you have more than a few thousand rows.
This is not emphasized often enough. Parquet is useless for anything that requires writing back computed results as in data used by signal processing applications.
Why would you need 7.2 mil row groups?
Row group size when stored in HDFS is usually equal to HDFS bock size by default, which is 128MB
7.2 mil * 128MB ~ 1PB
You have a single parquet file 1PB in size?
You can have compressed Parquet columns with 8192 entries being a couple of tens bytes in size. 600 columns in a row group is then 12K bytes or so, leading us to 100GB file, not a petabyte. Four orders of magnitude of difference between your assessment and mine.
I actually benchmarked this and duckdb CSV reader is faster than parquet reader.
CSV underperforms in almost every other domain, like joins, aggregations, filters. Parquet lets you do that lazily without reading the entire Parquet dataset into memory.
Yes, I think duckdb only reads CSV, then projects necessary data into internal format (which is probably more efficient than parquet, again based on my benchmarks), and does all ops (joins, aggregations) on that format.
With Parquet you almost never read in the entire dataset and it's fast on all the projections, joins, etc. while living on disk.
what? Why CSV is required to fit in memory in this case? I tested CSVs which are far larger than memory, and it works just fine.
The parquet file has metadata that allows duckdb to only read the parts that are actually used, reducing total amount of data read from disk/network.
this makes sense, and what I hoped to have. But in reality looks like parsing CSV string works faster than bloated and overengineered parquet format with libs.
Anecdotally having worked with large CSVs and large on-disk Parquet datasets, my experience is the opposite of yours. My DuckDB queries operate directly on Parquet on disk and never load the entire dataset, and is always much faster than the equivalent operation on CSV files.
I think your experience might be due to -- what it sounds like -- parsing the entire CSV into memory first (CREATE TABLE) and then processing after. That is not an apples-to-apples comparison because we usually don't do this with Parquet -- there's no CREATE TABLE step. At most there's a CREATE VIEW, which is lazy.
I've seen your comments bashing Parquet in DuckDB multiple times, and I think you might be doing something wrong.
original discussion was about CSV vs parquet "reader" part, so this is exactly apple to apple testing, easy to benchmark and I stand my ground. What you are doing downstream, it is another question which is not possible to discuss because no code for your logic is available.
> I've seen your comments bashing Parquet in DuckDB multiple times, and I think you might be doing something wrong.
like running one command from DuckDB doc.
Also, I am not "bashing", I just state that CSV reader is faster.
apt-cache search parquet
<nada>
Maybe later apt-cache search csv | wc -l
259 $ apt-cache search csv | wc -l
225
$ apt-cache search parquet | wc -l
0how? lossless compression? under what scenario?
vague headlines like this just beg more questions
Take Apache Iceberg for example. It is essentially a specification to how to store parquet files for efficient use and exploration of data, but the only implementation... depends on Apache Spark!
Deep centralization at the expense of simplicity and true redundancy; like renting a laser cutter when you need a boxcutter, a pair of scissors, and the occasional toenail clipper.
The OPs desired solution could have been found from probably some of those other candidates if asked “here is the challenge, solve in most McGuyver way possible”. Because if you change the second part, the correct answer changes.
“Here is a challenge, solve in the most accurate, verifiable way possible”
“Here is a challenge, solve in a way that enables collaboration”
“Here is a challenge, 6TiB but always changing”
^ These are data science questions much more than the question he was asking. The answer in this case is that you’re not actually looking for a data scientist.
One can't go wrong with DuckDB+SQLite+Open/Elasticsearch either with 6 to 8 even 10 TB of data.
[0]. https://duckdb.org/
My question is always "How much data do you actually have?" Many times you they reply with 500GB or 2TB. I tell that that isn't much data when you can get 1TB micro SD card the size of a fingernail or a 24TB hard drive.
My feeling is that if you really need to store petabytes of data that you aren't going to ask how to do it on reddit. If you need to store petabytes you will have an IT team and substantial budget and vendors that can figure it out.
I think it’s part of the BS cycle that’s prevalent in companies. You can’t admit that you are doing something simple.
- the developers want to get experience in a fancy stack to build up their resume
Everyone benefits from the collective hallucination
Data organisation can be more of an issue, but the general issue with this is often a lack of internal discipline on the data owners to carefully manage their data. But that isn't nearly as attractive for management. Putting 'brought in new cloud vendor / technology' looks better than 'improved data organisation' on a CV even if the new vendor was a waste of money.
It’s like a reverse signal.
If it's not a very write heavy workload but you'd still want to be able to look things up, wouldn't something like SQLite be a good choice, up to 281 TB: https://www.sqlite.org/limits.html
It even has basic JSON support, if you're up against some freeform JSON and not all of your data neatly fits into a schema: https://sqlite.org/json1.html
A step up from that would be PostgreSQL running in a container: giving you the support for all sorts of workloads, more advanced extensions for pretty much anything you might ever want to do, from geospatial data with PostGIS, to something like pgvector, timescaledb etc., while still having a plethora of drivers and still not making your drown in complexity and having no issues with a few dozen/hundred TB of data.
Either of those would be something that most people on the market know, neither will make anyone want to pull their hair out and they'll give you the benefit of both quick data writes/retrieval, as well as querying. Not that everything needs or can even work with a relational database, but it's still an okay tool to reach for past trivial file storage needs. Plus, you have to build a bit less of whatever functionality you might need around the data you store, in addition to there even being nice options for transparent compression.
Obviously we can get boxes with multiple terabytes of RAM for $50-200/hr on-demand but nobody is doing that and then also using awk. They’re loading the data into clickhouse or duckdb (at which point the ram requirement is probably 64-128GB)
I feel like this is an anecdotal story that has mixed up sizes and tools for dramatic effect.
Ramdisk would work.
Awk is imo a poor solution. I use awk all the time and I would never use it for something like this. Why not just use postgres. Its a lot more powerful, easy to setup and you get SQL which is extremely powerful. Normally I might even go with sqllite but for me 6TB is too much for sqllite.
For example, just logging stuff into a large text file is so much easier, performant and searchable that using AWS CloudWatch, presumably written by some of the smartest programmers who ever lived.
On another note I was once asked to create a big data-ish object DB, and me, knowing nothing about the domain, and a bit of benchmarking, decided to just use zstd-compressed json streams with a separate index in an sql table. I'm sure any professional would recoil at it in horror, but it could do literally gigabytes/sec retrieval or deserialization on consumer grade hardware.
That said, most open source tools have terrible performance and efficiency on large, fast hardware. This contributes to the intuition that you need to throw hardware at the problem even for relatively small problems.
In 2024, "big data" doesn't really start until you are in the petabyte range.
What do you use?
I wrote a comment on here the other day that some place I was trying to do work for was using $11k USD a month on a BigQuery DB that had 375MB of source data. My advice was basically you need to hire a data scientist that knows what they are doing. They were not interested and would rather just band-aid the situation for a "cheap" employee. Despite the fact their GCP bill could pay for a skilled employee.
As I've seen it for the last year job hunting most places don't want good people. They want replaceable people.
is not somewhat detailed requirements, as it depends quite a bit on the nature of the data.
You didn't, and indeed you have a point (missing specification of expected queries), so I expand it as a response here.
Among the MANY requirements I shared with the candidate, only one was the 6TiB. Another one was that it was going to be serving as part of the backend of an internal banking knowledge base, with at maximum 100 request a day (definitely not 10k people using it).
To all the upset data infrastructure wizards here: calm down. It was a banking startup, with an experimental project, and we needed the sober thinker generalist, who can deliver solutions to real *small scale* problems, and not the one who was the winner on the buzzword bingo.
HTH.
Processing and querying it is trickier.
Why?
That's the boring solution. If you don't have a use case, what kind of queries you would run then opt for maximum flexibility with the minimum setup of a managed solution.
If cost is prohibitive on the long run, you can figure out a more tailored solution based on the revealed preferences.
Fiddling with CSVs is the DWH version of the legendary "Dropbox HN commenter".
Whereas the "right answer" (I had a man on the inside) was to describe some wild tall and wide event based distributed system. For some nominal request volume that was nowhere near the limits of postgres. And they didn't even care if you solved the actual hard distributed system problems that would arise like distributed transactions etc.
Anyway, I said I failed the interview, really they failed my filter because if they want me to ignore pragmaticism and blindly regurgitate a YouTube video on "system design" FAANG interview prep, then I don't want to work there anyway.
That's quite a bit..
External or internal? Any examples?
"... it could be loaded (multimple times) to memory"
All 6TiB at once, or loaded in chunks?
I gave him the "find the phone numbers in 50,000 html files" question, and he decided to write a huge program with an ad-hoc state machine. When I asked how long it would take to write the program, he said he'd have to hit N files, with M lines per file, so... I interrupted him: no, WRITE. How long to WRITE the program? Oh. He estimated it at 5 days of work. At this point I was 50% ready to throw him out.
1) Mongo is a bad point of reference.
The one lesson I've learned is that there is nothing Mongo does which postgresql doesn't do better. Big data solutions aren't nosql / mongo, but usually things like columnar databases, map/reduce, Cassandra, etc.
2) Plan for success
95% of businesses never become unicorns, but that's the goal for most (for the 5% which do). If you don't plan for it, you won't make it. The reason to architect for scalability when you have 5 customers is so if that exponential growth cycle hits, you can capitalize on it.
That's not just architecture. To have any chance of becoming a unicorn, every part of the business needs to be planned for now and for later: How do we make this practical / sustainable today? How do we make sure it can grow later when we have millions of customers? A lot of this can be left as scaffolding (we'll swap in [X], but for now, we'll do [Y]).
But the key lessons are correct:
- Most data isn't big. I can fit data about every person in the world on a $100 Chromebook. (8 billion people * 8 bits of data = 8GB)
- Most data is rarely queried, and most queries are tiny. The first step in most big data jobs I've done is taking terabytes of data and shrinking it down to the GB, MB, or oven KB-scale data I need. Caveat: I have no algorithm for predicting what I'll need in the future.
- Cost of data is increasing with regulatory.
Is it really the general case or is it just a HN echo chamber meme?
My pet peeve is that patterns used by companies that in theory could become global unicorns are mimicked by companies where 5000 paying customers would mean an immense success
Lifestyle companies are fine, if that's what you're aiming for. I know plenty of people who run or work at ≈1-30 person companies with no intention to grow.
However, if you're going for high-growth, you need to plan for success. I've seen many potential unicorns stopped by simple lack of planning early on. Despite all the pivots which happen, if you haven't outlined a clear path from 1-3 people in a metaphorical garage to reaching $1B, it almost never happens, and sometimes for stupid reasons.
If your goal is 5000 paying customers at $100 per year and $500k in annual revenues, that can lead to a very decent life. However, it's an entire different ballgame: (1) Don't take in investment (2) You probably can't hire more than one person (3) You need a plan for break-even revenue before you need to quit your job / run out of savings. (4) You need much greater than the 1-in-10 odds of success.
And it's very possible (and probably not even hard) to start a sustainable 1-5 person business with >>50% odds of success, especially late career:
- Find a niche you're aware of from your job
- Do ballpark numbers on revenues. These should land in the $500k-$10M range. Less, and you won't sustain. More, and there will be too much competition.
- Do it better than the (likely incompetent or non-existent) people doing it now
- Use your network of industry contacts to sell it
That's not a big enough market you need to worry about a lot of competition, competitors with VC funding, etc. Especially ones with tall moats do well -- pick some unique skillset, technology, or market access, for example.
However, IF you've e.g. taken in VC funding, then you do need to plan for growth, and part of that is planning for the small odds your customer base (and ergo, your data) does grow.
most of the business plans are clear. acquire users and/or paying customers, build out the product, raise money and eventually raise prices too.
engineering side is even simpler, keep the lights on, work with product designers to find incremental steps, iterate.
as others have mentioned upthread if a startup finds a great niche (and manages to make it a whole new market for example) they simple cannot fail due to inadequate engineering. (because worst case you fire the whole engineering department and just acqui-hire a random startup and you'll be back to growth in no time.)
of course this take might be too hot, so I'm happy to read some counterarguments (maybe there are even counterexamples?)
> ...
> of course this take might be too hot, so I'm happy to read some counterarguments (maybe there are even counterexamples?)
Netscape.
(Plenty of others too)
And even when it failed it was worth 4.2B to AOL.
... if we're talking about browsers, Netscape's reincarnation Mozilla is again in the same tough spot. They have zero business sense, absolutely no plan, nada. At that point it's completely irrelevant how long it takes them to fix bugs or implement Web APIs or make it faster on various platforms. And still, they funded Rust and Servo.
Does a mining company want to become a "unicorn"?
A fish and chip shop?
Even within tech there is an extremely large number of companies whose goals are to steadily increase profits and return them to shareholders. 37 Signals is the posterchild there.
Maybe if you're a VC funded startup then yeah.
Obsessed with this "you must use PostgreSQL for every use case" nonsense.
And that anyone who actually has unique data needs is simply doing it for their resume or are over-engineering.
Pg fans are certainly here asking "why not PG?". Yet so are fans of other DBs; like DuckDB, CouchDB, SQLite, etc.
Postgres is the next obvious stepping stone after that, and the one where the vast majority of actual real-world cases that are not hypotheticals end up fitting.
Postgres is recommended because it's free, easy to maintain, install and it's has excellent query features.
> who actually has unique data needs
We are saying this is probably not true, and you just want to play with toys rather than ship working systems.
Google search, cannot be built in Postgres.
Startups are, however, atypical from new businesses, ergo the unicorn myth, meaning we see many attempts to follow such a path that likely stands in the way of many new businesses from actually achieving the more real goals of, well, being a business, succeeding in their venture to produce whatever it is and reach their customers.
I describe it as a unicorn "myth" as it very much behaves in such a way, and is misinterpreted similarly to many myths we tell ourselves. Unicorns are rare and successful because they had the right mixture of novel business and the security of investment or buyouts. Startups purportedly are about new ways of doing business, however the reality is only a handful really explore such (e.g. if it's SaaS, it's probably not a startup), meaning the others are just regular businesses with known paths ahead (including, of course, following in the footsteps of prior startups, which really is self-refuting).
With that in mind, many of the "real" unicorns are realistically just highly valued new businesses (that got lucky and had fallbacks), as they are often not actually developing new approaches to business, whereas the mythical unicorns that startups want to be are half-baked ideas of how they'll achieve that valuation and wealth without much idea of how they do business (or that it can be fluid, matching their nebulous conception of it), just that "it'll come", especially with "growth".
There is no nominative determinism, and all that, so businesses may call themselves startups all they like, but if they follow the patterns of startups without the massive safety nets of support and circumstance many of the real unicorns had, then a failure to develop out the business proper means they do indeed suffer themselves by not appreciating 5000 paying customers and instead aim for "world domination", as it were, or acquisition (which they typically don't "survive" from, as an actual business venture). The studies have shown this really does contribute to the failure rate and instability of so-called startups, effectively due to not cutting it as businesses, far above the expected norm of new businesses...
So that pet peeve really is indicative of a much more profound issue that, indeed, seems to be a bit of an echo chamber blind spot with HN.
After all, if it ought to have worked all the time, reality would look very different from today. Just saying how many don't become unicorns (let alone the failure rate) doesn't address the dissonance from then concluding "but this time will be different". It also doesn't address the idea that you don't need to become a "unicorn", and maybe shouldn't want to either... but that's a line of thinking counter to the echo chamber, so I won't belabour it here.
Nitpick but I cannot help myself: 8 bits are not even enough for a unique integer ID per person, that would require 8 bytes per person and then we are at 60GB already.
I agree with pretty much anything else you said, just this stood out as wrong and Duty Calls.
I wonder how they dealt with common storage issues like backups and disks having bad sectors.
Writing your mapping function would be tricky! But definitely theoretically possible.
We had spell checkers before computers had enough memory to fit all words. They'd probabilistically find almost all incorrect words (but not suggest corrections). It worked fine.
I think you're missing quite a few 9s!
That's exactly what every architecture astronaut everywhere says. In my experience it's completely untrue, and actually "planning for success" more often than not causes huge drags on productivity, and even more important for startups, on agility. Because people never just make plans, they usually implement too.
Plan for the next 3 months and you'll be much more agile and productive. Your startup will never become a unicorn if you can't execute.
I've come to the conclusion that the only strategy that works reliably is to build something that solves problems you have NOW rather than trying to predict the future.
Laying out adjacent markets, potential pivots, likely product features, etc. is a weekend-long exercise. That can help define both where the architecture needs to be flexible, and just as importantly, *where it does not*.
Over-engineering happens when you plan / architect for things which are unlikely to happen.
The suggestion upthread to use awk is awesome if you’re a bunch of Linux grey beards.
But if you have access to people with particular skills or domain knowledge… spending extra cash on silly infrastructure is (within reason) way cheaper than having that employee be less productive.
That's not my experience at all.
Architecture != implementation
Architecture astronauts will try to solve the world's problems in v0. That's very different from having an architectural vision and building a subset of it to solve problems for the next 3 months. Let me illustrate:
* Agile Idiot: We'll stick it all in PostgreSQL, however it fits, and meet our 3-month milestone. [Everything crashes-and-burns on success]
* Architecture Astronaut: We'll stick it all in a high-performance KVS [Business goes under before v0 is shipped]
* Success: We have one table which will grow to petabytes if we reach scale. We'll stick it all in postgresql for now, but maintain a clean KVS abstraction for that one table. If we hit success, we'll migrate to [insert high-performance KVS]. All the other stuff will stay in postgresql.
The trick is to have a pathway to success while meeting short-term milestones. That's not just software architecture. That's business strategy (clean beachhead, large ultimate market), and every other piece of designing a successful startup. There should be a detailed 3-month plan, a long-term vision, and a rough set of connecting steps.
Scaling up will bring its own challenges, with many of them difficult to foresee.
I think that in practice that’s counterproductive. A startup has a limited runway. If your engineers are spending your money on something that doesn’t pay off for years then they’re increasing the chance you’ll fail before it matters.
Planning is a weekend, or at most a few weeks.
a) It has a built-in and supported horizontal scalability / HA solution.
b) For some use cases e.g. star schemas it has significantly better performance.
> Big data solutions aren't nosql
Almost all big data storage solutions are NoSQL.
I think it's important to distinguish between OLAP AND OLTP.
For OLAP use cases (which is what this post is mostly about) it's almost 100% SQL. The biggest players being Databricks, Snowflake and BigQuery. Other tools may include AWS's tools (Glue, Athena), Trino, ClickHouse, etc.
I bet there's a <1% market for "NoSQL" tools such as MongoDB's "Atlas Data Lake" and probably a bunch of MapReduce jobs still being used in production, but these are the exception, not the rule.
For OLTP "big data", I'm assuming we're talking about "scale-out" distributed databases which are either SQL (e.g. cockroachdb, vitess, etc) SQL-like (Casandra's CQL, Elasticsearch's non-ANSI SQL, Influx' InfluxQL) or a purpose-built language/API (Redis, MongoDB).
I wouldn't say OLTP is "almost all" NoSQL, but definitely a larger proportion compared to OLAP.
Most I've seen aren't. NoSQL means non-relational database. Most big data solutions I've seen will not use a database at all. An example is hadoop.
Once you have a database, SQL makes a lot of sense. There are big data SQL solutions, mostly in the form of columnar read-optimized databases.
On the above, a little bit of relational can make a huge performance difference, in the form of, for example, a big table with compact data with indexes into small data tables. That can be algorithmically a lot more performant than the same thing without relations.
Sources ?
It's a second system syndrome + survivor bias thing I think: people who had to clean up the mess of a good MVP complaining about what wasn't done before. But the companies that DID do that planning and architecting before did not survive to be complained about.
Layers of abstraction make code harder to reason about and work with, so it's a lose lose when trying to iterate quickly, but there's also the idea of architectural "mise en place" vs "just dump shit where it's most convenient right now and don't worry about later" which will result near immediate productivity losses due to system incoherence and disorganization.
It's a bit annoying how the design of Django templates works against this by not allowing free functions...
If you have a product gaining that much traction, it’s usually because of some compound effect based on the existence and needs of its userbase. If on the way up you stumble to add new users, the userbase that’s already there is unlikely to go back to the Old Thing or go somewhere else (because these events are actually rare). For a good while using Twitter meant seeing the fail whale every day. Most people didn’t just up and leave, and nothing else popped up that could scale better that people moved to. Making a product that experiences exponential growth in that way is pretty rare, and struggling to scale those cases and having a period of availability degradation is common. What products hit an exponential growth situation failed because they couldn’t scale?
I think that was exactly their point. If new architectures were actually necessary, we would have seen a greater rise in Mongo and the like. But we didn't, because the existing systems were perfectly adequate.
I suppose there are only 256 unique types of people on the planet? :)
"If you can't do your statistical analysis in 1 to 5 TB of data, your methodology is flawed"
This is probably more about human limitations than math. There's a clear ceiling in how much flexibility we can use. That will also change with easier ways to run new kinds of analysis, but it increases with the logarithm of the amount of things we want to do.
It sounds interesting, but it's totally new to me.
You can do (lossy) compression on rows of vectors (treated like a matrix) by taking the top N eigenvectors (largest N eigenvalues) and using them to approximate the original matrix with increasing accuracy (as N grows) by some simple linear operations. If the numbers are highly correlated, you can get a huge amount of compression with minor loss this way.
Personally I like to use it to visualize linear separability of a high dimensioned set of vectors by taking a 2-component PCA and plotting them as x/y values.
Like: "Look, boss, I can compute all those averages for that report on just my laptop, by ingesting a SAMPLE of the data, rather than making those computations across the WHOLE dataset".
Boss: "What do you mean 'sample'? I just don't know what you're trying to imply with your mathmo engineeringy gobbledigook! Me having spent those millions on nothing can clearly not be it, right?"
The amount of salesman hype and chatter about big data, followed by the dick measuring contests about whose data was big enough to be worthy was intense for awhile.
It was extremely difficult to get > 64gb on a machine for a very long time, and implementation complexity gets hard FAST when you have a hard cap.
And it's EXTREMELY disruptive to have a process that fails every 1/50 times, when data is slightly too large, because your team will be juggling dozens of these routine crons, and if each of them breaks regularly, you do nothing but dumb oncall trying to trim bits off of each strong.
No, Hadoop and MapReduce were not hyperefficient, but it was OK if you write it correctly, and having something that ran reliably is WAY more valuable than boutique bit-optimized C++ crap that nobody trusts or can maintain and fails every thursday with insane segfaults.
(nowdays, just use Snowflake. but it was a reasonable tool for the time).
And I don't understand why you're reading "boutique bit-optimized C++ crap" into "most basic and obvious optimizations".
One of those most basic and obvious optimizations is to avoid reading a dataframe into memory in its entirety, when the math that you want to do on top of it can actually be done as a running accumulator while reading the data from a stream. This is possible in 90% of realistic use cases, but the fraction of software written back then that took advantage of this was shockingly small. Solving the problem by buying more machines, chopping up the dataframe into smaller pieces, and farming out the payload through Hadoop had management buy-in. Yet, for some reason, doing the sane thing, namely rewriting poorly-written software, didn't.
Originally big data was defined by 3 dimensions:
- Volume (mostly what the author talks about) [solved]
- Velocity, how fast data is processed etc [solved, but expensive]
- Variety [not solved]
Big Data today is not: I don't have enough storage or compute.
It is: I don't have enough cognitive capacity to integrate and make sense of it.
So AI currently, probably ;)
One of the reasons big data systems took off was because enterprises had exports out of third party systems that they didn't want to model since they didn't own it. As well as a bunch of unstructured data e.g. floor plans, images, logs, telemetry etc.
It means that data comes in a ton of different shapes with poorly described schemas (technically and semantically).
From the typical CSV export out of an ERP system to a proprietary message format from your own custom embedded device software.
Highly recommend this and related talks by him, most of them are in YouTube.
[1] https://www.youtube.com/watch?v=KRcecxdGxvQ
[2] https://amturing.acm.org/award_winners/stonebraker_1172121.c...
It is for me. Six times per year I go out to the field for two weeks to do data acquisition. In the field we do a dual-aircraft synthetic aperture radar collection over four bands and dual polarities.
That means two aircraft each with one radar system containing eight 20TiB 16-drive RAID-0 SSD storage devices.
We don't usually fill up the RAIDs so we generate about 176TiB of data per day and over the two weeks we do 7 flight, or 1.2PiB per deployment or 7.2PiB per year.
We can only fly every other day because it takes a day between flights to offload the data via fiber onto storage servers that are usually haphazardly crammed into the corner of a hangar next to the apron. It is then duplicated to a second server for safekeeping and at the end of the mission everything is shipped back to our HQ for storage and processing.
The data is valuable, but not "billions" valuable. It is used for resource extraction, mapping, environmental and geodetic research, and other applications (but that's not my department) so we have kept every single byte for since 2008. This is especially useful because as new algorithms are created (not my department) the old data can be reprocessed to the new standard.
Entire nations finally know how many islands they have, how large they are, how their elevations are changing, and how their coasts are being eradicated by sea level change because of our data and if you've ever used a mapping application and flown around a city with 3d buildings that don't look like shit because they were stitched together using AI and photogrammetry, you've used our data too.
We have to use hard drives because SSDs would be space and most certainly cost prohibitive.
We stream 800GiB-2TiB files each representing a complete stripe or circular orbit to GPU-equipped processing servers. Files are incompressible (the cosmic microwave background, the bulk of what we capture, tends to be a little random) and when I started I held on to the delusion that I could halve the infrastructure by writing to tape until I found out that tape capacities were calculated for the storage of gigabyte-sized text files of all zeros (or so it seems) that can be compressed down to nothing.
GPUs are too slow. CPUs are too slow. PCIe busses are too slow. RAM is too slow. My typing speed is too slow. Everything needs to be faster all of the time.
Everything is too slow, too hard, and too small. Hard drives are too small. Tuning the linux kernel and setting up fast and reliable networking to the processing clusters is too hard. Kernel and package updates that aren't even bug fixes but just changes in the way that something works internally that are transparent to all users except for us break things. Networks are too slow. Things exist in this fantasy world where RAM is scarce so out-of-the-box settings are configured to not hog memory for network operations. No. I've got a half a terabyte of RAM in this file server use ALL OF IT to make the network and filesystem go faster, please. Time to spend six hours reading the documentation for every portion of the network stack to increase the I/O to 2024-levels of sanity.
I probably know more about sysctl.conf than almost every other human being on earth.
Distributed persistent object storage systems for people who think they are doing big data but really aren't either completely fall apart under our workload or cost hundreds of millions of dollars-- which we don't have. When I tell all of the distributed filesystem salespeople that our objects are roughly a terabyte in size they stop replying to my emails. More than one vendor has referred me to their intelligence community customer service representative upon reading my requirements. I am not the NSA, buddy, and we don't have NSA money.
Every once in a while we get a new MBA or PMP who read a Bloomberg article about the cloud and asks about moving to AWS or Azure after they see the costs of our on-premises datacenter. When I show them the numbers, in terms of both money and time, they throw up in their mouths and change the subject.
To top it all off all of our vendors are jumping on the AI/cloud bandwagon and discontinuing product lines applicable to us.
And now I've got to compete for GPUs with hedge funds and AI startups trying to figure out how to use a LLM to harvest customer data and use it to show them ads.
I do not have enough storage or compute, and the storage and compute I do have is too slow.
DPUs/IPUs look interesting but fall on their face when an object is larger than a SQL database query or compressed streaming video chunk.
What’s preventing you from buying a second set of SSD drives and swapping them into planes end of day?
The planes would then be able to fly every day while the previous day’s data is being offloaded on the ground, condensing the two weeks into one. Swaps could even be quite easy if you replace full enclosures rather than each drive individually.
I'm probably missing something, though!
The arrays are custom parts made to interface directly with the radar and are VPX modules. VPX edge connectors are not rated for very many matings, believe me we've tried.
Every-other-day flights also keep the pilots comfortably within their sleep limits.
My boss was smart enough to have stakeholder meetings where they regularly discussed what to keep and what to throw away, and with some smart algorithms we were able to compress all that data down into like 200MB per day.
We loaded the last 2 months into an sql server and the last 2 years further aggregated into another, and the whole company used the data in excel to do queries on it in reasonable time.
The big data is rotting away on tape storage in case they ever need it in the future.
My boss got a lot of stuff right and I learned a lot, though I only realized that in hindsight. Dude was a bad manager but he knew his data.
In my experience, doing this will yield the query results faster than some other systems even starting the query execution (yes, I'm looking at you Athena)...
I even think that a lot of queries can be run from a browser nowadays, that's why I created https://sql-workbench.com/ with the help of DuckDB WASM (https://github.com/duckdb/duckdb-wasm) and perspective.js (https://github.com/finos/perspective).
AI also use all the data, just with a magick neural network to figure out what it all means.
- The "hallucination" factor means every result an AI tells you about big data is suspect. I'm sure some of you who really understand AI more than the average person can "um akshually" me on this and tell me how it's possible to configure ChatGPT to absolutely be honest 100% of the time but given the current state of what I've seen from general-purpose AI tools, I just can't trust it. In many ways this is worse than MongoDB just dropping data since at least Mongo won't make up conclusions about data that's not there.
- At the end of the day - and I think we're going to be seeing this happen a lot in the future with other workflows as well - you're using this heavy, low-performance general-purpose tool to solve a problem which can be solved much more performatively by using tools which have been designed from the beginning to handle data management and analysis. The reason traditional SQL RDBMSes have endured and aren't going anywhere soon is partially because they've proven to be a very good compromise between general functionality and performance for the task of managing various types of data. AI is nowhere near as good of a balance for this task in almost all cases.
All that being said, the same way Electron has proven to be a popular tool for writing widely-used desktop and mobile applications, performance and UI concerns be damned all the way to hell, I'm sure we'll be seeing AI-powered "big data" analysis tools very soon if they're not out there already, and they will suck but people will use them anyway to everyone's detriment.
> I used to joke that Data Scientists exist not to uncover insights or provide analysis, but merely to provide factoids that confirm senior management's prior beliefs.
I think AI is used for the same purpose in companies: signal to the world that the company is using the latest tech and internally for supporting existing political beliefs.
So same job. Hallucination is not a problem here as the AI conclusions are not used when they don't align to existing political beliefs.
AI / ML means more than just LLM chat output, even if that's the current hype cycle of the last couple of years. ML can be used to build a perfectly serviceable classifier, or predictor, or outlier detector.
It suffers from the lack of explainability that's always plagued AI / ML, especially as you start looking at deeper neural networks where you're more and more heavily reliant on their ability to approximate arbitrary functions as you add more layers.
> you're using this heavy, low-performance general-purpose tool to solve a problem which can be solved much more performatively by using tools which have been designed from the beginning to handle data management and analysis
You are not wrong here, but one challenge is that sometimes even your domain experts do not know how to solve the problem, and applying traditional statistical methods without understanding the space is a great way of identifying spurious correlations. (To be fair, this applies in equal measure to ML methods.)
It started with Hadoop which was inspired by what existed at Google and became popular in enterprises all around the world who wanted a cheaper/better way to deal with their data than Oracle.
Spark came about as a solution to the complexity of Hive/Pig etc. And then once companies were able to build reliable data pipelines we started to see AI being able to be layered on top.
Data models generated by intentional human action e.g. clicking a link, sending a message, buying something, etc are universally small. There is a limit on the number of humans and the number of intentional events they can generate per second regardless of data model.
Data models generated by machines, on the other hand, can be several orders of magnitude higher velocity and higher volume, and the data model size is unbounded. These are often some of the most interesting and under-utilized data models that exist because they can get at many facts about the world that are not obtainable from the intentional human data models.
Everyone was talking about data being the new "oil".
Enterprise could basically deploy petabyte scale warehouse run HBase or Hive on top of it and build makeshift data-warehouses. It was also when the cloud was emerging, people started creating EMR clusters and deploy workloads there.
I think it was solution looking for problem. And the problem existed only for a handful of companies.
I think somehow, how cloud providers abstracted lot of these tools and databases, gave a better service and we kind of forgot about hadoop et al.
As to the topic, IMO, there is a contradiction. The only way to handle big data is to divide it into chunks that aren’t expensive to query. In that sense, no data is "big" as long as it’s handled properly.
Also, about big data being only a problem for 1 percent of companies: it's a ridiculous argument implying that big data was supposed to be a problem for everyone.
I personally don’t see the point behind the article, with all due respect to the author.
I also see many awk experts here who have never been in charge of building enterprise data pipelines.
What the article talks about is more like a particular type of database architecture (often collectively called "NoSQL") that was a fad a few years ago, and as all fads, it went down. It doesn't mean having lots of data is useless, or that NoSQL is useless, just that it is not the solution to every problem. And also that there is a reason why regular SQL databases have been in use since the 70s: except in specific situation most people don't encounter, they just work.
<cough> Electron?
The most interesting applications for "big data" are all (IMO) in the scientific computing space. Yeah, your e-commerce business probably won't ever need "big data" but load up a couple genomics research sets and you sure will.
We were having scaling issues with one particular customer, so I looked into the various solutions. They all looked pretty cool, but I didn't like that the developer communities were small compared to MySQL and PostgreSQL.
So I ended up sticking with MySQL and building a simple sharding mechanism that allowed us to move big customers to their own server if necessary. It worked great.
While I was working on that, MixPanel published a long, highly detailed blog post about why they had moved to MongoDB. That led me to have second thoughts. "If these guys are doing it, maybe we should too?" But I was pretty far down the sharding development path and couldn't justify making such a wholesale change based on one blog post.
Fast forward about nine months, and MixPanel published a new blog post about why they had moved off MongoDB. Vindication!
Ever since then I've had conservative disposition when it comes to deploying new technologies.
However what ML systems and in particular LLMs rely on having access to millions (if not billions or trillions) of examples. The underlying infra of which is based on some of these tools.
Big Data isn't dead, just this weird idea that the tools and usecases around querying databases has been finally recognised as being mostly useless to most people. It is and always has been about training ML models.
[2]https://motherduck.com/_next/image/?url=https%3A%2F%2Fweb-as...
This isn't a value judgement about whether that's a good idea, just an observation from talking with many tech companies doing ML. This definitely feels like a bubble that will burst in due time, but for now ML is turbocharging Big Data.
There are some larger data-sizes, and query patterns, where either BigQuery Capacity compute pricing, or another vendor like Snowflake, becomes more economical.
https://cloud.google.com/bigquery/pricing
BigQuery offers a choice of two compute pricing models for running queries:
On-demand pricing (per TiB). With this pricing model, you are charged for the number of bytes processed by each query. The first 1 TiB of query data processed per month is free.
Queries (on-demand) - $6.25 per TiB - The first 1 TiB per month is free.Spark in general was a great tool to get out of the Big Data era.
With GDPR, we went from keeping everything by default unless a customer explicitly requested it gone to deleting it all automatically after a certain number of days after their license expires. This makes opaque data lakes completely untenable.
Don't get me wrong, this is all a net positive. The customers data is physically removed and they don't have to worry about future leaks or malicious uses, and we get a more efficient database. The only people really fussed were the sales team trying to lure people back with promises that they could pick right back up where they left off.
For lack of data, they generate random bytes collected on every mouse movement on every page, and every packet that moves through their network. It doesn’t tell them anything because the only information that means anything is who clicks that one button on their checkout page after filling out the form with their information or that one request that breaches their system.
That’s why big data is synonymous with meaningless charts on pointless dashboards sold to marketing and security managers who never look at them anyway
It’s like tracking the wind and humidity and temperature and barometer data every tenth of a second every square meter.
It won’t help you predict the weather any better than stepping outside and looking at the sky a couple times a day.
Actually, now that I think about it, we should have two products for the users, one let them to query single records as fast as possible without hitting the production OLTP transactional database, even from really big data (find one record from PB level data in seconds), one to power the dashboards that ONLY show aggregation. Is Lakehouse a solution? I have never used it.
And just in the same fashion, there was massive hype around Big Data 1.0. From 2013: https://hbr.org/2013/12/you-may-not-need-big-data-after-all
Everyone has so much data that they must use AI in order to tame it. The reality is, however, is that most of their data is crap and all over the place, and no amount of Big Data 1.0 or 2.0 is ever going to fix it.
https://www.aftonbladet.se/senastenytt/ttnyheter/inrikes/a/8...
Which, of course, was discussed on HN back then: https://news.ycombinator.com/item?id=6398650
However you can aggregate user level info into just the features you need which will get you a looooonnnnggggg way.
Big data folks typically do sampling and such, but that doesn’t eliminate the need for a big data environment where such sampling can occur. Just as a compiler can’t predict every branch that could happen at compile time (sorry VLIW!) and thus CPUs need dynamic branch predictors, so too a sampling function can’t be predicted in advance of an actual dataset.
In a large dataset there are many ways the sample may not represent the whole. The real world is complex. You sample away that complexity at your peril. You will often find you want to go back to the original raw dataset.
Second, in a large organization, sampling alone presumes you are only focused on org-level outcomes. But in a large org there may be individuals who care about the non-aggregated data relevant to their small domain. There can be thousands of such individuals. You do sample the whole but you also have to equip people at each level of abstraction to do the same. The cardinality of your data will in some way reflect the cardinality of your organization and you can’t just sample that away.
You wont be partitioned for this case, but the compute you need is just filtering out this set.
But sampling wont get what you want especially if you are doing QC at the business team level about whether the CX is behaving as expected.
The pros:
-- Samples are small and fast most of the time.
-- can be used opportunistically, eg in queries against the full dataset.
-- can run more complex queries that can't be pre-aggregated (but not always accurately).
The cons:
-- requires planning about what to sample and what types of queries you're answering. Sudden requirements changes are difficult.
-- data skew makes uniform sampling a bad choice.
-- requires ETL pipelines to do the sampling as new data comes in. That includes re-running large backfills if data or sampling changes.
-- requires explaining error to users
-- Data sketches can be particularly inflexible; they're usually good at one metric but can't adapt to new ones. Queries also have to be mapped into set operations.
These problems can be mitigated with proper management tools; I have built frameworks for this type of application before -- fixed dashboards with slow-changing requirements are relatively easy to handle.
This is the thesis for https://rowzero.io. We provide a real spreadsheet interface on top of these data sets, which gives you all the richness and interactivity that affords.
In the rare cases you need more than that, you can hire engineers. The rest of the time, a spreadsheet is all you need.
[1] I made these up.
Big data is dead - https://news.ycombinator.com/item?id=34694926 - Feb 2023 (433 comments)
Big Data Is Dead - https://news.ycombinator.com/item?id=33631561 - Nov 2022 (7 comments)
more accurately, "big data" or its main paradigms are not for everyone. on the other hand, most teams are still unable to reasonably consolidate their data to leverage the benefits of these tools.
"AI" will have a similar fate in a few years time for similar reasons.
AI scientists will propose all sorts of elaborate complex solutions to problems using LLMs, and the dismissive responsive will be “Your problem is solvable with a couple if statements.”
Most people just don’t have problems that require AI.
Similarly, a lot of companies talk about how they have tons of data, but there’s never any real application or game changing insight from it. Just a couple neat tricks and product managers patting themselves on the back.
Setting up a good database is probably the peak of a typical company’s tech journey.
E.g. I would argue that its translation capabilities alone make GPT-4 worthwhile, even if it literally couldn't do anything else.
It seems to conflate somewhat with SV companies completely dismantling privacy concerns and hoovering up as much data as possible. Lots of scenarios I'm sure. I'm just thinking of FAANG in the general case.
There are still bags to be made if you can scare up a CTO or one of his lieutenants working for a small to medium size Luddite company. Add in storage on the blockchain and a talking AI parrot if you want some extra gristle in your grift.
Long live the tech grifters!
The talks were all concentrated around topics like: ingesting and writing the data as quickly as possible, sharding data for the benefit of ingesting data, and centralizing IoT data from around the whole world.
Back then I had questions which were shrugged off — back in the day it seemed to me — as extremely naïve, as if they signified that I was not the "in" crowd somehow for asking them. The questions were:
1) Doesn't optimizing highly for key-value access mean you need that you need to anticipate, predict, and implement ALL of the future access patterns? What if you need to change your queries a year in? The most concrete answer I got was that of course a good architect needs to know and design for all possible ways the data will be queried! I was amazed at either the level of prowess of the said architects — such predictive powers that I couldn't ever dream of attaining! — or the level of self-delusion, as the cynic in me put it.
2) How can it be faster if you keep shoving intermediate processing elements into your pipeline? It's not like you just mindlessly keep adding queues upon queues. That had never been answered. The processing speeds of high-speed pipelines may be impressive, but if some stupid awk over CSV can do it just as quickly on commodity hardware, something _must_ be wrong.
Also larger enterprises don’t even use gcp all that much to begin with
My take on the killer application is the climate change for example earthquakes monitoring. For a case study China has just finished building world's largest earthquake monitoring system with the cost of around USD1 Billion across the country with 15K stations [1]. Somehow at the moment is just monitoring existing earthquakes. But let's say there is a big data analytics technique can reliably predicts impending earthquake within a few days, that can probably safe many people and China now still hold the records of the largest mortality and casualty numbers due to earthquakes. Is it probable, the answer is a positive yes based on our work and initial results it's already practical but in order to do that we need integration with comprehensive in-situ IoT networks with regular and frequent data sampling similar to that of China.
Secondly, China also has the largest radio astronomy telescopes and these telescopes together with other radio telescopes collaborate in real-time through e-VLBI to form a virtual giant radio telescopes as big as the earth to monitor distance stars and galaxy. This is how the black hole got its first image but at the time due to logistics one of the telescope remote disks cannot be shipped to the main processing centers in US [2]. At that moment they are not using real-time e-VLBI onky VLBI, and it tooks them several months just to get the complete sets of the black holes observation data. With e-VLBI everything is real-time and with automatic processing it will be hours instead of month. These radio telescopes can also be used for other purposes like monitoring climate change in addition to imaging black holes, their data is astronomical pardon the pun [3].
[1] Chinese Nationwide Earthquake Early Warning System and Its Performance in the 2022 Lushan M6.1 Earthquake:
https://www.mdpi.com/2072-4292/14/17/4269
[2] How Scientists Captured the First Image of a Black Hole:
https://www.jpl.nasa.gov/edu/news/2019/4/19/how-scientists-c...
[3] Alarmed by Climate Change, Astronomers Train Their Sights on Earth:
https://www.nytimes.com/2024/05/14/science/astronomy-climate...
> There are some cases where big data is very useful. The number of situations where it is useful is limited
Even though there are some great use-cases, the overwhelming majority organisations, institutions, and projects will never have a "let's query ten petabytes" scenario that forces them away from platforms like Postgres.
Most datasets, even at very large companies, fit comfortably into RAM on a server - which is now cost-effective, even in the dozens of terabytes.
Another upcoming example is the latest 5G DECT NR+ standards (the first non-cellular 5G), it will only fuels these massive accumulation of datasets and these monitoring sensor devices do not even get connected to the Internet (think of private factory networks) [1].
Apparently there are limited number of human using or having sensors but for non-human based devices the sky is the limit. For human based communication the data is very limited, we rarely communicate with each others and most of our data now is based on our intermittent media consumptions while streaming audio/video [2]. For IoT sensors devices, they mainly have regular and frequent interval sampling that probably in the ranges of every seconds, minutes, hours, etc. Some if not most of this data is not clean data, there are raw data, and raw data is inherently big and huge compared to the data, for example raw image data vs JPEG data, where the former can be several time bigger in size and processing requirements.
[1] DECT NR+: A technical dive into non-cellular 5G:
https://news.ycombinator.com/item?id=39905644
[2] 50 Video Statistics You Can’t Ignore In 2024: