Does my data fit in RAM?
yourdatafitsinram.net
yourdatafitsinram.net
Raw RAM space is not the issue... it's indexing and structure of the data that makes it process-able.
If you just need to spin through the data once, there's no need to even put all of it in RAM - just stream it off disk and process it sequentially.
If you need to join the data, filter, index, query it, you'll need a lot more RAM than your actual data. Database engines have their own overhead (system tables, query processors, query parameter caches, etc.)
And, this all assumes read-only. If you want to update that data, you'll need even more for extents, temp tables, index update operations, etc.
But it works so well and it's so easy. It's a really difficult habit to kick.
Obligatory joke: "Oh boy, virtual memory! Now I can have a really big RAM disk!"
Their changes did not do anything, everything was getting cached. It was cool, but one needs to know what is happening with their data. I am just glad they asked before going to management saying their tests only took 5 minutes.
We eventually had one server entirely replaced at HP's cost after yelling at them long enough, and that one never worked well enough to ever use in production, either. I'd say we had maybe a 70-80% success rate with those servers. They were beasts, though, with 4TB of RAM as I recall, and 288 cores.
There are tons of problems that need to process large data, but touch each item just once (or a few times). You can go a really long way by storing them in disk (or some cloud storage like S3) and writing a script to scan through them.
I know, pretty obvious, but somehow escapes many devs.
Needless to say I wasn't recommended for the job, and it taught me a valuable lesson: if you don't first give them what they want, you can't give them what they actually need.
As soon as you're touching it more than once, sticking it in RAM upon reading makes everything much faster.
The lowest price floor per GB has been similar for the past decade. Roughly at $2.8/GB in 2012, 2016, and 2019. And all DRAM manufacturers has been enjoying a very profitable period.
And yet our Data size continue to grow. We can fit more Data inside memory not because DRAM capacity has increase, but we are simply increasing memory channels.
https://thememoryguy.com/dram-prices-hit-historic-low/
You've selected out the low points on the graph: 2012, 2016, and 2019, most of the time DRAM has not been available at these prices. Now is definitely the time to load up on RAM.
That’s not what we in computing refer to as a “collapse”.
And it is predicted to climb back up this year, due to manufacturers dropping wafer starts at a point in time when a large launch of next-gen consoles is drastically increasing consumption.
Nope, definitely no collusion there. /s
I'm afraid I had no idea that had happened at all.
> since early this year,
Since..when?..now? Feb 12th is pretty early in the year AFAIAC
It would be better to reference this as quoted from the article which was written in November 2019. So not really last week
>most of the time DRAM has not been available at these prices.
I did said price floor.
Lowest massively-available price, please and thank you.
>most of the time DRAM has not been available at these prices.
should make it not the floor. If the floor doesen't have to be available, then what's the exact point it becomes relevant? Otherwise the price is simply misleading
A price has to be massively available at a point in time to matter. It doesn't have to be available forever to matter. It feels like you're conflating the two.
The price is on a downward trend, but there are hitches and setbacks. One fair way to measure it is to use some kind of average. Another also-fair way to measure it is to go by the lowest "real" price, where "real" means you can buy something like a million sticks on the open market.
When we're talking about whether we should be impressed by a price, using the lowest historical price for comparison makes sense.
(And just to be absolutely clear, you would need to adjust the metric for a product that goes up in price over time. But for something on a downward trend, this metric works fine.)
The trouble is that there are limits to how much memory you can fit on a motherboard.
[1] https://www.supermicro.com/en/Aplus/system/1U/1113/AS-1113S-...
It doesn't seem to have brought prices down much.
Can you partition the data in any useful way? For example if queries use separate ranges of dates, then you can partition data so that queries only need to touch the relevant date range. Can you pre-process any computations? Sometimes tricky things done within the context of multiple joins can be done once and written to a table for later use. Can you materialize any views? Do you have the proper indexes set up for your joins and filters? Are you looking at execution plans for your queries? Sometimes small changes can speed up queries by many orders of magnitude.
Smart queries + properly structured data + a well tuned postgres DB is an incredibly powerful tool.
> What tool should I use it to load this dataset on RAM and run these queries?
The question should really be: What tool should I use to make this fast?
Postgres can be pretty fast when used correctly and you can make your data fit.
(mongodb I believe uses direct mmap access instead of a pagecache, and lmdb does this as well)
For lots of databases, most of their time is spent locking and copying data around, so depending on your workload you might getsignificant speedups in Pandas/Numpy if it's just you doing manipulations, and there are multicore just-in-time compilers for lots of Pandas/Numpy operations (like Numba/Dask/etc).
If you have lots of weird merging criteria and want the flexibility of SQL I'd say use a modern Postgresql with multicore selects on that 4TB box.
People like to talk about the elasticity or compute, but startup is not free (or even cheap in most cases).
OP is probably talking about an x1e.32xlarge. According to Daniel Vassalo's S3 benchmark [1], it can do about 2.7GB/sec.
So your 4TB DB might take ~30min to fetch.
Bandwidth is free, you'd pay $2 for the 30 min of compute, and some fractions of pennies for the few hundred S3 requests.
I'd give Google BigQuery a shot. Should work fast [seconds] and scale seamlessly to [much] larger datasets and [many] more users. For a 1 TB dataset, I have a hard time imagining crafting a slow query. Maybe something outlandish like 1000[00?] joins. They also have an in-memory "BI Engine" offering, alas limited to 50GB max.
On premise, there is Tableau Data Engine. I don't think they offer a SQL interface, you have to buy into the entire ecosystem.
Long shot: I've been working on "most expressive query system over multiple tables" as an offshoot of some recent NLP work. Your use case piqued my interest. I'd love to help / understand it better. My contact is in my profile.
This makes for a good laugh: Queries for the masses. Contact: f'info@${user}.com'
Putting the data on a ramdisk just becomes entirely redundant because it's still going to create a second memory cache that it uses.
Many operations do a local cache warming by running the common queries over the database before they it is brought online for processing. As a secondary note, people often under-estimate the size of their data because they don't account for all of the keys, indexes and relationships that also would be memory cached in an ideal situation.
Google's BigQuery, AWS Redshift, Snowflake (are all hosted), or MemSQL, Clickhouse (to run yourself). Other options include Greenplum, Vertica, Actian, YellowBrick, or even GPU-powered systems like MapD, Kinetica, and Sqream.
I recommend BigQuery for no-ops hosted version or MemSQL if you want a local install.
MemSQL uses rowstores in memory combined with columnstores on disk. Both can be joined together seamlessly and the latest release will automatically choose and transition the table type for you as data size and access patterns change.
Spark will naturally use RAM first and then disk as needed.
In my experience, it's nearly always the case that pulling in all data is not necessary, and that thinking through your goals, data, and processing can often reduce both the amount of data touched and the processing run on it massively. Look up the article on why GNU grep is so fast for a bunch of general tricks that can be employed, many of which may apply to data processing generally.
Otherwise:
1. Random sampling. The Law of Large Numbers applies and affords copious advantages. There are few problems a sample of 100 - 1,000 cannot offer immense insights on, and even if you need to rely on larger samples for more detailed results, these can guide further analysis at greatly reduced computational cost.
2. Stratified sampling. When you need to include exemplars of various groups, some not highly prevalent within the data.
3. Subset your data. Divide by regions, groups, accounts, corporate divisions, demographic classifications, time blocks (day, week, month, quarter, year, ...), etc. Process chunks at a time.
4. Precompute summary / period data. Computing max, min, mean, standard deviation, and a set of percentiles for data attributes (individuals, groups, age quintiles or deciles, geocoded regions, time series), and then operating on the summarised data, can be tremendously useful. Consider data as an RRD rather than a comprehensive set (may apply to time series or other entities).
Creating a set of temporary or analytic datasets / tables can be tremendously useful. As much fun as it is to write a single soup-to-nuts SQL query.
5. Linear scans typically beat random scans. If you can seek sequentially through data rather than mix-and-match, so much the better. With SSD this advantage falls markedly, but isn't completely erased. For fusion type drives (hybrid SSD/HDD) there can still be marked advantages.
6. Indexes and sorts. The rule of thumb I'd grown up with in OLAP was that indexes work when you're accessing up to 10% of a dataset, otherwise a sort might be preferred. Remember that sorts are exceedingly expensive.
If at all possible, subset or narrow (see below) data BEFORE sorting.
6. Hash lookups. If one table fits into RAM, then construct a hash table using that (all the better if your tools support this natively -- hand-rolling hashing algorithms is possible, but tedious), and use that to process larger table(s).
7. "Narrow" the data. Select only the fields you need. Most especially, write only the fields you need. In SQL this is as simple as a "SELECT <fieldlist> FROM <table>" rather than "SELECT * FROM <table>". There are times you can also reduce total data throughput by recoding long records (say, geocoded names, there are a few thousands of place names in the US, using Census TIGER data, vs. placenames which may run to 22 characters ("Truth or Consequences", in NM), or even longer for international placenames. You'll need a tool to remap those later. For statistical analysis, converting to analysis variables may be necessary regardless.
The number of times I've seen people dragging all fields through extensive data is ... many.
Some of this can be performed in SQL, some wants a more data-related language (SAS DATA Step and awk are both largely equivalent here).
Otherwise: understanding your platforms storage, memory, and virtual memory subsystems can be useful. Even as simple a practice as running "cat mydatafile > /dev/null" can often speed up subsequent processing.
In the machine learning world, some of the algorithms that are industrial workhorses will require you to have your dataset in memory (ie: all the common GBM libraries), and will walk over it lots of times.
You may be able to perform some gymnastics and allow the OS to swap your terabyte+ dataset around inside your 64GB of RAM, but the algorithms are now going to take forever to complete as you thrash your swap constantly while the training algorithm is running.
tl;dr - a terabyte dataset in the machine learning context may very well need that much RAM plus some overhead in terms of memory available to be able to train a model on the dataset.
I just ran `time cp /dev/nvme0n1 /dev/null` on the 1TB 970 Pro. The result:
real 4m50.724s
user 0m2.001s
sys 3m10.282s
So with literally zero optimization effort, we've hit the spec (and saturated a PCIe 3.0 x4 link).https://www.newegg.com/samsung-970-pro-1tb/p/N82E16820147694
What will cause you serious and unavoidable trouble is if you cannot structure things to have any spatial locality. If you only want one 64-bit value out of the 4kB block you've fetched, and you'll come back later another 511 times to fetch the other 64b values in that block, then your performance deficit relative to DRAM will be greatly amplified (because your DRAM fetches would be 64B cachelines fetch 8x each instead of 4kB blocks fetched 512x each).
Having 4/16TiB servers or "memory db servers" as I thought of them solved a lot of problems outright. Still need huge i/o but less of it depending on your workload.
> The kernel and your other processes need some space to malloc, and you dont want to page in/out.
Some space, like "most of 20-50 gigabytes"?
You want to take into account how exactly the space used by joins will fit into memory, but 2-5% of a terabyte is an extremely generous allocation for everything else on the box.
This was a fairly salient point, and remember circa 2012/2013 struggling to fit large bioinfomatics data into an older iMac with base R.
In practice, an awk script frequently ran circles around processing. On a direct basis, awk corresponds quite closely to the SAS DATA Step (and was intended to be paired with tools such as S, the precursor to R, for similar types of processing).
The fact that awk had associative arrays (which SAS long lacked, it's since ... come up with something along those lines) and could perform extremely rapid sort-merge or pattern matches (equivalent to SAS data formats, which internally utilise a b-tree structure) helped.
With awk, sort, unique, and a few hand-rolled statistics awk libraries / scripts, you can replace much of the functionality of SAS. And that's without even touching R or gnuplot, each of which offer further vast capabilities.
And at an aggreable annual license fee.
25 years ago i was suggesting clients to upgrade from 4MB to 6-8MB as it was improving their experience with our business software, these days i've already suggested a couple of customers to upgrade from 6TB and 8TB respectively ... as it would improve their experience with our business software. What's funny is that customer experience with business software back then was better than today.
I’m no fan of distributed systems that sit idle at 2% utilisation when a single node would do : BUT, reducing it down to “cost” and “does it fit in RAM” is way too reductive
These days I’m firmly of the opinion that if you can make it run on a single machine, you absolutely should, and you don’t get more machines until you can prove you can fully utilise one machine.
When a single box cannot handle your load, you start one more, and then you start worrying about scalability.
(Of course, YMMV, I guess there are cases where "one box" -> "two boxes" requires an enormous quantum jump, like transactional DB. For everything else, there's a box.)
Do you know how quickly disks fail if you force them at 100% utilisation, 24/7?
Then what happens when this system dies? How much downtime do you have because you have to replace then hardware then get your hundreds of gigabyte dataset back in RAM and hot again?
I’ve worked as a lead developer at companies where I’ve been personally responsible for hundreds of thousands of machines, and running a node to 100% and THEN thinking about scaling is short sighted and stupid
I mean "when you've proven that the application you've written can fully, or near fully utilise the available power on a single machine, and that when running production-grade workloads, actually does so, then you may scale to additional machines.
What this means is not getting a 9-node spark cluster to push a few hundred gb of data from S3 to a database because "it took too long to run in python" because it's a single threaded, non-async, non-performance tuned.
How is that any different? You just backed off a tiny amount by saying “fully or near fully” - you still shouldn’t burden a single host to “fully or near Fully” because:
It puts more strain on the hardware and will cause it to fail a LOT faster
There’s no redundancy so when the system fails you’ll probably need hours or maybe days to replace physical hardware, restore from backup, verify restore integrity, and resume operations - which after all this work, will only put you in the same position again, waiting for the next failure
Single node systems make it difficult to canary deploy because a runaway bug can blow a node out - and you only have one.
Workload patterns are rarely a linear steam of homogenous tiny events - a large memory allocation from a big query, or an unanticipated table scan, or any 5th percentile type difficult task can cause so much system contention on a single node that your operations effectively stop
What about edge cases in kernels and network drivers - many times we have had frozen kernel modules, zombie processes, deadlocks and do on, again, with only one node something as trivial as a reboot means halting operations.
There’s just so many reasons a single node is a bad idea, I’m having trouble listing them
You're missing the word "can". It's a very important part of that sentence.
If your software can't even use 80% of one node, it has scaling problems that you need to address ASAP, and probably before throwing more nodes at it.
> It puts more strain on the hardware and will cause it to fail a LOT faster
Unless you're hammering an SSD, how does that happen? CPU and RAM should be at a pretty stable amount of watts anywhere from 'moderate' load and up, which doesn't strain anything or overheat.
> redundancy
Sure.
I’m not against redundancy/HA in production systems, I’m opposing clusters of machines to perform data workloads that could more efficiently handled by single machines. Also note here that I’m talking about data science and machine learning workloads, where node failure simply means the job isn’t marked as done, a replacement machine gets provisioned and we resume/restart.
I’m not suggesting running your web servers and main databases on a single machine.
Also did you know that Hadoop is open source. So that $250K is purely for hardware.
(Someone did it recently with 1.5TiB of ram on a Mac Pro. Then Linus from Linus tech tips did it on Windows with 2TiB)
You don't have to run it full time tho'...
Add to that small army of people, because, you know, you need specialists of variety of professions just to debug all integration issues between all those components that WOULD NOT BE NEEDED if you just decided to put your stuff in memory.
Frankly, the proportion of projects that really need to work on data that could not fit in memory of a single machine is very low. I work for one of the largest banks in the world processing most of its trades from all over the world and guess what, all of it fits in RAM.
Story time: I worked on one project where a single (large) building's internal sensor data (HVAC, motion, etc. 100k sensors) would fill a 40TB array every year. They had a 20 year retention policy. So Dell would just add a new server + array every year.
I worked with another company that had 2000 oracle servers in some sort of franken-cluster config. Reports took 1 week to run and they had pricing data for their industry (they were a transaction middleman) for almost 40 years. I can't even guess the data size because nobody could figure it out.
This is not a FAANG problem. This is an everage SME to large enterprise problem. Yeah, startups don't have much data. Most companies out there aren't startups.
By the way, memory isn't the only solution. In the past 15 years, I've rarely worked on projects where everything was in memory. Disks work just fine with good database technology.
That's a lot of data, but what do you even do with it other than take minuscule slices or calculate statistics?
And for those uses, I'd put whether it fits in RAM as not applicable. It doesn't, but can you even tell the difference?
Agreed that 20 year retention is silly. We thought it was silly, but the policies reflected the need for historical analysis for audit purposes.
It does in fact matter what you can fit in RAM though. We had to adapt all our systems to a janky SQL Server setup that was horrible for time series data and make our software run on those servers. RAM availability for working sets was a huge bottleneck (hence the cost of analysis).
Yes your company's data may fit in RAM. But does every intermediate data set also fit in RAM ? Because I've also worked at a bank and we had thousands of complex ETLs often needing tens to hundreds of intermediate sets along the way. There is no AWS server that can keep all of that inflight at one time.
And what about your Data Analysts/Scientists. Can all of their random data sets reside in RAM on the same server too ?
$100K has always been "cheap" for a "business computer" and today you can get more computer for that money than ever.
$100K of hardware (per year or so) is small-fry compared to almost every other R&D industry out there. Just compare with the cost of debuggers, oscilloscopes and EMC labs for electronic engineers.
Buy them a machine each at a cost of $40-60 billion ?
Or would it make more sense to buy one Spark cluster and then share the resources at a fraction of the cost.
Still expensive, but much less than 40 billion.
We have a Spark cluster which supports all of those users for $10-$20k a month.
There is a trend, when the application is inefficient, to spend huge amount of resources on scaling it instead of making the application more efficient.
I mean, I also prefer doing things on a single machine, but if that machine gets expensive enough, or writing a program that can actually use all that power gets too difficult, why not switch to a cloud database?
Look at it this way: This is for the person that's already going to get enough ram sticks to fit the entire data set or multiple of it, across many machines, and deal with the enormous overhead from doing queries across many machines. The revelation is that you can fit that much ram inside a single machine for a much cheaper and faster experience.
We have way too many fucking copies of our DB.
A quick search shows that you can get at least 24TiB from AWS: https://aws.amazon.com/ec2/instance-types/high-memory
I have chosen for the cloud options to only select virtual instances that can be spun up on demand. The high-memory instances you link to are purpose-build.
On the other hand, it is true, they do exist.
I remember jokingly telling my team to spend the millions of dollars we did expanding our clusters with one of these and just processing in RAM. I honestly don't think I did the analysis to make sure it would be actually better.
It was 'jokingly' because we couldn't afford three of these machines anyway. The clusters had the property that we could lose some large fraction of nodes and still operate, we could expand slightly sub-linearly, etc. which are all lovely properties.
It would have been neat, though. Ahhhh imagine the luxury of just loading things into memory and crunching the whole thing in minutes instead of hours. Gives me shivers.
You didn't have to have all the capital. They'd rent you the machine on a long lease and you could say you'd bought a million dollar computer.
Sun tried to get into similar markets but it was always a tougher deal on minis and micros.
The website needs updating.
May I suggest that a) they consider applications need RAM too and b) that if the price is POI, it's probably better to mention "But you know, 188 m5.8xlarges might be cheaper"
I would definitely imagine that most workloads rarely need more than a few to a few hundred TB in memory, since you may have petabytes of data but you probably touch very little of it.
Also, the site that remained wasn't ever updated. So I took some time and updated some links.
I liked updating/creating this because I love the simplicity of the concept / thought behind it.
"Inspired by this tweet" which links to https://twitter.com/garybernhardt/status/600783770925420546
For a few given system configurations, and a given definition of "fits".
Hope you read this before you spent $50,000,000 dollars/month (although you might be able to negotiate a discount).
1. If you have to ask, then either it doesn't now, or it doesn't sometimes. So assume it doesn't.
2. If you can use a cluster, then maybe.
3. In some senses, it doesn't matter. How so? Reading and writing from RAM is very slow, latency-wise, for today's processors. If I can bend the truth a little, it's a bit like a fast SSD. So, if you can up the bandwidth to disk enough, it becomes kind of comparable. Well, if you can use 16 PCIe 4.0 lanes, it's roughly 24 GB/sec effective bandwidth, which is roughly half of your memory bandwidth. Now it's true that in real-life systems it's usually just 4 lanes, but it's very doable to change that with a nice card.
4. DIMM-form-factor non-volatile memory may increase memory sizes much more.