Does my data fit in RAM?
yourdatafitsinram.net
yourdatafitsinram.net
* Python list of N Python floats: 32×N bytes (approximate, the Python float is 24 bytes + 8-byte pointer for each item in the list)
* NumPy array of N double floats: 8×N bytes
* Hey, we don't need that much precision, let's use 32-bit floats in NumPy: 4×N
* Actually, values of 0-100 are good enough, let's just use uint8 in NumPy and divide by 100 if necessary to get the fraction: N bytes
And now we're down to 3% of original memory usage, and quite possibly with no meaningful impact on the application.
(See e.g. https://pythonspeed.com/articles/python-integers-memory/ and https://pythonspeed.com/articles/pandas-reduce-memory-lossy/ for longer prose versions that approximate the above.)
Efficient representation should be something you build into your data model, it will save you time in the long run.
(Also if you have 100s of columns you're hopefully already benefiting from something like NumPy or Arrow or whatever, so you're already doing better than you could be... )
This is the argument I've been having my whole career with people who claim the better way is "too hard and too slow" .
I'm like "gee, funny how the thing you do the most often you're fastest at... could it be that you'd be just as fast at a better thing if you did it more than never?" .
Good thing the status quo requires no evidence, but any change we want to propose? Impossibly high standards.
Promotable projects at Google even have (or at least used to have) a complexity requirement. You can guess where the incentives lead.
The data I work with is messy, from hand written notes, multiple sources, millions of rows, etc etc. A single point that's written as "one" instead of 1 makes your whole idea fall on its face.
But, I have workarounds for these issues by loading everything into postgres under TEXT columns in a "raw" schema, then do some typecast tests in a descending list of types to get the smallest possible type to transfer to a new table in a "prod" schema. It's read-only data, so it's not a big deal to run it once, and builds out a chain of changes from csv -> sql.
Something like this could be done with pickling to avoid having to re-type every time I run the code (and I've done that for some past projects, but it's... ehhh).
"Oops, messed up my import slightly. Gotta run it again and wait ages.... again"
"Oops, loaded the dataframe twice on accident and had OOM. Gotta restart from the beginning... again"
"Oops, forgot to .head(5) my dataframe and jupyter's crashed... again..."
Doing everything in SQL solves so many problems. And OOM is practically a non-issue.
If per column then that is hopelessly slow. 500+ minutes per dataset and I may have dozens of datasets.
Store at full precision, process at fractional precision, a story as old as time.
https://github.com/faster-cpython/ideas/discussions/138
Victor Stinner's experiment showed some performance regressions, too:
https://github.com/vstinner/cpython/pull/6#issuecomment-6561...
There have definitely been attempts to modernize; the HPy project (https://hpyproject.org/), for instance, moves towards a handle-oriented API that keeps implementation details private and thus enables certain optimizations.
I wish they'd had another go. Call it Python 4.
But getting most people away from Python 2 to Python 3 already took a really long time.
This is one reason projects like Twitter popularized serializations like json-stream in the past, to make it even easier to incrementally load a large file with basic software. Formats like TSV and CSV are also trivially easy to load with streaming IO.
I think the mark of good data formats and libraries is that they allow for this. They should not force an in-memory all or nothing approach, even if applications may want to put all their data in memory. If for no other reason, the application developer should be allowed to commit most of the system RAM to their actual data, not the temporary buffers needed during the IO process.
If I want to push a machine to its limits on some large data, I do not want to be limited to 1/2, 1/3 or worse of the machine size because some IO library developers have all read an article like this and think "my data fits in RAM"! It's not "your data" nor your RAM when you are writing a library. If a user's actual end data might just barely fit in RAM, it will certainly fail if the deep call-stack of typical data analysis tools is cavalier about allocating additional whole-dataset copies during some synchronous load step...
The realization that modern servers can easily persist multiple terabytes of data is profound.
The fact that some datasets are just floats and you can quantize some floats from 32-bits down to 8-bits is true but not a helpful observation.
I also don’t know where you get “Python float is 24 bytes + 8-byte pointer for each item in the list”. Wat.
Latency (hand wavey) L1: 1 ns L2: 2.5 ns L3: 10 ns RAM: 50 ns SSD: 50,000 ns / 50 us
Those are very approximate and specifics vary.
RAM is 5 to 50 times more latent than cache. SSD is ~1000 times slower than RAM. And a spinning HDD is I think ~100 times slower than an SSD.
If your database fits in RAM you’re likely in a happy place. If your DB is so massive it needs to spill to disk you wind up with a mountain of complexity. Multiple machines, sharding, hot/cold etc.
The point of the article is “modern servers have a lot of RAM and you might be able to delete a lot of complexity if you throw money at a server with 4 terabytes of RAM. This option is more practical than you might have realized!”
Not sure how many ways there are to reword that. A CPython float takes 24 bytes of memory, and storing them in a list means 8 bytes per item for the pointer. So in CPython, a list of n floats takes 32n bytes of memory.
>>> sys.getsizeof(1.0)
24
>>> l = list(map(float, range(62_500_000))) # memory use goes up by >2 GB
>>> del l # memory use goes down by >2 GB
(no need to go straight to NumPy to avoid this when relevant, though – array.array is built in.)8 bytes for the 64-bit double 8 bytes for object type pointer 8 bytes for reference count
Lists are dynamic arrays with object pointers, some another 8 bytes per pointer.
TIL and I stand corrected on that point. I thought CPython was more clever than that for lists of floats, but apparently not!
The desktop application, written in an early 2004-ish version of C# and .NET regularly brought the workstations to their knees and thrashed memory and pegged the CPU when large 1080p images were being loaded in or moved around.
Each RGBA image, stored as a PNG on the HDD, was loaded in, each RGBA 32-bit unsigned integer was unpacked into its attendant R, G, B, and alpha components into 32-bit unsigned integers themselves, which were then stored in an Int (boxed the native to a non-native type) which were then appended to a dynamically allocated at read time ArrayList. Deep copies were made of each image any time one of them was resized or manipulated, keeping the original untouched image in RAM in case it was needed. A copy of the image before transformation was stored on the Undo stack. And the 1080p working surface of the screen was super-sampled at 8x resolution to support defringing of images when layered. All stored as 8-bit RGBA components in boxed 32-bit integers in a dynamically allocated ArrayList.
- The PNG specification contains plenty of information including width and length that allow you to use 2d arrays rather than ArrayLists
- The RGBA 32bit integer doesn’t need to be unpacked.. you can just perform shifts and transformations of each of the 8bit channels.
- If you do unpack, you should use native 8bit unsigned ints
- You don’t need an unmodified copy.. that’s your file
- Undo stack should contain deltas, not the whole image over and over again.
Edit: formatting
I expect that you would get quite a performance boost with just that.
go on amazon and buy another stick of RAM
There's one severely under-appreciated factor that favors the first option: computers are commodities that can be acquired very quickly. Almost instantaneously if you're in the cloud. Skilled lower-level programmers are very definitely not commodities, and growing your pool of them can easily take months or years.
And if the problem is big enough, buying hardware will cause operational problems, so you'll need more people. And most likely you're not gonna wanna spend on people, so you get a bunch of people who won't fix the problem, but buy more hardware.
That's why people love the cloud.
And yet, people still regularly choose to go down a path that leads there. Because business decisions are about satisficing, not optimizing. So "I'm 90% sure I will be able to cope with problems of this type but it might cost as much as $10,000,000" is often favored above, "I am 75% sure I might be able to solve problems of this type for no more than $500,000," when the hypothetical downside of not solving it is, "We might go out of business."
You can go from 10 -> 12, where a better design would get you from that same 10 -> 50
It's too difficult to renegotiate with computers, too easy to renegotiate with people. When you don't actually know what you need to do, you need people. When you think you know what you need to do, but you're wrong, then you really need people. Most of us are in the latter category, most of the time.
Did the results of the calculations drive that decision?
This is why I recreated the site when it went down quite a while ago.
The recent article "Use One Big Server"[0] inspired me to (re)submit this website to HN because it addresses the same topic. I like this new article so much because in this day and age of the cloud, people tend to forget how insanely fast and powerful modern servers have become.
And if you don't have budget for new equiment, the second-hand stuff from a few years back is stil beyond amazing and the prices are very reasonable compared to cloud cost. Sure, running bare metal co-located somewhere has it's own cost, but it's not that of a big deal and many issues can be dealt with using 'remote hands' services.
To be fair, the article admits that in the end it's really about your organisation's specific circumstances and thus your requirements. Physical servers and/or vertical scaling may not (always) be the right answer. That said, do yourself a favour, and do take this option seriously and at least consider it. You can even do an experiment: buy some second-hand gear just to gain some experience with hardware if you don't have it already and do a trial in a co-location.
Now that we are talking, yourdatafitsinram.net runs on a Raspberry Pi 4 which in turn is running on solar power.[1] (The blog and this site are both running on the same host)
[0]: https://news.ycombinator.com/item?id=32319147
[1]: https://louwrentius.com/this-blog-is-now-running-on-solar-po...
I have a few second-hand HP/Dell/Supermicro systems running colocated. I find that for all software issues, remote management / IPMI / KVM over IP is perfectly sufficient. Remote hands are needed only for actual hardware issues, most of which is "replace this component with an identical one". Usually HDD, if you're running those. Overall, I'm quite happy with the setup and it's very high on the value/$ spectrum.
Remote hands is for the inevitable hardware failure (Disk, PSU, Fan) or human error (you locked yourself out somehow remotely from IPMI).
P.S. I have a HP Proliant DL380 G8 with 128 GB of memory and 20 physical cores as a lab system for playing with many virtual machines. I turn it on and off on demand using IPMI.
It's definitely worth the effort to script starting the KVM, and maybe even the sol. If you've got a bunch of servers, you should script the power management as well, if nothing else, you want to rate limit power commands across your fleet to prevent accidental mass restarts. Intentional mass restarts can probably happen through the OS, so 1 power command per second across your fleet is probably fine. (You can always hack out the rate limit if you're really sure).
[1] I don't need a whole server, but for $30/month when I wanted to leave my VPS behind for a few reasons anyway...
My most recent prototypes use a hybrid mechanism that dramatically increases the supported working set size. Any property larger than a specific cutoff would be a separate read operation to the durable log. For these properties, only the log's 64-bit offset is stored in memory. There is an alternative heuristic that allows for the developer to add attributes which signify if properties are to be maintained in-memory or permitted to be secondary lookups.
As a consequence, that 2TB worth of ram can properly track hundreds or even thousands of TB worth of effective data.
If you are using modern NVMe storage, those reads to disk are stupid-fast in the worst case. There's still a really good chance you will get a hit in the IO cache if you application isn't ridiculous and has some predictable access patterns.
SQLite or PostgreSQL can be given some configuration/hints to be more aggressive about using RAM while still having their built-in capability to spill to storage rather than hit a hard limit. Or on Linux (at least), just allowing the OS page cache to sprawl over a large RAM system may make the IO so fast that the database doesn't need to worry about special RAM usage. For PostgreSQL, this can just be hints to the optimizer to adjust the cost model and consider random access to be cheaper when comparing possible query plans.
Once you do some sanity check benchmarks of different systems like that, you might find different bottlenecks than expected, and this might highlight new performance optimization quests you hadn't even considered before. :-)
Oh I absolutely have gone down this road as well.
The biggest thing for me is taking advantage of the other benefits you can get with in-memory working sets, such as arbitrary pointer machine representations.
When working with a traditional SQL engine (even one tuned for memory-only operation), there are many rules you have to play by or things will not go well.
[0] https://15721.courses.cs.cmu.edu/spring2020/slides/02-inmemo...
[1] https://15721.courses.cs.cmu.edu/spring2020/papers/02-inmemo...
Around 2000, a guy told me he was asked to support very significant performance issues with a server running a critical application. He quickly figured out that the server ran out of memory. Option 1 was to rewrite the application to use less memory. He chose option two: increase the server memory, going from 64 MB to 128 MB (Yes MB).
At that time, 128 MB was an ungodly amount of memory and memory was very expensive. But it was still cheaper to just throw RAM at the problem than to spend many hours rewriting the application.
For example, four years earlier, it would be >$3k. A year before that, >$6k.
This was an amazing, and terrible, time to be into computers.
I want one of these.
a system with 1TB of ram is 133k, 8.5mil for a system with 64TB of ram?
256GB DIMMs go for ~3k, so 133k for just 1TB is still a bit much.
Significant % of top 500 fortune used it
Sometimes sharding will yield better results. It will still take a long time to read through 64 TB of memory, even if you are just NOPing on it.
Our dataset was ~150 GiB, I think? All in RAM. Took a while to start the server, as it all came off disk. Could have been faster. (It borrowed Redis's query language, and its storage was just "store the commands the recreate the DB, literally", IIRC. Dead simple, but a lot of slack/wasted space there.)
Overall not a bad database. Latency serving out of RAM was, as one should/would expect, very speedy!
[1]: https://tile38.com/
http://www.h2database.com/html/features.html#in_memory_datab...
[0] https://dev.mysql.com/doc/refman/8.0/en/memory-storage-engin...
[0] https://twitter.com/garybernhardt/status/600783770925420546
That's the amount of RAM on an IBM Power E980 System.
I though the issue with ram was the much higher costs.
Also, could you put all of postgres dB in ram, like with redis? Thus no need for separate redis server since your entire dB is fast in ram.
Though at scale, sure, one could even run a RAM disk for speeding up lots of traditional software, with some caveats.
Does my data fit in RAM? - https://news.ycombinator.com/item?id=22309883 - Feb 2020 (162 comments)
But in common use it's a singular mass noun.
Like "information", "data" is correctly used as a singular mass noun.
Even if it were done as a separate page, it should still be shown or hidden via the CSS display property in order to avoid this reloading.