24 Gigabytes of Memory Ought to be Enough for Anybody
codinghorror.com
codinghorror.com
I once did several hours of work trying to optimize my use of Redis to avoid having to upgrade my VPS (I was nearing the limits of the physical memory at 1.5 GB). I even asked Thomas for advice on how to decrease memory usage. His reply: "How much to the next tier?" "$30 a month." "Why are we having this conversation?" And, of course, he was right.
Especially keep this in mind with regard to meetings. A half-dozen people in a meeting burns up about a day of developer time in cost every hour they meet.
I head this a lot from people who write O(n^2) algorithms instead of O(log n) :-P
Sometimes O(n^2) algorithms are ok even when O(log n) alternatives exist. If it takes a trivial amount of time to write the slow algorithm whereas it would take a lot more time to write the better one and if you can guarantee the input complexity the algorithm will handle won't exceed a certain range, then it's fine.
Consider a factorial function for integer input & output, for example. What's the best way to implement such a thing?
It actually doesn't matter. The most important aspects are to get the error handling and input bounds assertions right, because 21! is larger than 2^64, and the difference between recursion, iteration, and caching at that scale is probably just noise unless you are calling the function thousands or millions of times a second.
P.S. This is a good reason why clear documentation outlining design decisions and their rationale is important. When someone decides to take some pre-existing code that "works great" and massively expand how it's used without changing it the result can be disastrous if people haven't taken into account the inherent limitations of that code.
On the other hand, I once saw a project where people seemed to want to use a non-SQL database to store about a hundred megabytes of metadata a year. They also want to use S3 for a couple terabytes of content instead of a filesystem.
Throwing money on a problem that doesn't need to be solved, just because you can, is tempting.
I would think RAM preservation would be a much greater concern in the VPS world as the hard limits are ever present, and quickly become enormously expensive.
In the dedicated server world it isn't quite as prohibitive. An R810 can be had with 512GB. Just got several new servers to add to the mix, each with 144GB of memory.
Though of course I eminently disagree with Atwood's statement about algorithms. In fact I think he doesn't believe that either, but it was just a bit of color on the entry.
Once your data set doesn't fit in memory, it is indeed a scarce resource. And 24 GB is really not very much data.
There will never be a substitute for good algorithms. In my field (bioinformatics), thanks to next-gen sequencing technology, the data growth actually outpaces Moore's law. I imagine the same is true for data collected from social networks.
(Referring to Andy Grove of Intel CEO fame 87-98, still there in some capacity, and Bill Gates.)
1. Your smarter competitor will both throw RAM at the pb and optimize their code, and end up outperforming you.
2. Your quad-socket machine already has RAM maxed out and adding more RAM (ie. adding a second server) will need significantly more complex code to distribute the workloads on more than 1 machine. Optimizing your memory usage is probably easier.
3. Increasing RAM is just not an option on a mobile platform (eg. cellphone).
4. You will lose customers by upping the minimum RAM requirements for your app.
5. Once you have added more RAM, you have no more tricks left in your bag to quickly scale up in case of emergencies. Instead work on optimizing your code while you have time.
You can chuck as much RAM as you like at your problem but it's not going to help once your data set fits in memory - and maybe way before that if you're thrashing those piddly L3 and L2 cache's.
Are you perhaps missing that it's a deliberate inversion of the more obvious statement "needing more RAM is for people who don't know how to use Algorithms" in order to make the point that RAM is cheaper than an engineer's time?
And that's why knowing your hardware and your workload is important. I remember once we had a performance problem with a x86 server. Over lunch, the guy with the problem told me the working dataset was about 8 megs per batch or so. I told him to move the workload to an Itanium server we had for testing because I knew that server had 11 megs of L2 cache, besides being slightly faster per instruction than the Xeon that was running the workload. The throughput more than tripled.
It breaks my heart when people just assume more memory/GHz/cores is always faster.
Why am I in school?
- turn the screw: US$ 5
- knowing which screw to turn: US$ 5000
Algorithms are a poor substitute for us poor engineers that don't have money to throw :)
You need to get out more if you think (as your comment would seem to suggest) that only Google has efficiency problems...
Google has to add $xx * the number of its' servers vs. $xxxx fixing it by one or two developers.
But you do bring up an interesting point, how does Google handle the heterogeneous server problem? One assumes that not all nodes are created equal (if only because it would be cost prohibitive to upgrade all at once), so do you just use dynamic task scheduling and hope for the best or is there something cooler going on there?
I like your blog but please never write something like algorithms are for people who can't afford more ram. Memory management and efficiency should always be a goal, or else you end up with browsers using 2gb of ram with 5 tabs open, and more crap like that.
Yes, memory is cheap. That shouldn't promote terrible code. Vostok4 on January 21, 2011 12:09 AM
That was my first thought.
I still have a laptop I bought in '07 with just 1G of RAM. It works fine and I've never felt compelled to drop a grand or two every couple of years just so that I can.. what? Use more system resources to do the same things I always do?
The problem I've been having with NOT keeping up in the tech arms race at home is that this poor laptop can easily start paging memory with only a couple of Gnome apps open and FF running. It used to be a pretty decent machine.
There are contexts where it's "cheaper" to throw hardware at the problem. I don't think the desktop is one of those areas. Not everyone using them is an engineer being paid over a hundred dollars an hour.
As much more of a hardware guy myself, I've always been frustrated by this path of thinking. Sure, it's convenient for you as the programmer if you don't have to think about resources, but it always feels like the growth in hardware performance is mostly consumed by more and more needless resource usage.
Always reminds me of how cars are getting more and more efficient, powerful and clean-burning, and yet mpg is only barely creeping upwards.
Alternatively, an anecdote from my father- he was once tasked with upgrading the capacity of a company NAS, which was a few hundred megabytes, that was closing in on max capacity. He arranged to bring on some 20x that capacity, thought "well, this should do the job for a while", and it was filled in a matter of weeks.
Inexpensive motherboards are a terrible idea, they can cause really weird problems that look like issues with other components. I've built a few computers and since sworn never to buy a cheap motherboard again.
3 years ago, I purchased an Asus Republic of Gamers motherboard for 400 dollars. It was the most expensive component in my PC, and I hate it every minute of every day.
All my work is in bootloaders, but this motherboard has a terribly, terribly buggy BIOS. Restarting the PC is not a fun thing to do with this 400 dollar PoS.
Of course, the only reason BIOS upgrades failed to address this is because only a few thousand were ever sold... and mostly to gamers and not hard-core computer engineers.
Ten years ago a typical workstation had around 32GB of disk storage. Today, 32GB of RAM in a workstation is perfectly ordinary.
Take the size of your local disk space today. Ten years from now that will be the amount of RAM in your computer.
We're scheduled to hit 11nm; after that I'm not sure. See the Wikipedia article on quantum tunnelling for an explanation that's mostly above my comprehension.
http://en.wikipedia.org/wiki/Quantum_tunnelling
edit: To get an idea of the scales we're talking about, 11nm is only about 50 silicon atoms wide.
http://www.wolframalpha.com/input/?i=11nm+/+width+of+silicon...
http://en.wikipedia.org/wiki/Jim_Gray_(computer_scientist)
It's held true for a long time. Like any rule of thumb, at some point it may no longer be correct, but we're not at that point yet.
So, do your part, and upgrade soon. ;-)
So all memory will be the same.
We really don't have very long left with the current paradigm. Past about a decade or so, we'll need something new if we want to continue growth.
I look around me and most of new systems have 4 GB of RAM. It's possible to have 32 GB of RAM, sure, but it is absolutely uncommon, and blatantly unnecessary (none of my own desktop machines have more than 2 GB).
All of this is documented in the Chromium wiki, of course. But the point is, 4GB memory is their minimum requirement, and insufficient for doing a lot of work on a project like Chromium.
20GB about 2 years later.
Of course, it also has 32 cores... It's not exactly your standard i7 workstation.
Almost all of our machines are either 32GB or 16GB.
As btmorex pointed out earlier, "Once your data set doesn't fit in memory, it is indeed a scarce resource. And 24 GB is really not very much data."
However, I am a scientist, and my machine at work has 96GB RAM with 24Cores, so at the end it comes to around 2GB/Core, which isn't that much anymore. In order to run algorithms on big data and not to bother about disk accesses (SSD or not) more RAM is just crucial. My previous machine had only 8GB ram and it was a big problem to stress algorithms with big data sets on them. So in my case, there are never enough RAM ;)
His is bigger than yours.
And whilst I really don't want to engage in this kind of dick waving, since it's Jeff Atwood I will.
I have a HP Z800 ( http://h10010.www1.hp.com/wwpc/us/en/sm/WF06a/12454-12454-29... ) at home with 48GB RAM currently and it's upgradeable to 192GB RAM.
Of course, having hardware that you're not using is just dumb... why have such a large capex for home hardware if it's not required. I only have this beast because I worked on a project in my spare time around distributed data. For this I built 17 virtual machines with which to test my work. I calculated the cost of the opex of renting this in the cloud would exceed the cost of purchasing over the duration of the project. And because spending money on this kinda hardware is still dumb I then put it to work on a piece of software I sold to a local consulting company. The result is that at least the machine I have has been and is used, and has paid for itself a couple of times over.
BTW, this computing power really helps to build Chromium in 25 minutes rather than hours that most people experience.
Also... since I'm now thinking of the Blackbird ground speed check story ( http://groups.google.com/group/rec.aviation.stories/browse_t... ), someone else has got to step forward and obliterate my home machine with some more impressive dick waving.
It pegs a couple of cores but the others have relatively low load. I've followed the optimisation advice ( http://www.chromium.org/developers/how-tos/build-instruction... ) but can't manage to get it below 25 minutes.
The bottleneck certainly isn't disk or RAM.
Anyway, I like Jeff Atwood's writing because it's true geek celebration and entertainment. If you don't like him, maybe you can just pass on?
Definitely not "blah".
Nailing down the improvement in build time to any single change would take more time substituting things than I'd care to expend right now. The extra memory for file cache no doubt helped quite a bit, as did the SSD, but different parts of the build are I/O heavy, and other bits are CPU dominated. My MacBook Air 13" (4GB RAM, SSD etc., but slow processor, running Windows 7) isn't particularly fast at the build.
And I think the more exciting issue is the falling cost of SSDs. I was recently reading about in-memory databases being much faster not just because of faster access, but because they spend so much less CPU time working around disk access issues. Can the whole OS be written that way? We are quickly heading toward exciting times.
Also, SO-DIM became quite cheap: it is really affordable to get 8 Gb in a laptop (IMO, the difference between 4 and 8 is big because @ 4 Gb, running more than one vm on mac os x is not so practical, whereas at 8 it is a no brainer).
However, most people don't seem to be able to tell where that stops making sense, and the end result is software that's slow no matter how much memory you throw at it.
Anyone that has worked for any amount of time with "enterprise" software knows what I'm talking about.
Hypothetical example: Boss asks if you can fit several millions of products into the database. Keep them up-to date by setting related prices on them. Each synchronization task needs to be aware of product deltas: is it available, sold out, on sale and so on. Oh and there are many suppliers, catalogs. All the while you're running set intersect, difference and so on on multi-million count objects, you still need to be aware that access to this data needs to be consistent and sane. Access times of several hundreds of milliseconds are not acceptable.
Obvious solutions is to distribute the load, counting and updating amongst machines. Load publicly available data from high performance memory tables. At some point, it all comes unravelled when you get a request for custom sub-sets where each one can live queries by price, text, brand ...
Memory bound problems are just that. At some point you hit the wall. I think planning for that impact is the best insurance a shop can make.
I predicted this happening in a piece I posted in 2008.
http://clubtroppo.com.au/2008/07/10/shared-hosting-is-doomed...
And the HN discussion. http://news.ycombinator.com/item?id=241952
Also consider that though the article indicates a ~$900-1000 for a full 24GB system, it may be only ~$3000 for a S5520SC-based system with 96GB. I suppose one would be good for horizontally-scaled web caches and the other for big databases or whatever funky app might make tons of random seeks over a large dataset.
Sure I am really glad that my work PC has more than that, so I can open several IDEs with large solutions and not worry about memory, but for home - I still don't see a need to upgrade.
Why am I not surprised to hear that from Mr. Atwood?
Yeah, 24GB is nice and stuff but the "not care about memory"-part rings my alarm bells. I give him 2 or 3 months of careless coding till we see a "64gb ram ought be enough" post ...