When did we stop caring about memory management?
hanselman.com
hanselman.com
Not having to worry about memory is now a luxury we can afford in the age of plenty. While writing an efficient spellchecker would have been a great feat in the 80s, there's no point of that when you can load the whole dictionary file in memory.
In other words, memory management is a #highclassproblem for many developers--you worry about efficiency once you have a lot of adoption. (Within reason; if your app is unbearably slow that isn't good.)
And why settle for "without optimizing it's slightly faster than things used to be when people had to optimize", when you can have "optimized, it's orders of magnitude faster than things used to be"?
> you worry about efficiency once you have a lot of adoption.
Sometimes things can be tricky to refactor if you didn't move carefully at every step. Initially not worrying about efficiency at all might mean having to rewrite from scratch once you do.
Of course, this can't be generalized. Sometimes it really, really doesn't matter, unless you care about craftsmanship as a value in and of itself -- taking 10-20% more time to write code that makes you happy and proud might very well make you more effective in the long run.
But in the web-focused marketplace, adding 6 months to your launch calendar to get something fast to market may mean someone else's good-enough solution becomes the entrenched, de-facto standard that you are now fighting to prove you're superior enough to that it justifies people's time cost to move their data to your solution.
That's the risk tradeoff.
It's easy to think why not just do it right in the first place. Aside from the timing you, as a startup, don't really know what "it" is or "it" is constantly moving. Committing to perfection at that stage can be a sure way to kill the company. Even if "it" is known and it's not changing you still need to keep in mind the Gall's Law.
However, it has heaped to performance and code separation of concerns and structure debt. Of course, AirBnb can now afford to move things to Java/Scala and things that are better for multinational companies.
For AirBnb, it's really hard to fix a lot of these things since the product requires good uptime.
For a lot of problems, once you have this off-the-shelf solution, you're just not going to have a memory issue. It's like cooking without having to manage where the pots, pans, and plates are. Say you hire some guy to clean up for you, and the guy just implements some simple strategy. It's a lot more pleasant, and if you have a large enough kitchen, it won't make a difference to the quality (and timeliness) of the food. In some cases you may have to consider the guy running around, but most of the time you net win anyway.
For things that are truly performance sensitive, you still need to think about cache issues and such, and you definitely still care about memory management.
I can think of two projects in my past five years that really showed performance sensitivity. I knew that at design time so I chose to care about where data sat in memory and how it would behave in multilevel caches.
All other projects, couldn't care less where the kitchen dude stored my pots, and he never lost one.
None of those projects jumped from performance agnostic to performance sensitive overnight.
Typical L1 cache size for the x86 family is only a few times larger than on the 486. Cache is getting faster, but transistor shrinkage should affect cache speed and instruction execution time in roughly equal proportion.
Of course, the marginal value of spending development time improving performance can vary a great deal; but when performance matters, the key considerations are going to remain location, location, and location for the foreseeable future.
In my experience, many of the issues with debugging performance issues are not related to memory management, but understanding JIT heuristics which are moving targets (especially in the dynamically typed languages).
I think we need to collectively separate the idea of some kind of automated (or semi-automated) collection strategy, from a language level JITed Virtual Machine. You may have one without the other.
[1]https://www.cs.virginia.edu/~cs415/reading/bacon-garbage.pdf
I've not worked on performance critical code running under a JIT, but have worked on performance critical code running under a GC, and having an intimate knowledge of memory can be vital to understanding the performance of your application.
I'd never claim that understanding memory usage/allocation/freeing isn't important. It's important to understand how memory is used whether that be understanding how your collector handles generations of memory in a GC'd language, or whether it be understanding how your programming patterns cause a waterfall of destructor chains in a manually managed language. Often, understanding these things is difficult.
My claim was that, in my experience, the biggest difficulties in debugging performance were in understanding JIT heuristics in language level Virtual Machines. This is because (especially in dynamic languages) the heuristics aren't documented anywhere, they're subject to change, and are likely not consistent between particular runtimes you're hosted within. (Those optimizations you made just for V8, could even hurt perf if you run your library in JSC).
There's also the issue with language level virtual machines in that they usually impose some system wide tax on many operations (a layer of indirection), so I could see the people mentioned in the article having a hard time squeezing out that last portion of performance due to that indirection - again, that might be the fault of there being a language level Virtual Machine - and not necessarily as much related to memory (it depends on your application, and amount you allocate, of course).
In my experience, optimizing collection behavior has been relatively straightforward in comparison and the techniques you learn for one collector seem to translate to others. Perhaps my experience is biased towards JIT/VM/GC in dynamic languages where the JITs are so much more complex in comparison.
I have seen core hot paths in critical client side JS libraries be slowed down 9x because of obscure JIT deopts caused by properties names beginning with a dollar sign. I'm not just talking about a 9x slowdown of field lookups - our entire end to end benchmarks tanked (since fixed). Perhaps the VM architects decided that something with a leading dollar sign would not likely benefit from being made into a hidden JIT class? I've never had anything so obscure happen with any kind of garbage collection, whether that be reference counted ARC, or a tracing collector. Edit: spelling
So, yes, we've traded performance for automatic memory management. But we've gained reliability, safety, and security. For the vast majority of server-side apps, that's a worthwhile trade.
† Unless you use the processes like NASA and the like use, which for most apps are infeasible from a cost point of view.
If course it's possible to write stuff using manual memory management in a manner such that it's nearly impossible to do well. Doctor, doctor, it hurts when I do that...
I wrote a 'C' based thing around 2012 that:
- is about 50,000 lines; 50ish by 1000ish lines. - uses no dynamic memory. - does very "dynamic memory" things.
Each type/shape/class has its own statically allocated 'C' array of structs that is declared to worst-case, and has a "new". Behind the "new" is a lot of behavior, including pointing to other such things/objects.
A set of "delete" operations would be also relatively easy.
But because it's an array and it is in 'C', this is relatively little code. Yes, there is duplication involved but the object creation code is probably 20% of the total thing and each object creation is darn near provably correct. It's built on RAII principles ( perhaps foolishly, since there's no deletion ). Dependencies between objects are completely explicit in the code. Object creation must be in the correct order to manage dependencies
The magic happens because scripts* are used to create all things, once at startup. it also helps because by its very nature, the system is fully serialized - after a creation script, the state of the tables of objects can be written to a text file.
*scripts are, in this case nothing but sequences of verbs which send text files full of XML to a service at a TCP port. It could have been JSON or attribute-value pairs, but instead it's XML.
I do take extra care to make a very comprehensive test suite or it.
but since the allocation scheme is 1) cheap and 2) fully exposed as state in the system, there's nothing actually opaque about it. I don't have to do "trust falls" with a garbage collector or even the heap. You can't have a malloc() or new() failure - it'll violate a constraint and quit/exit(). You then make the table larger.
From the frame of reference of the other objects, it's abstracted and opaque - after the initial setup. From the "scripting" interface, it's fully transparent.
I've never seen a large C or C++ codebase written by a large team that didn't have serious memory management problems.
I just smell a lot of despair in these threads.
What you're likely a victim of is the Dunning-Kruger effect or Schneier's Law depending on which you prefer:
https://www.schneier.com/blog/archives/2011/04/schneiers_law...
On the other hand, you could've in fact designed something as good as you claimed. In that case, I seriously suggest you publish your methods in full for peer review as they could save millions spent by the likes of NASA, seL4 project, and so on. Any method that might hit full correctness with just basic coding or whatever is a potential game-changer. Very exciting stuff.
I doubt anybody is remotely interested in this - the problem domains are pretty boring and highly constrained. These tend to be realtime embedded ( both soft and hard realtime ). It seems wildly inappropriate for gaming, graphics or Web stuff. Perhaps I'm wrong there; not really all that clued in.
There is much less to this than you might imagine. The problem is that a book is a team effort, and I have no idea who to bother with it. Not asking for a volunteer here.
For an idea of the basic philosophy, it might be worth looking at Bruce Powell Douglass.
There's just so many ways that language can get you I have a doubt even a handful of methods would cover them all. It takes Work to get anywhere. Best of luck to you on that.
"For an idea of the basic philosophy, it might be worth looking at Bruce Powell Douglass."
Appreciate the reference. Can't recall if he's in my collection or not. I'll look into his work. :)
And I appreciate the heads up that I may be repeating myself/others, and that this subject is really for long form stuff.
LOL. I have to ask: do you "know" this from experience or someone else told you?
Thousands of applications were made not so long ago with manual memory management. Many of them are still the most used desktop apps in the world.
Methods to secure them are not rocket science, really. I would gladly concede it's cheaper to do CRUD apps with GC, but you're going hyperbolic with NASA ;-)
And pretty much all of them are unreliable garbage that we as an industry should have been too embarrassed to ship.
(of course, vulnerabilities and unreliability in those apps came from more places than just manual memory management, just as today's slowness comes from more places than the GC)
So you have experienced thousands of unreliable garbage apps. I feel really sorry for you. My experience has been the opposite: lots of wonderful and reliable programs. And when I found one that wasn't good, I stopped using it and replaced it with a better one.
I mean as a user. As a programmer I've seen all kind of programs, from "critical mission" to "quick hack" and from excelent to unreliable garbage. And you know what? It's not language features but ludicrous incompetence what makes an app unreliable garbage. There are programmers out there that can make unreliable garbage in any language.
The "secret" to make some app reliable is not reserved to NASA and the like. Follow two simple patterns:
1. For every user action, validate input, try executing the code, catch the occasional exception and show an error message.
2. For every resource, allocate, try{ using it } finally { release it }
Put each of these skeletons in a function or method, so the resource is a local variable (no colateral effects) and it's easy to debug.
That's all. Well, also: don't be too "creative" unless you know what you're doing (most times you don't).
I'm fifty-something, spent a lot of painful time mechanically applying that recipes to rotten codebases, and when I read the kind of criticisms you made I assume they come from people too young to have real world experience on what they're criticizing. The alternative is worse: to be the kind of people that created that rot and then blamed the language.
Just in case: I don't believe that manual memory languages are "better". I love Python and (most of the time) JavaScript.
Actually, every single year I've worked on security-critical software I've understood more and more how impossible it is to secure large C and C++ codebases worked on by large teams with just "careful coding". The picture changes a lot when you have people actively searching for vulnerabilities.
Do you believe that security software or browsers are better writen in other languages than C/C++? I won't dispute that, even if I'd have some questions about relative merits. Please think of some well known Perl and PHP applications and their security record.
Anyway the unqualified:
That's because correct C-style memory management never† scales up to large software projects in a way that doesn't result in crashes and, worse, vulnerabilities.
... and talking about manual memory management in general...
And pretty much all of them are unreliable garbage that we as an industry should have been too embarrassed to ship.
... are simply not true. Worse, they're insulting for everyone who has worked in this kind of application.
Yes. And lots and lots of apps (most public network facing things) are security sensitive.
> Please think of some well known Perl and PHP applications and their security record.
Nobody is saying that "memory safe == no vulnerabilities". I do think it's true that being memory unsafe results in many more vulnerabilities. I've measured this precise statistic in our bug tracker.
> That's because correct C-style memory management never† scales up to large software projects in a way that doesn't result in crashes and, worse, vulnerabilities.
It's true. And I work in native code myself!
It's only an insult if you presuppose that introducing vulnerabilities reflects badly on the programmer who introduced them. I don't believe that; most major vulnerabilities I've seen were introduced by very good programmers. It's a problem with the tools they're working with.
Where did you find reliable software, and what does it do? After twenty years, it all seems to be the shit it always has been.
Yes, I have worked on C++ apps for years (browsers).
What I am saying is what nearly everyone in the security community will tell you, from what I've seen.
Such thinking is why MULTICS used PL/0, why GEMSOS used Pascal, why Wirth used Modula and Oberon for safer OS's, and why Altran-Praxis used Ada/SPARK for a CA. Results were better than C (or predecessor) every time. Deal with root cause, symptoms fade away naturally. Solutions like Eiffel and Rust take that philosophy into modern IT.
Things that went well:
1) RAM got ridiculously cheap: 2x 8GB is now what $60-70? So it would seem, in that regard, yeah you shouldn't care about memory.
2) GC languages and runtimes. Believe or not there was heavy skepticism that GC runtime could be ever be a serious strategy in practice. It was an exotic technology that many thought would left for academia and fringe research area. But most popular frameworks and languages today are GC based.
Some things were a bit un-expected:
1) Memory access speed is fast but it didn't keep up with CPU speed. A cache miss will costs hundreds of nanoseconds. CPU speeds may have been doubling every 2 years or so, but memory access speeds have only been increasing by 10-20% or so.
In the 80s and 90s you just didn't care as much about cache hits and misses. But those are the forefront of performance optimization strategies. "Align your struct members","use structs of arrays instead of arrays of structs","think about cache lines and strides when you loop through your data arrays..." etc.
And because CPU speeds go so much faster, memory access also stands out on performance charts. So GC languages and runtimes can still seems slow and sluggish for applications that don't pay attention to memory.
2) Applications got more concurrent and servers are expected to handle a larger number of clients / connections / requests. That means each one needs a context, so it needs memory. So memory is again at the forefront. Accessing it and its latency, GC pause times are important.
3) Mobile happened. Mobile devices have proliferated so the needs for fast, small footprint applications never went away. Even as there are serves witih terabytes of RAM around, there are even more little arm cores and then everyone talks about IoT. That means langauges/runtimes which compile ahead of time will never be 100% displaced. Heck even on Android, ART (AoT copiling) is replacing JIT.
You can still micro-optimise for memory and cpu performance. That's certainly a thing that is possible to do.
Alternatively, you can buy more cpus and memory.
What has changed in the past few years is that for many workloads, the cost of buying more cpus and memory is significantly lower than the cost of all the development work to optimise away the need. Hardware is cheap compared to developers, for many applications.
There are still some applications where this is not the case, and in those cases people still care about memory management.
On modern CPUs, cache misses take so long and unaligned access is nearly free in comparison, so that padding can become detrimental. Of course ordering members so that they're all aligned is the best option, but if that can't be done, it's better to have misaligned and pack more in the same space than leave wasted padding.
Under the right circumstances packed structs can improve performance, but they come with a lot of technical debt since they can cause sporadic extreme slowdowns on certain common CPUs, and performance problems, aborts, or even data corruption on non-x86 platforms.
The author's example is interesting. He talks about a webserver that after optimizations can reach 1.2 million requests per second. In reality though, your gigabit ethernet interface can't handle that much anyway.
This bites us in the ass, because I can't make Wistia (for example) stop eating all the memory on this thing with my app running next to it.
So, personally, I stopped caring when it was not a priority and didn't, on average, affect user experience. But; it was a mistake and I am watching memory/etc more and more these days. This is for various reasons (old hardware, more than one app on a page, growing amount of bundled media content, etc).
I am working on an Android app, and rewrote the part that handles the connection via a binary protocol over a socket with a remote server. Rewriting that to take advantage of NIO already improved connection times (build up a connection, initialize, etc) from almost 2 minutes to only 8 or 9 seconds. RAM usage of the backend was reduced by 50% – and it is easier readable.
More often than not it’s not that hard to add it, if one carefully thinks about the requirements and future extensions at every step.
Is that argument to think about memory from the start or to not think about it until absolutely necessary? :)
In addition, extra tools and profiling / static code analysis has helped me to write better code.
wxWidgets is the same - when a wxWindow gets deleted, it will remove its children in the view hierarchy, which are likely not children of it as data members. I think a lot of this comes from the fact that underneath native pointers are being used (eg handles on Windows) and widgets/controls are not necessarily going to be data members of a parent. It also means you can reparent widgets/controls to other GUI items at runtime. I do this (for example) when switching notebook pages to rehouse a GUI element in order to save RAM; I just assign it new data to display and it means less RAM usage and creation time for each notebook page.
For your own data controls, we can still ensure that they clean themselves up alright and safely, just use pointers for the GUI items. In some sense, the GUI toolkit takes care of deleting its own members, eg. wxWidgets will delete all of a parent wxWindow children at teardown, so it isn't a leak situation.
Since I was looking at it for hobby/educational projects, I basically decided to not do something with a GUI. I think I had also looked at wxWidgets and FLTK, hoping that one of them was cleaner and more modern; FLTK managed the "cleaner" part, at least. The last time I looked at C++ GUI libraries it seemed like JUCE was the most promising, but it also had a bit of weirdness (GPL license without linking exception, heavy bias toward audio applications).
CMS GC doesn't choke your all for 100ms
In practice, the GC Apple was using prior to ARC could collect cycles, but it consumed a nontrivial number of CPU cycles compared to ARC.
Apple's tools people seem to be unconvinced that GC is efficient enough for production use on slow CPUs.
It makes sense for Swift because of Swift's unique constraints—interoperability with an atomic-reference-counted world (Objective-C) and, less importantly, mobile UIs putting demands on latency. But most applications would be better off with a tracing GC.
I know this to be true for ubiquitously heap allocated systems, but I'm curious if this has been studied extensively for stack+heap reference counted systems, where only a relatively small number of items are heap allocated? Seems like reference counting would gain in this kind of environment, as the stack effectively gives you many of the advantages generational gcs enjoy from their nursery, and you don't have to worry about write boundaries, tracing and the like.
There's still plenty of cases where you absolutely do care: low level firmware, kernels, absolute performance, and high reliability. It's more a case of not worrying about it when there's much bigger risks in a project than having loose guarantees about allocation performance or fragmentation.
Try making your own particle engine and a simple game running at 60fps. I guess "avoiding GC" is only a very small subset of memory management, but if you're a noob like me, it surely will open your eyes about some things you would never have guessed could cause problems.
> How low should we go? How useful is it to know about C-style memory management when you're a front-end JavaScript Developer?
There are things in between as well (such as RAII). That's about using it. Usefulness of knowing it is another thing. Someone will have to implement things on the lower level either way. I.e. garbage collector itself is going to manage memory after all.
2. Java and C# JITs perform many more optimizations than Go 6g/8g do.
3. Objective-C message passing is a lot slower than method calls in Java and C#.
I think it's due to LuaJIT being one of the very few tracing JITs that are still in development, since they are notoriously difficult to write and maintain.
If I care about performance I'd check the generated assembly and modify accordingly till I get decent result.
Baring that: do you have any reputable source claiming LuaJIT outperforms C#/java/C?
That would imply a horrid baseline in order to get it that fast. Merely I don't believe a tracing JIT would have a great impact in a real world scenario.
PHP with its shared nothing design makes it very scalable, the 2016 implementations are very fast and consume little memory. LuaJIT see other post. Object-C (there is C underneath, read https://en.wikipedia.org/wiki/Object-C ). Check out Julia, it's great. You can write pretty low level code in modern languages like Go, Rust without loosing everything (don't come with Sing# and Java related misc things, yada yada). Also my point was that these languages consume less memory. I have yet to see a Java or dotNet application that doesn't consume too much memory for my feeling (not to mention the GC idioms). When I read that people put monolithic Java enterprise apps in Docker containers or wonder why the new Startmenu crashes that often, it reminds me why I care about computer resources. After all we are still stuck with 3.8GHz single core performance, same as 2004! (memory bandwith,etc got faster, I know) Give me a 10 GHz single core CPU and we can start talking.
Most of the RAM of my application is used by the app and the goddamn UI and its fancy xxxhdpi and xxhdpi and xhpdi and hpdi and mdpi images, and because Android doesn’t support svg everywhere yet, we have to package tons of ressource twice and thrice and five times.
In the Windows world especially that's not even slightly true.
Run "Process Explorer", Options -> Color Selection -> check ".Net Processes" (Yellow).
In all consumer Windows up to Windows 7 no dotNet processes run (by default) at all. [Longhorn was meant to bring a dotNet shell (explorer, etc) but failed big time.] I have lots of common applications and all are coded in C/C++ with the exception of IDEs (VS/Eclipse/IDEA) which eat memory like no tomorrow. I couldn't care less about the brave Win8/10 world.
Still supports C++ ;)
So XP and Vista then.
Most of the software that comes pre-installed on Windows machines (by the OEM) is .NET. That fact by itself means the amount of software written in .NET is far far far above 1%.
> In Windows world you are very wrong.
So you're saying after the shedloads of money, effort and development Microsoft has pumped into .NET it hasn't even amounted to 1% of applications running it? That's preposterous.
DotNet and Java are very big in enterprise software and server software. Java is also big in open source thanks to Apache foundation (if you need a full text search engine, Lucene is still the only feature complete in town; and all the projects based around Lucene like Hadoop, Solr, etc) There was an C++ port of Lucene (CLucene) that was a lot faster and used little memory and compatible with Lucene file format. Though no native fork of Lucene keep up with the development speed of Apache Java Lucene.
All non-technical people, so 90% of users. I get where you're coming from (and you have valid points about speed/memory usage) but to say <1% of software on Windows is .NET is frankly wrong. Sure it's not 100% and might not even be 50%, but it's way way way above 1%.
I remember a Fry's newspaper ad that said something like: "$50 1GB DIMM, $0 after in-store discount. (Limit 2 per customer.)"
Kind of similar to the old (and possibly apocryphal) Sun Microsystems legend: Frustrated that too many domestic programmers were spoiled with ample CPU speeds, they hired a bunch of Russian programmers who dealt with the export restrictions on computing gear and thus had learned to write tight, elegant code.
I guess the answer is obvious: it is nice, but might not worth the effort. 1.2 million requests/s is pretty good, but in reality, you probably would still need 2-3 other hosts as a failsafe backups, let alone there are other factors, like memory constraint, in which case, less optimized solution might be just acceptable because it is sufficient. So under common usage cases, the need for scalability outweighs the urge to squash the maximum performance from a single machine.
That kind of scalability difference does still matter.
My parents told me to "Clear your plate." That was because they lived through the depression. I remember writing code in assembler and 64k and feel the say way now. Programs I wrote in that space now take 64m and I think "What a waist".
b) The quote isn't talking about data structuring/memory layout, but rather micro-optimizations of instruction sequences (removing goto statements, in the case of the original article)
At least in image processing and robotics, from my experience.
It's very annoying when you have optimized all obvious hotspots 10x only to get 3x speedup because the remaining 30% code consistently abuses « simple ways », e.g. return big objects by value, use std::vector instead of proper Eigen arrays, constantly serialize objects everywhere, etc.