On Stack Overflow's recent battles with the .NET Garbage Collector
samsaffron.com
samsaffron.com
This article highlights why (good) C/C++ have such a hard time moving to C#/.Net. In .Net memory management is, 99.9999999999% of the time, fire and forget, whereas a big chunk of learning and programming C/C++ is about memory management. Most C, and some C++, programmers would have baulked at having a massive object graph in memory - releasing it becomes a massive headache, and that's before you even get into synchronisation issues, and the .Net team have done a really good job of taking this headache away (yes, I realise they 're not the first). However, we're now at the point where there are a lot of C# programmers who've never heard of a linker (by default, csc compiles and links in one go), have little concept of static vs dynamic linking, and, more importantly, don't have the first clue about memory management. This manifests itself in things like not using IDisposable properly/consistently, using the string.+ operator rather than using StringBuilder/StringWriter, and developers have no idea about the difference between class and struct, especially when it comes to parameter handling. Even books like (the excellent) C# in Depth are pretty light on the memory management side of things, deferring this to books about the CLI and DNR. Whereas, the truth is, if you're going to write a big application, you're going to have some concept of this.
Often we think the magic memory fairy is going to make everything OK, it is easy to fall in this trap even if you understand the underlying concepts, this is clearly not the case at high scale.
Much like marc said, it's absolutely fine for small amounts, but when you're doing a lot of string operations involving a loop, using StringBuilder will be faster and more memory efficient.
This is a good read about it, I highly encourage it: http://www.dotnetperls.com/stringbuilder-1
On the other hand I am not sure if this is a problem in .NET. I know the semantics look bad but no one said the execution model had to match the semantics. If I remember correctly the JVM has the capability to rewrite the code to the faster version, though I am not sure if the CLR does.
String result = "";
for (String str : myStrings) {
result += str + " "; // on each execution, a new StringBuilder is created
}
as opposed to StringBuilder result = new StringBuilder();
for (String str : myStrings) {
result.append(str).append(" ");
}
The JVM can't hoist the generated StringBuilder out of loops to produce the second, faster piece of code. I believe the CLR is the same.But wasn't slashdot the first high volume site sharing their solutions to performance issues?
Seemed like /. was one of the first millions-of-hits-per-day site that was heavily dynamic content. All with Perl and MySql, right? I still don't really understand how they do it, but I remember feeling grateful that they were describing it openly.
http://meta.stackoverflow.com/questions/19244/community-podc...
If your Google-Fu is good.
After reading this article I felt extremely humbled. As a developer doing relatively "simple" work day to day, seeing this level of programming proficiency is very impressive. There isn't a chance I'd be able to figure out a problem such as this. At least certainly not within the timeframe that the SO team fixed it in.
It's articles like this that make me want to continue pushing my limits as a developer.
Don't sell yourself short. You generally don't know what your actual limits are until you really push hard against them or break through them. I've had to hunt down and fix some very unusual problems - with few or no references or no results from google.
Those kind of mind-bending problems actually give a great deal of satisfaction. I'm serious that you never really know your limits until you break them. I think the trick is to keep growing and expanding your abilities and knowledge.
I think that I have gotten to where I am by constantly pushing myself, but reading articles like this just reminds me that I still have a ways to go.
I think this would be a cleaner approach than changing the entire code base.
Hell, you could argue the "root cause" is the use of a garbage-collected environment in the first place. All popular GC implementations have latency issues like this. All of them. If you can't deal with occasional high latencies, you should identify that requirement before choosing Java or C#.
All the solutions provided are just patching around that issue. I don't see anything in the post that looks like a root cause, and the GP's post has the advantage of being much simpler to implement.
What would be more useful to them is to be able to provide either hints to the GC, or to actually have some control over memory management. That is, keep the application the same, but let it influence GC behavior. They're forced to go at it the other way: keep the GC the same, but change the application so what the GC does becomes the right thing.
they modify the code in ways that don't represent the semantics of the problem to work around a limitation of the implementation.
That is always true when you start optimizing for performance. And nick_craver explains why the proposed suggestion may be easier to say, but there are hidden complexities. Personally, once you start introducing external control to your application like that, alarm bells go off in my head.
I can not stress how important it is to have
queryable web logs.
I'd love to read some (actionable) articles on how to set up some of the benchmarking / logging stuff that you guys have. For sure I'll check out .NET Memory Profiler and mini profiler linked in the article, but I'd be grateful if anyone wants to post more comprehensive info on where to start.But it certainly seems like to me that if you want to build a high-performance, scalable piece of software, you don't want to have to bang your head against wall and fight against deficiencies within the language implementation itself in order to get your software working properly.
(NB, languages like Ruby and PHP are lovely, but are over 10x slower than C# and Java.)
That being said, having to rewrite an entire engine because of the underbellies of the CLR seems to be dubious, especially in the case of a development environment that purports to make life easier for the developer.
I have nothing against C#, one of the trading frameworks I use uses it exclusively (Ninjatrader) and it's pretty nice. It's just that I would hate to have invested weeks/months into a software project, using what I imagine would be tried-and-true software techniques, and then have to rework them to get over the particulars of the implementation.
That being said, I guess all languages have similar problems. For example, when I was working on a backend server back in 2001, we found that our product was "leaking" memory on HP-UX, even though none of memory leak detectors found any problems, and the memory leak problem didn't occur in Solaris, AIX and Windows (we had a cross-platform server). It turns out on high-load systems, our app was allocating a bunch of small strings all over the place, and the base memory allocator for HP-UX was causing memory fragmentation over time. By switching to another memory allocator that handled this situation, it fixed the problem.
I guess there aren't any C/C++ books that I've come across that say "don't allocate a millions of small memory blocks because your memory might fragment, even though you free them properly." So maybe I misspoke, I guess to some degree it does manifest itself in all implementations of languages.
Most business app work is string manipulation. If you think Ruby is slow, look at the timings for "sequential" operation in the table on this page (yes, it's my site):
http://roboprogs.com/devel/2009.12.html
If you think that I have badly blundered in my methodology, get the benchmark code, fiddle with it, and run it on your own hardware:
https://github.com/roboprog/mp_bench
FYI: I'm not trying to say that "Ruby is the most awesomest language, like ever, d00dz!!!". I like different tools for different jobs. I'm just tired of seeing people judge implementation speeds based on bit twiddling benchmarks, rather than stuff that at least churns through a large number of strings, if not other object types, and does some I/O.
Interesting comment about the date formatter. I might put in something some time to compare the relative time spent in formatting vs the other concatenation. I suppose I'd be grateful if somebody made a fork with versions showing the split of time between those steps.
Perl: print $local_process_var, " ", &gen_pg_template(), "\n"
Python: print local_process_var, " ", gen_pg_template()
Java: System.out.println( timestamp + " " + genPgTemplate() );
I'm not sure, but I think the Perl/Python code sends the strings as arguments to print that will just output them directly. In Java you will concatenate 3 strings before sending it to "print". This adds some extra CPU usage that the other versions don't have
Still, I could modify the scripting output lines to bring them down to Java's level :-)
At least Ruby has the option to mark some strings as outside of garbage collection -- to make them permanent symbols/atoms.
We need an option in these modern GC languages to select a reference counting option, or perhaps for portions of code to kick it old school and simply ask for some objects to allocated outside the GC heap on the assumption that they will be around for a while and that the programmer will manually throw that subset of objects away if they become unwanted.
TMTOWTDI. (and all ways seem to have their place at times)
In C# you can use the unsafe keyword to get more or less back into the C/C++ world. While your new'd objects are still GC'd, you have the option of just grabbing an allocator and getting unmanaged memory if you feel like it.
It's worth noting that the .NET GC works quite well for typical cases, and only noticeably chokes on this one use case we have and only under load that most websites will never see (and most users don't notice even on SO, but we watch the logs religiously to make sure these things don't become noticeable). There are also some improvements coming in .NET 4.5's GC that may mitigate this, but we're naturally not willing to wait for it.
I worry people may be taking "don't use .NET, its GC sucks" away from this blog post, which really isn't the point (or accurate).
I get the feeling some already came in with that opinion, so they will take away what they want to. I personally got the impression that this is something that would have caused pain in almost any language with a generational GC.
Admission: I'm am reluctant to want to work on Windows, though. But since that is the Stack Exchange environment, C#.NET seems hands-down the best platform for it.
http://www.mono-project.com/FAQ:_ASP.NET
Seems to validate the IIS bit.
Our initial tag engine design was a GC nightmare. The worst possible type memory allocation. Lots of objects with lots of references that lived just enough time to get into generation 2 (which was collected every couple of minutes on our web servers).
Interestingly, the worst possible scenario for the GC is quite common in our environment. We quite often cache information for 1 to 5 minutes.
(Curiously, this is under a header mentioning "abuse of the garbage collector". I'd put it the other way around ...)
I must admit I am a little surprised that tags are stored as a free text field.
Tags are stored both in a posttags table and free text. Fts is faster for many queries.
Another option I was surprised they didn't mention was to run GC.Collect periodically to clear out gen2.
Since nothing is actually getting collected, if you ran GC.Collect periodically, you would just be incurring more lag with no benefit. I think your idea might help more in a different situation, one where you do have large amounts of memory getting reclaimed with each GC sweep.