You Got Your Web Browser in my Compiler
randomascii.wordpress.com
randomascii.wordpress.com
> Without access to a lot of source code I can’t tell exactly what is going on, but...
> Until Microsoft fixes their compiler to not load their web browser it seems impossible to avoid this problem when doing lots of parallel builds.
In my past life developing applications and services on Windows, this scenario was too common — something is broken, you can't quite tell why because you don't have the source, and you can't do anything but wait for it to be fixed or find some hacky workarounds (which is tougher to do without the source).
This has been the real win on switching to an open source stack for me: you can take your tools apart, look inside them, and fix any issues. I know Microsoft has come further in this in the past few years, but it still sucks to be stuck in that kind of situation.
https://github.com/sparklemotion/nokogiri/pull/554
Gave me the warm fuzzies, even though I broke everything temporarily :)
Granted, MongoDB is not meant to be used for small databases (the name comes from hu-mongo-us), but the app I was using was written against Mongo and changing that would have been a whole lot harder than modifying some constants.
Changes here, BTW: https://github.com/kentonv/mongo/commit/14f391a000134e5d9d65...
Fixing bugs in GCC for an unitiated person is probably super hard.
It's not like you have to read all 7 million lines of code. The author of the article already has a stack trace; all he'd need to do is skim the functions along that trace. In this case it sounds highly likely that the fix is to comment out some code somewhere that is initializing a resource that isn't actually needed.
Frankly it sounds pretty easy.
Until recently, I had thought that the only real tools out there to do this kind of analysis lived only on Solaris (DTrace). (As I recall, SystemTap gets part of the way, but does not reach far enough into userspace to really get a good view of what's going on -- is that still true?) It would be very interesting to augment Windows's performance tracing tools with some of the other things that DTrace has to offer -- pervasive low-overhead trace scripting seems like it would have made some of the other subsequent analyses that Bruce was working on much easier.
FYI, there are ports of dtrace to systems other than Solaris.
There's no reason why any file parsing library would end up fetching remote data without being explicitly asked to do so. Actually, it shouldn't even be the library's concern to fetch those files, an XML library has no business with networking. It's a security concern and a maintenance hell.
It's not the API, it is the part of many of the specs, as it was thought to be a good idea once, specifically many standardizations involved additional definitions located at the http servers, example:
http://msdn.microsoft.com/en-us/library/aa468557.aspx
Which was "poised to play a central role in the future of XML processing, especially in Web services where it serves as one of the fundamental pillars that higher levels of abstraction are built upon."
So you had to implement it to be "conforming," and then to avoid overheads as "optimizations." Ironically, "/optimize" feature isn't optimized.
The reason this it's not discovered earlier is that the Visual Studios which contain that option were priced more thousands of dollars (I don't know the what the currently cheapest version containing "/optimize" is -- anybody knows?).
I would agree. But the XML spec. authors clearly disagreed with both of us:
External Entities:
XML 1.0: http://www.w3.org/TR/2008/REC-xml-20081126/#sec-external-ent
XML 1.1: http://www.w3.org/TR/2006/REC-xml11-20060816/#sec-external-e...
This is "XML speak" for an "include" statement to something like cpp, with the exception that this "include" could end up performing remote network fetches to acquire that which is being included.
So, technically, to be a proper, standards compliant XML parser, the parser has to at least submit requests to "fetch" these entities to the higher level code using the library, and let that code decide what to do about the "includes".
As to why Microsoft's implementation is the way it is, absent a Raymond Chen blog post explaining the why, we can only guess.
<?xml version="1.0" encoding="UTF-8"?>
<office:document-content
(...)
xmlns:xlink="http://www.w3.org/1999/xlink"
xmlns:dc="http://purl.org/dc/elements/1.1/"
xmlns:math="http://www.w3.org/1998/Math/MathML"
xmlns:ooo="http://openoffice.org/2004/office"
xmlns:ooow="http://openoffice.org/2004/writer"
xmlns:oooc="http://openoffice.org/2004/calc"
xmlns:dom="http://www.w3.org/2001/xml-events"
xmlns:xforms="http://www.w3.org/2002/xforms"
xmlns:xsd="http://www.w3.org/2001/XMLSchema"
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
xmlns:rpt="http://openoffice.org/2005/report"
xmlns:xhtml="http://www.w3.org/1999/xhtml"
xmlns:grddl="http://www.w3.org/2003/g/data-view#"
xmlns:officeooo="http://openoffice.org/2009/office"
xmlns:tableooo="http://openoffice.org/2009/table"
xmlns:drawooo="http://openoffice.org/2010/draw"
It's not Microsoft specific bug that ate some brains.https://web.archive.org/web/20060613090946/http://textuality...
"Given that namespaces have definitive material, and that such definitive material is typically available on the Web, and that namespace names may be "http:"-class URIs, it is a grievous waste of potential if it is not possible to use the namespace name in retrieving the definitive material."
And in order to do all the processing and transformations popular at the time somewhere there should be the copies of the documents specified with the URI's. Bruce detected some loads from some documents stored in the DLLs, locally.
Two favourites: 1) App never started because it couldn't access the internet to fetch a DTD/XSD. 2) Sun/Oracle removed XSDs and the app refused to start.
Also you need to get using the precompiled headers if you can -- sometimes this requires structuring your projects right to avoid lots of little projects.
Even with large projects I usually have really low compilation times for C++.
Here is a blog post I wrote about the right way to do this:
The linked article demonstrates some ignorance as well by noting that command-line MSBuild builds are slower than from Visual Studio. I'd hazard a guess and say that they didn't find the /m option.
[0] Well, it used to be. I assume they've moved on to more demanding tests now, since hardware has far outstripped kernel build complexity.
The issue here alone is that he was trying to compile 12 projects simultaneously each with every logical processor he has. There might be some slack in the system, but is each project 92% wait time?
How do you think threads were implemented by the OS back when we all had one processor with one core?
When you kickoff a unlimited parallel build with `make -j` you get one gcc process per .c file.
As you can imagine this over subscribes the CPU by little bit.
Maybe this isn't exactly a great argument since I know Linus has been known to take the side of not ever breaking user mode wrt kernel changes (probably rightfully so) but it is possible to write buggy code, including hangs, that runs on Linux. That would be the analogous scenario.
Strictly speaking, the direct cause of OP's problem appears to be outside the kernel, but it is system-level. Microsoft has created a stack in which low-priority tasks can trivially destroy the responsiveness of interactive tasks unrelated to the build process.
I looked at your blog post and it recommends using /MP and setting num. parallel project builds to num. cores.
In other words, you recommend exactly what I was doing!
We do use precompiled header files. And we get great build times. It's only with /analyze builds that the system locks up.
Yes, 144 parallel compiles is excessive, but as your article says, 12-way parallel compile and 12-way parallel project builds are both necessary to expose all possible parallelism. Otherwise there will be times when many CPUs are idle. Until VS gets a global scheduler to avoid over subscribing the CPUs there is no way to get full parallelism without sometimes over subscribing.
I'm surprised that you find SSDs that helpful. With enough RAM they should offer only modest improvements to compilation, because source files and header files get cached.
(ie that MSXML is calling URLMon, with unpredictable results). There's a comment at the bottom of that thread that suggests how to fix the bug if it's in your own code, but not if it's in Visual Studio...
edited to add: related bugs crop up in many programs that use xml parsers, not just MS. Some combination of disabling validation, loading of external entities, or adding an entity resolver is usually needed (see eg these options for Xerces http://xerces.apache.org/xerces-j/features.html). The symptom there was usually errors in production that didn't occur on the developer's box because in production the network connection to grab the schema didn't work...or worse, a SOAP service that will fetch external entities in messages.
One solution would be to load the resource manually using Win32 APIs, convert to BSTR, and then calling http://msdn.microsoft.com/en-us/library/ms754585%28v=vs.85%2... instead. And yes don’t forget to disable loading of external entities too.
[1] N. Zakas: High Performance Javascript, p. 35 [2] http://en.wikipedia.org/wiki/Trident_(layout_engine)
From your report this is 'use a builtin XML parser' and that xml parser has a dependency on IE?
But this... this takes the dependency tangle to whole new levels of comedy.
I can understand your argument in the general case, certainly, but it doesn't apply very well to COM.
(Edit: seems to have been taken away at least as of 2.10)
I wonder if the behavior in the article also applies to C# compiles?
C# compilation is seriously fast compared to C++ and the likes (due to its proper staticly typed nature) and this allows the compiler to assert certain things quickly and not waste any more time on this.
Given how .NET developers loves separating things into tons of DLLs and libraries (and how .NET makes this easy), I would be surprised if this part of the compilation wasn't as optimized as the rest of the process.
I'm not sure what you mean by this - C++ compilation isn't slow because of a lack of static typing :)
Here is some actual info on the subject: http://www.drdobbs.com/cpp/c-compilation-speed/228701711
Go is fast to compile as well and statically typed
See the post on the C++ preprocessor that was on HN some days ago as well
"The core tenant of C++ has been to move work from runtime to to compile time."
Both RAII (avoiding use of new & delete) and generics are strengths of C++ but also require heavy compile time work.
RAII needs decent analysis and optimization or your program will thrash copying struts by value. Generics and templates force your compiler to run a second compilation in your frontend.
This on top of C's copy/paste based include system and macros gives you what should be one of the slowest to compile procedural languages designable.
That's what I thought anyway?
C++ compilation is sort of screwy in that to compile a new C++ file the compiler actually has re-compile all the definitions of the libraries you are using. It is insanely inefficient compared to just using explicitly defined and immutable API interfaces on Assemblies/Modules.
Right now the closest thing is FxCop or a third-party tool.
Doctor, doctor, it hurts when I do this!