Compiler bug? Linker bug? Windows Kernel bug
randomascii.wordpress.com
randomascii.wordpress.com
But on the other hand reporting feature suggestions to Microsoft is much more pleasant than anything open source has e.g. https://office.uservoice.com/forums/285186-general
In Open Source world, because by the nature of it you're non-paying, asking for features either gets back "submit your pull request" or "go away." Which makes sense, those who contribute the most time get the most say, and that would be fine if those that contribute the most time had the same goals and taste as those that didn't (see the failure of Linux Desktop as an example).
I guess by now Emoji fill that role to some extent. Including the overused part ;-)
None for the Linux kernel, though, but it does have >20,000 forks on GitHub: https://github.com/torvalds/linux/network/members
(disclaimer: I've been running Linux and other open-source OSes for over two decades. I like OSS but it's not perfect and staffing for projects is an under-solved problems.)
Oh, and if you have more than about 512 bytes of local variables in a stack frame and try to use CE remote debugging, that stops working as well.
Issue tracking at Microsoft (and presumably all large tech companies) includes things like:
* localization errors
* typos in user-visible text
* suggestions for rewording user-visible text
* issues related to icons and other graphics
* issues related to sounds that ship with the product
* feature requests
* feature enhancements
* tracking tickets for work to be done or work in progress
* tickets for unverified, incorrect, and unreproducable bugs
So a. 65,000 tickets, not 65,000 defects. b. "65,000 tickets" is meaningless without knowing more about them.
The key is to hit the milestones and get the product working as much as possible.
Enterprise Software Motto: "Sometimes good enough is good enough". Amen.
https://github.com/chromium/chromium/commit/052a09014b2018e2...
https://github.com/llvm-mirror/llvm/commit/8b2a3e8b203a332bb...
I imagine that might be reason enough for them to treat it as a low-priority.
I had the impression that poor documentation forced a lot of Windows programmers to understand Microsoft's API's via experimentation.
Not always. Various standard library functions can cause UB, and not just things like C's div function. In C++, calling front() on an empty std::vector<T> causes UB, for instance. http://www.cplusplus.com/reference/vector/vector/front/
I agree that from a system perspective, it doesn't look like good behaviour.
std::vector et al. are parts of the C++ language. "Language" != pure syntax.
For what it's worth, I read that POSIX does have UB in the spec (as in, distinct from how the compiler is allowed to break things in terms of language-level UB), at least in some ghastly signal handling situations. I wasn't able to find at a glance whether Windows does too, or if it categorically doesn't.
> According to POSIX, the behavior of a process is undefined after it ignores a SIGFPE, SIGILL, or SIGSEGV signal that was not generated by kill(2) or raise(3).
Are you familiar with Windows or just assuming all OSes works like *nix systems? That's definitely not at all close to how Windows file systems behave.
It is implementation defined on the UNIX OS variant and mounted file system.
The beauty of POSIX is how it appears to be portable, while leaving quite a few things being implementation defined.
> The mmap() function adds an extra reference to the file associated with the file descriptor fildes which is not removed by a subsequent close() on that file descriptor. This reference is removed when there are no more mappings to the file.
I'm sure there have been unix variants that got it wrong and might not even have been in posix from the beginning (it is there since at least SUSv2); still I think this is traditional unix behavior.
This is the cause of a lot of woes on windows, such as random access denied errors when attempting to delete stuff, and what forces a reboot when you update nearly anything on windows.
One of my favourite things related to this is that the Windows XP explorer had a bug where clicking on a file would generate a dynamic preview of an audio/video file in the sidebar, which would lock the file, rendering it impossible to delete because you had to click on it first....
i haven't actually tested this with a real library. this is what i've assumed happens from playing around with mmap(PROT_READ|PROT_EXEC, MAP_PRIVATE|MAP_DENYWRITE).
O_TRUNC for the reasons you stated and unlink because it causes a race between the unlink and the copy when a new process that tries to use the library will either fail because it doesn't exist or get a copy of the library that isn't completely written yet.
The correct way to do it is to write the new file to the same filesystem at a different path and then, once it's completely written, move it to the intended path. Move (i.e. rename(2)) is atomic on sane filesystems so anything that opens the file will get one complete copy or the other.
HANDLE file = CreateFile(TEXT("\\\\.\\C:\\mmap.bin"), GENERIC_ALL | SYNCHRONIZE, FILE_SHARE_READ | FILE_SHARE_WRITE | FILE_SHARE_DELETE, NULL, OPEN_ALWAYS, 0 /*can also try: FILE_FLAG_DELETE_ON_CLOSE*/, NULL);
if (file != INVALID_HANDLE_VALUE)
{
HANDLE mapping = CreateFileMapping(file, NULL, PAGE_EXECUTE_READWRITE, 0, 1, NULL);
void *ptr = MapViewOfFile(mapping, FILE_MAP_ALL_ACCESS, 0, 0, 0);
CloseHandle(file);
*static_cast<unsigned char *>(ptr) += 1;
}
with /NXCOMPAT and even /DYNAMICBASE and with the embedded manifest <?xml version="1.0" encoding="UTF-8" standalone="yes"?>
<assembly manifestVersion="1.0" xmlns="urn:schemas-microsoft-com:asm.v1" xmlns:asmv3="urn:schemas-microsoft-com:asm.v3">
<assemblyIdentity type="win32" name="Microsoft.Windows.Foo" version="1.0.0.0"></assemblyIdentity>
<trustInfo xmlns="urn:schemas-microsoft-com:asm.v3">
<security>
<requestedPrivileges>
<requestedExecutionLevel level="asInvoker" uiAccess="false"></requestedExecutionLevel>
</requestedPrivileges>
</security>
</trustInfo>
<compatibility xmlns="urn:schemas-microsoft-com:compatibility.v1">
<application>
<supportedOS Id="{8e0f7a12-bfb3-4fe8-b9a5-48fd50a15a9a}"></supportedOS>
</application>
</compatibility>
</assembly>
but it seems to still modify the file after closure even after the CloseHandle call.Edit (in response to comments): Actually, I don't think it has anything to do with PE files being optimized in some special fashion, because they say loading via LoadLibrary(Ex) can cause this issue too, and those are very much user-mode constructs, and hence can't mess with the caching/paging behavior (which are lower, in the kernel). AFAIK the only kernel-mode component to them is to just create a file mapping (NtCreateSection) and after that all they do is a bunch of user-mode juggling (fixing up relocations, loading other libraries that are depended on, calling entrypoints, etc.).
So my speculation is that this bug is a race condition between the cleanup of a dirty memory mapped section of a file and the reading of the same file contents in a separate mapping.
Memory mapped I/O is very much normal I/O. Normal, as in "common" or "usual".
> Memory mapped I/O is very much normal I/O. Normal, as in "common" or "usual".
You know well what I meant.
So while I did know what you meant, I found your comment misleading to someone who isn't familiar with the subject.
Too many programmers misunderstand and are scared of memory mapping as it is.
The linkers are getting a compatibility hack (FlushFileBuffers call) but this is not supposed to be necessary, by the file-system contract.
Been there since Windows 98, at least.
From the TFA:
if a program writes a PE file (EXE or DLL) using memory mapped file I/O and
if that program is then immediately executed (or loaded with LoadLibrary or LoadLibraryEx), and
if the system is under very heavy disk I/O load
then a necessary file-buffer flush may fail> if a program writes a PE file (EXE or DLL) using memory mapped file I/O and
Write
> if that program is then immediately executed (or loaded
Read immediately
> then a necessary file-buffer flush may fail
You get bad data in memory.
The parent made it sounds like you could just write a loop with for(;;) { write_file("x"); read_file("x"); } and expect discrepancies. That's not the case and it's wilfully misleading to suggest it is.
* The files being written probably won't be using memory mapped IO (I'd expect them to use the equivalent of fwrite instead) as the installer won't need to make complex changes to the binaries in the same way that a compiler/linker would
* There will likely be a significant delay between writing the file and running it compared with a build process that uses its own build output as part of the build (even in the build process case, the bug went away for a year, possibly due to an additional delay between the write and execute steps)
* Once the installer has run, the system will probably not be under heavy load (the author was running a highly parallel build on a 24-core machine, and this bug was still only hit in 3% of builds)
Really the earlier post was just misleading, as it implies the bug is much easier to hit than it really is.
That’s rare (usually there you’d read it back in and execute in your process space), but it’d be at least a second realistic scenario.
Not only compilers. Decompressors (.zip, etc.), installers and auto-updaters. There are more use cases for writing out executables than just compilers.
Although these cases probably write sequentially and don't need to memory map the file and update sections in a random access fashion.
Here the implication is that there is a fundamental design flaw in the OS that may cause widespread issues, when really this is a bug in quite a narrow edge case: it may happen when building colossal projects repeatedly on very high spec computers, the article suggests that the team at Microsoft investigating the bug couldn't reproduce it since they couldn't build Chrome as quickly as the author could.
Coherency is covered by the WinAPI documentation (check the remarks section). Always RTFM.
https://msdn.microsoft.com/en-us/library/windows/desktop/aa3...
> [ex-coworkers at microsoft] confirmed that my fix should mitigate the bug (I’d already noted that it had allowed ~600 clean builds in a row), and promised to create a proper fix in Windows.
would have saved you the bother of nerd-sniping.
The coherency notes in the article you link are irrelevant because they're about concurrent access. In TFA, the file is created (through mapping), closed, then executed. Sequentially. Also no ReadFile or WriteFile involved (in fact though I don't know how Windows implements it one would expect PE loading to use memory mappings, and thus be covered under coherence guarantees)
The coherency issues are your "tread lightly" warning. The docs regarding FlushViewOfFile expand on this specific issue:
Flushing a range of a mapped view initiates writing of dirty pages within that range to the disk. Dirty pages are those whose contents have changed since the file view was mapped. The FlushViewOfFile function does not flush the file metadata, and it does not wait to return until the changes are flushed from the underlying hardware disk cache and physically written to disk. To flush all the dirty pages plus the metadata for the file and ensure that they are physically written to disk, call FlushViewOfFile and then call the FlushFileBuffers function.
So, after carefully reading this documentation, we can clearly conclude OP's pull-request is actually missing a call to FlushViewOfFile prior to calling FlushFileBuffers.
The whole point of a cache is that if A writes something, and then B reads it after A finishes writing it, B should see the results of A's modification. Whether it has been flushed to the physical media is irrelevant, because if it hasn't the cache manager can still serve the contents directly from the cache. And a call to Flush is not necessary for that.
So, the OP's change is not missing anything. As the article explains, it is in fact the Windows kernel that is behaving incorrectly (which was also confirmed by people who work on the windows kernel).