OS X LevelDB Corruption Bounty: 10.00 BTC + 200.2 LTC
bitcointalk.org
bitcointalk.org
But the donors could have backed out and did not, and there are other similarly-sized bounties in the bitcoin community (I currently make my living off of community donations as a bitcoin-core developer).
Anyway, I think the attention is well deserved and I hope this contributes to a real fix. The current situation makes me not want to use the official bitcoin client under osx, at all. And this will be the case for a long time now, since I have no idea whether the bug will be correctly fixed.
I wouldn't be shocked if it ultimately turned out to be due to some setting that gurus would never have enabled. :)
https://developer.apple.com/library/mac/documentation/Darwin...
"For applications that require tighter guarantees about the integrity of their data, Mac OS X provides the F_FULLFSYNC fcntl. The F_FULLFSYNC fcntl asks the drive to flush all buffered data to permanent storage. Applications, such as databases, that require a strict ordering of writes should use F_FULLFSYNC to ensure that their data is written in the order they expect. Please see fcntl(2) for more detail."
https://developer.apple.com/library/mac/documentation/Darwin...
"F_FULLFSYNC - Does the same thing as fsync(2) then asks the drive to flush all buffered data to the permanent storage device (arg is ignored). This is currently implemented on HFS, MS-DOS (FAT), and Universal Disk Format (UDF) file systems. The operation may take quite a while to complete. Certain FireWire drives have also been known to ignore the request to flush their buffered data."
OS X has aggressive file buffering in memory, and it's getting more aggressive all the time. For example, cfprefsd, introduced in 10.8 (https://developer.apple.com/library/mac/releasenotes/DataMan...) made it so that when a system application read a preferences file, it stayed in memory and ignored the disk version, until cfprefsd eventually synced it back to disk. In 10.9, the behavior is much worse to the point that as soon as a pref is in cfprefsd, it's unlikely to leave it until the user logs out / the machine reboots.
In this instance, OS X has, for quite some time, had "defrag on the fly" for files under 20MB in size. On access of the file, it's read into memory and kept there in its entirety until memory pressure from other processes triggers a sync it back to disk. When it comes to writing a small file back to disk, OS X will "get around to it" when it's damned well ready unless you force its hand using the fcntl options above.
Unfortunately, the bit about "This is currently implemented on HFS, MS-DOS (FAT), and Universal Disk Format (UDF) file systems" covers pretty much the range of filesystem types that OS X can natively read+write on - but one that might get past this is ExFAT. I'd be surprised if that was the case, but it is natively supported read+write on OS X and would be something quick and easy to test (set up an ExFAT volume for the database) and possibly verify this is the root cause.
(Additionally, third-party read+write access to filesystems like NTFS via Paragon / Tuxera may be able to confirm this as well.)
More reading material (MySQL has been dealing with this since 2005): http://lists.apple.com/archives/darwin-dev/2005/Feb/msg00072...
https://code.google.com/p/leveldb/issues/attachmentText?id=1...
https://github.com/sipa/bitcoin/commit/b28d8b423bddc860c5858... https://github.com/gmaxwell/bitcoin/commit/e7bad10c12ce9b5d4...
But again - I'd point to the work of other longstanding database projects that are available on OS X as a source of "how we ensured data correctness".
Something a tester experiencing corruption at startup may want to try is using the 'purge' command from the Terminal.
While a restart will indeed trigger caching files back to disk before the pending restart of the system, the 'purge' command will simulate a "cold boot"-like empty disk buffer by dropping existing file caches.
https://developer.apple.com/library/mac/documentation/Darwin...
If using the command solves the problem of the database corruption without a restart, then you're definitely suffering from disk cache. Easy to confirm.
Additionally, since this is an error on boot up of LevelDB, this is probably in regards to file reading - especially since on second bootup after a restart, no error is detected. F_FULLSYNC is for ensuring that a particular file's changes are 'fully written to disk' ...
... but the cache works both ways. A program could also end up reading the disk cache (which sounds like what's happening here), unless you used F_NOCACHE or F_GLOBAL_NOCACHE. Mind you, these don't prevent the accessing of files already in disk cache - they prevent a file from getting into the disk cache in the first place.
http://lists.apple.com/archives/darwin-dev/2009/Oct/msg00165...
(Disclaimer: I can't think of a situation where the disk cache would get out of sync with the on-disk version of the file if F_FULLSYNC is used when accessed by a subsequent launch of a program except for in the case of faulty RAM on a machine which flipped bits. Your average file operation done by your average application in OS X isn't performing a checksum of data and generally wouldn't notice a single bit flip. It would be interesting to see which of these machines are using ECC RAM.)
The fsync() function is intended to force a physical write of data from the buffer cache, and to assure that after a system crash or other failure that all data up to the time of the fsync() call is recorded on the disk. Since the concepts of "buffer cache", "system crash", "physical write", and "non-volatile storage" are not defined here, the wording has to be more abstract.
From what I understand, that behaviour is in spec (for me, borderline, at best, but I don't make that spec) according to http://pubs.opengroup.org/onlinepubs/009695399/functions/fsy... ("physical write from the buffer cache", not "physical write to the disk") and, AFAIK, is what others do, too (http://ridiculousfish.com/blog/posts/mystery.html)
Edit: http://lists.apple.com/archives/darwin-dev/2005/Feb/msg00072..., referenced from that ridiculous fish post, gives more background info.
As a proprietor of such a system you also end up with a whole category of new bad behaviors, as people try to maximize their hourly pay by finding ways to game the system, getting the most payout for the least input.
Wikipedia does have miscellaneous scores and badges, of which I find the badges awarded by community members as recognition of a contribution most useful: https://en.wikipedia.org/wiki/Wikipedia%3ABarnstars
There's also just raw counts of contributions, which everyone takes with a large grain of salt:
https://en.wikipedia.org/wiki/Wikipedia:List_of_Wikipedians_...
https://en.wikipedia.org/wiki/Wikipedia:List_of_Wikipedians_...
You can also click "thank" next to specific edits, which just sends the person a notice that someone appreciated their edit. That I think is useful, but I'm not sure how useful it would really be to keep a score of "number of thanks" or whatever. The point is just to say "hey someone noticed you did a good job here and appreciates it" to give some encouragement, not to keep score of who got thanked more.
And finally many people just collect vanity "hey look at what I've done!" lists, which can be a nice way of reflecting on your contributions and feeling good about them. Many people's User Pages are like that, or you could do it externally like e.g. http://www.gwern.net/Wikipedia%20resume
All it takes is one developer who happens to be an expert in this area to see the bug and say, "I know exactly what is causing this!" and go fix it. What's the saying - "with enough eyes, all bugs are shallow" or something.
I don't know much about LevelDB or the Bitcoin client, but I am currently taking a closer look at OS X first.
[1] https://developer.apple.com/library/mac/releasenotes/General...
[2] https://developer.apple.com/library/mac/releasenotes/macosx/...
When ever I need something written to disk _immediately_ I go straight to the drivers:
if (ioctl(fd, BLKFLSBUF, NULL))
perror("BLKFLSBUF failed");
that should work.