Sometimes the bug isn't in your code, it's in the CPU
leaf.dragonflybsd.org
leaf.dragonflybsd.org
I work with some very talented developers who, when they try something and it doesn't work, try something else. I am fundamentally incapable of that. If it doesn't work, I MUST KNOW WHY. Even if that requires building a debug version of my entire stack, adding all sorts of traces, and wolf-fence debugging until I have a minimal fail case.
It's a real limitation; if I hit an undebuggable brick wall, I have no ability to attack the problem from a different angle. Luckily, there are few things that are fundamentally undebuggable.
I found I knew of it under a different name: binary search debugging. Git includes built-in support under the bisect command.
> The "Wolf Fence" method of debugging time-sharing programs in higher languages evolved from the "Lions in South Africa" method that I have taught since the vacuum-tube machine language days. It is a quickly converging iteration that serves to catch run-time errors.
http://dl.acm.org/citation.cfm?id=358695 (if you have access)
Anyone know what the "Lions in South Africa" method is? I couldn't find it via Google, it just kept turning up references to the same paper.
http://coreygoldberg.blogspot.com/2008/12/wolf-fence-debuggi...
It stipulates that the state of Alaska has got exactly one wolf, so you build a fence across the middle of the state to find on which side wolf would howl, then subdivide the problem, etc...
I'm assuming in your case the wolf got replace with a lion and Alaska with South Africa.
Second, I would say that over the course of my 10 year career in managing developers, I've heard many, many times that the bug was in the kernel, or in the hardware, or in the complier, or in the other lower level thing the developer had no control over. This has been the correct diagnosis exactly once. If I had to guess, I would say about 5%.
Perhaps we need to incentivize software developers with fear of execution, or something.
There certainly are a lot of hardware bugs in cpus -- it's just that most of them get fixed before anyone outside the cpu company ever sees them.
[1] I work on Visual Studio so I have found compiler/debugger bugs as I am generally using 'in development' bits, but far more often than not the bug turns out to be mine and mine alone :)
But really, I have only ever experienced one bug in a compiler (that I hadn't written) but it was such an odd experience, like the patient having Lupus.
Usually bugs would involve unlikely sequences of operations or operations in unexpected states. But there have been very serious bugs involving wrong math (Intel Pentium) or cache failures leading to complete crashes (AMD Phenom). These two made it to production and were show-stoppers because OS devs could do very little about them (in the Phenom bug, they could, but with a noticeable performance hit). I don't think I've seen any production CISC chip completely free of bugs. OS devs have to do the testing and the circumventing.
I mean... typically x86 chips have DOZENS of documented bugs.
http://en.wikipedia.org/wiki/Pentium_FDIV_bug http://en.wikipedia.org/wiki/AMD_Phenom
Always expect you are doing it wrong. It will so rarely be the case that this expectation is wrong that you can discount it as insignificant.
Want to find bugs in Sun's Java 6 compiler for X64 Linux , use annotations (yeah, I found one in their V30 release last week). Want to find bugs in MS' C++ compiler, write your own templates (this was a few years ago, maybe it's better?). The best programmers push the limit of their tools because they know what's "supposed to happen".
Poor programmers hit something that doesn't work, and just try something else, cause, well they're just trying shit. I would go so far to say that poor programmers, in fact, are unable to find compiler, optimizer, OS, or hardware bugs because, by definition, they probably don't have a firm handle on what's "supposed to happen".
I know these things can and do happen. I've come across one or two of these strange ones before, but too often I've seen people jump to the conclusion that someone / something else was to blame. Without any other real evidence other than that they have exhausted their shallow back of talent.
http://pragprog.com/the-pragmatic-programmer/extracts/tips
However, I did once work with someone who did find a bug in select on SunOS... (this was '89 or so).
The documentation was a completely wrong for several years before I started programming there. Once I realized that the documentation was lying to me, I started methodically examining how things actually worked by writing lots of very simple test programs and documenting the actual behavior.
The others didn't care why it was broken. Most of them were just trying to build cool areas, they weren't really programmers at all. They would just tweak things until it appeared to work or they were frustrated enough to give up.
[1] I'm sure a lot of HN knows this, but to save the non-gamers the hassle of looking it up, a MUD is a type of text-based online game and an MPROG is a script used to control the actions of the characters in the game.
Apparently I wasn't the only one, because within a week all the labs were back on GCC 2.95.
Many distributions rolled back so that the default "stable" compiler matched the one they had to use to build the packages - i.e. common sense.
Once the packages were updated to deal with the gcc 3.x language changes, the compiler and packages started appearing together.
Now, blaming without reproduction on the dev stack is a sign of a lamer. :-)
These legitimate bugs do exist, and some of us have a talent for finding them with annoying frequency.
Less experienced programmers often "want" to find bugs in the compiler/OS/whatever because that way it's not their fault, and they lack the skill to track down difficult problems in their own code. More experienced programmers realize that finding bugs outside of your own code is often a disaster because frequently there's nothing you can do to fix it.
My first programming job was doing VB programming in Access 2 programs that had to run on Windows 3.1. (Yes, this was in the last millennium.) I kept on running into bugs that I could demonstrate were in Access, not in my code. It was very frustrating.
My next job was in Perl. I went several years before I found an actual bug in the language. Which then went unfixed for years because someone might be using it. Despite the fact that in every significant Perl code base that I've seen since, there are real bugs in the code that nobody has noticed which trace back to the bug that I found. Why do you ask whether I am bitter?
So your suggestion failed glaringly for me when I was using VB, but since has worked much better.
Access 95 had the ability to upgrade from Access 2, and that included the ability to migrate from Access Basic to VBA. The tool was not flawless (very little from Microsoft is), but mostly worked pretty well.
I have personally found bugs in Linux (kernel, libc), Oracle, various JVMs, etc, usually cases in which algorithms optimized for "normal" loads became pathological under extreme load. It's much more common than perhaps you'd think.
Compiler bugs are another story entirely. I have found dozens of them (confirmed), and I can find more whenever I feel like it.
(I've found several missed-optimization bugs in gcc, but I found them while working on a project where I examine assembly frequently; I have no idea how I'd go about looking for a compiler bug).
While I don't consider missed optimisations bugs as such, they are easy to find. Simply compile some non-trivial function and look at the output. There's usually something that could be done better, especially if some exotic instruction can be used.
Perhaps you'll give me a little credit :) if I mention that I found missed optimization bugs in extremely trivial functions. One of them involved gcc generating several completely useless stores even at -O3: http://gcc.gnu.org/bugzilla/show_bug.cgi?id=44194
The link is nice for everyones education, but I, for one, would appreciate a little less condescension.
And, yes, I would expect everyone in this community to recognize who these people are, and roughly what their contributions have been.
One of these guys is not like the others.
http://daringfireball.net/projects/markdown/
Not that that puts him on the same level as the others, but it is true (as the GP suggested) that his name is well-known here.
> And, yes, I would expect everyone in this community to recognize who these people are, and roughly what their contributions have been.
I'm sorry but I (and everyone else) don't owe no one any such things. This is excepting too much from people. Yes, not knowing who is Linus, or Bill Gates or Woz would be strange, but its ridiculous to say that Madd Dillon is as well known. I can't (and don't want to) know every significant linux/bsd contributor. I have enough information filling up my limited brains as it is.
Within the Hacker community, anybody who's been "in the industry from the late 80s" should have awareness of around a couple hundred or so major figures, most of whom my mother would not know. Pundits like John Gruber (MG Siegler, Sarah Lacy, Michael Arrington) are perhaps better known in the HN community than those outside of it - but I would expect everyone who has been around any hacker community for a couple decades would recognize Names like Dennis Ritchie, Guido Van Rossum, Richard Stevens, Larry Wall, Tannenbaum, Bill Joy, etc...
In addition to this group of people, People like Matt Dillon should be on your peripheral Radar, even if you don't follow BSD that closely. I've never run FreeBSD/DragonFly BSD, and even I recognized his name.
I'm not suggesting you memorize every last hacker/pundit of any note, but there is just a core _canon_ of people that we reasonably should be expected to know - it's knowledge like this that ties the broader hacker communities together.
if I google Matt Dillon, I only get things about an American actor. So I can't really think of a any reason that I would have heard of him, since this is the first thing that I have read (afaik) that has mentioned him at all. If he was writing a popular technical blog on top of his work then I'm sure I would have.
Of course that is not to disparage his work at all, there are many people who do extremely impressive, valuable things but receive little fame. Even amongst educated people you would probably struggle to find someone who could tell you who invented the combustion engine (Étienne Lenoir) but I'm sure many could name the members of the latest fad pop music group.
It's just that some people are pushing into the spotlight more than others, or more likely they get there on purpose. Some people are happy doing good work in the background and only want recognition amongst a small group of peers.
- Dennis Ritchie of course, and Kernighan too (though I had to google for how to write it) - Gruber and Arrington - yes - Siegler and Lacy - no - Van Rossum, Wall - of course - Tannenbaum - yes, - Stevens - no idea - Joy - rings a bell, but still no idea.
you see, I'm not really into remembering names, especially of the people I'm not in close contact with. I only remember then if they are pounded on me for a while ;)
That said - what is it about the hardware manufacturers that makes them relatively immune to this sort of thing? Is it formal verification and rigid engineering process? Is it that they spend so much money developing these things that they better do them right, god dammit?
Sometimes I think that the whole industry would be much better off if everyone up the stack was held to these kinds of standards. If that were the case though, where would we be? We'd have rock solid systems, but how sophisticated would they be? Would UNIX exist? What about (a more bulletproof and less feature complete) Java?
All of the above- with the minor correction that it's not about money spent developing per se. Producing silicon masks is obscenely expensive, so catching a bug before tape-out vs. after tape-out can be a difference of hundreds of thousands of dollars. So, think of it as "you better get it right the first time, god dammit"
(nitpick: "per se" http://en.wiktionary.org/wiki/per_se)
lol
Seriously, that was very cool. Thanks!
IMO that's an advantage that easily overcomes any instability that new software brings.
That said there was an interesting article/interview (cant find it sorry maybe someone else can) with one of the creators of hotmail. Off the top of my head he said that because he came from the hardware side when creating the hotmail software it never broke due to the processes and practices he followed. Perhaps there is a middle ground that we can take and get benefits all around.
Agreed that the rapid development cycles and malleability of the product is part of what's made software great - but my feeling is that maybe we've let that slip a little too much - leading to the bloated, slow, buggy software that everyone runs all the time.
In the hardware world at the chip level, the environment is fairly rigid. A designer has a pretty good idea of what the environment is going to be for their chip. Linux is designed to operate in all sorts of environments: multiple cpu architectures, multiple versions of those architectures, the myriad peripherals, and all the other software that is going to run on linux.
That's not really the case with hardware at the chip level. We know what chip we are going to be talking to, or what family of chips. And most of those are designed in house, so there can be a lot of give-and-take going on.
It's just not a lot of combinations (compared to software, in my opinion.)
An exception would be memory, which is usually made by a third party. And it's also where you find incompatibilities... some manufacturer's ram doesn't work in macbook pros, etc. That's a bug.
And because the hardware can take a long time to iterate through design->implementation->verification->fabrication->prototype verification, there is a strong force driving designs toward being the simplest possible implementation that accomplishes the goals.
So you have these very well defined blocks of functionality that get pieced together and form a chip. The interfaces between blocks are very rigidly defined, and the amount of functionality in a given block is usually pretty limited.
At the level I work at, design is done in VHDL or Verilog (and now SystemVerilog is becoming a viable option). So that's a limited space. That could possibly be the biggest contributor to "less bugs" in hardware (if that is a true supposition). I get flustered with all the different software programming languages.
A big disadvantage of hardware design is that testing is first done in event based simulation. The tools available for this are pretty awesome, but it's still extremely slow. My last design would take about 100 hours to simulate about 300 ms of hardware time. That's a large part of why the iteration times are so long.
Now, with AMD and Intel, and designing these ridiculously highspeed CPUs, it's a whole different ball of wax. I expect a lot of their design is transistor level, full custom. Maybe they prototype in a higher level language, but you aren't going to 3GHz doing standard cell designs in VHDL.
Apparently Bulldozer, AMD's latest chip, is their first to start using automated design tools; one ex engineer claims that it resulted in 20% bigger and 20% slower designs:
http://www.xbitlabs.com/news/cpu/display/20111013232215_Ex_A...
One of my comp sci professors found a bug in an Intel chip and got his name in the errata. I think that gives you +100 to nerd credibility :)
The errata documentation for AMD Family 10h Processors (Athlon, Opteron, Phenom, etc.) is here: http://support.amd.com/us/Processor_TechDocs/41322_10h_Rev_G...
The errata for AMD Family 12h Processors (A-Series APU, etc.): http://support.amd.com/us/Processor_TechDocs/44739_12h_Rev_G...
I found this out when an AMD engineer confirmed an AMD CPU bug for me: http://stackoverflow.com/questions/7004728/is-this-should-no...
It would be interesting if he has accidentally triggered a backdoor, such as mentioned in this post.
http://theinvisiblethings.blogspot.com.au/2009/03/trusting-h...
http://thread.gmane.org/gmane.os.dragonfly-bsd.kernel/14471
(Check out in particular the section "EFFORTS AT FINDING A KERNEL BUG THAT WASN'T A KERNEL BUG")
If all else fails, issue a product recall or downplay the bug's severity.
On a more serious note, I wonder what auto manufacturers do. (There are many CPUs in modern automobiles, and auto manufacturers are often compelled to do recalls.)
They're certainly not pushing the bleeding edge at all like the 3 GHz desktop/laptop processors.