Fixing all the bugs won't solve all the problems – Deming's path of frustration
shermanonsoftware.com
shermanonsoftware.com
https://www.cs.utexas.edu/users/EWD/transcriptions/EWD03xx/E...
https://www.cs.utexas.edu/~EWD/transcriptions/EWD10xx/EWD103...
There are several takeaways here, but three of my favorite highlights:
* Dijkstra actually did not like calling such errors "bugs", as it is a framing problem.
* Likewise, he believed that we should avoid anthropomorphizing software or identifying with it.
* Finally, he believed in building a formal specification of software and then proving that the software matched this specification. Part of this was to avoid "bug fixing" hunts in which the focus was on "debugging" instead of correctness. But also, this was to ensure that there was a proper system view of the software, which ties back into the conclusion of the article here.
It is a tragedy that programmers nowadays don't read Dijkstra nor think about what he wrote and meant. I tell people to think of him as their "Cus D'Amato" if they want to be a "Mike Tyson" in their field i.e. a man who has thought deeply about the subject, knows all the angles and can "train" one in the "correct" manner of writing programs.
Of course, there can always be errors in specification, but this is supposed to be the first place we implement SAT solvers or proof assistants. During Dijkstra's time, such technologies were not yet available as they are now.
There is work to be done until this is all practical and the overhead to do this is within the current overhead of software engineering, but we are quickly reaching that point in time.
Given unlimited time and budget, plus the guarantee that things won‘t need to change in unexpected ways, that would be great.
The overall system and application changes in unexpected ways, but as you get lower down the stack, the amount of churn is reduced. Hence, there are strategies to design systems and applications so the most high-assurance pieces are those least subject to change, and that the application code itself runs at the lowest assurance level.
Of course, Dijkstra was more of an academic, but this can be managed in real systems with engineering process. A system that is 20% formally verified is safer and has fewer defects than one that is 0% formally verified. The key is applying engineering to this to ensure that the time and budget spent on this specification provides the most benefit.
These issues are fundamental, and persist, regardless of the technology, or state of the art.
plus ça change
[0] https://stevemcconnell.com/articles/software-quality-at-top-...
[1] https://stevemcconnell.com/wp-content/uploads/2017/08/art04-...
The clear specification really is the solution to resolve this issue. That you don't declare behavior as buggy that is according to specification.
Fixing bugs also seems to be limited to try catching all bugs instead of making changes to have the program do the correct thing like only loading the necessary data.
Gerald Weinberg talks about "Software Stalins," managers who grab onto one indicator and think that driving it to zero (or 100, or 11 for you Spinal Tap Fans) will resolve all other problems.
If only things were so simple.
Perhaps a lot of products needed bug fix campaigns because most corporations tend to reward SWE productivity based on so-called "evidence" like:
- KLoC
- # of commits
- adding features
- so-called "impact"
, rather than:
- reducing the cost to maintain code
- making code more robust
- making features actually work as users expect
- rip-replacing what shouldn't be salvaged because it's too expensive
So instead of adopting a funny mustache and making inflexible commandments, it would be less disruptive to have targeted bug fix campaigns where needed to boost nonfunctional requirements like quality, reducing support costs, and improving user satisfaction to their desired target levels.
https://www.macrotrends.net/stocks/charts/CSCO/cisco/stock-p...
I don't mean to take away the truth of your experience or dismiss that you might better know the disconnect between what you describe and what outsiders might see in hindsight -- it's just the disconnect itself that's fascinating.
> If only things were so simple.
Indeed!
Part of the challenge in a "freeze spec" approach implicit in "only fix bugs" is that it assumes the market requirements are well established. That was definitely not the case for Internet protocols/functionality in 1992. HTTP traffic did not exist yet (along with a couple of dozen other things that would quickly become dominant by 1997.
The effect of the bug fix stall was to give our more nimble competitors more growth and more runway. The fact that Cisco did very well was despite the "just fix bugs" mandate.
I offered the story because I thought some folks reading the linked post would find it unbelievable that a company would decide to fixate on a single simple metric when operating in a complex and evolving environment.
Though: version control ... 1992 ... haha.
Hey, that's when that unwieldy behemoth ClearCase was released.
Tell me more about how you would subscribe to a MS "bugfix only channel of updates" ? This is my first time hearing about that.
Isn't that what's meant by LTS?
Who was competing with Cisco in 1992? I struggle to remember. They were so dominant in the "dotcom" era (late 90s) that it felt like they were the only choice, the "IBM" of enterprise networking hardware.
Incompetent management is something of the norm in software engineering. Large successful companies have a deepish management hierarchy (4+ levels) and the competence of the org is capped by the least competent manager in the chain - who tends to be nontechnical and confused. Small companies have the ability to be more competent, but correspondingly tend not to have the resources to be influential players in the market.
I would expect that typical - maybe even above average - performance by software companies is accompanied by some breathtaking stories of management failure. If it were possible for a company to not make any mistakes in its software development it'd probably be a never-before-seen fountain of wealth creation.
I'm wondering what things are like in the alternate universe where OS/2 became the dominant PC OS instead of Windows.
Paul O’Neill famously did exactly this at Alcoa starting in 1987, focusing solely on worker safety resolved many other problems and multiplied profitability.
>The company's market value increased from $3 billion in 1986 to $27.53 billion in 2000, while net income increased from $200 million to $1.484 billion.
[0] https://www.forbes.com/sites/roddwagner/2019/01/22/have-we-l...
I think it makes a very good story that a focus on worker safety--to the exclusion of any other objectives--is all that you need. But I don't think anything is that simple.
If you have someone in your team like that, give them latitude.
> The software runs slowly because large amounts of unnecessary data are being sent to the users.
The manager manager was convinced that fixing all of the bugs would lead to perfect, 100% successful execution. The devops team lead spent an hour trying to explain why the software couldn’t achieve better reliability than AWS (the system spread across AWS zones, but each instance was contained within a single zone).
1. You can guarantee no bugs in your system
2. You can guarantee no bugs in AWS's systems
3. You can guarantee there's no nuclear attacks on us-east-1 within the next year or so
It's more changeable vs unchangeable. Part of that is technical cost, but in my experience most is perceived risk, which is largely a function of understanding and team (leader) confidence.
The scariest things are existential: telling people the system they believe in has fatal flaws introduces a gap that can only be filled with a lot of confidence, which is typically lacking when growing fast or with new systems.
The main remedy is to be sure you have an actual view and understanding of the system, and address things from that perspective.
Otherwise, bugs are just make-work.
/s/to/from