A short fable of software engineering vs. regular engineering
researchblogs.cs.bham.ac.uk
researchblogs.cs.bham.ac.uk
We can manufacture billions of copies of our product with error rates that are almost too small to measure. Six sigma? Hah!
Our design processes are far in advance of the automotive industry. Continuous design is normal practice. It's only the most careful (or most backward) shops that'll take months to update a design.
The automotive industry, by comparison, regularly takes years to update a design even by a minor version. The best shops at this i.e. the F1 race teams can fix bugs in days, minor releases in months and, if this season is anything to go by, struggle horribly with major version upgrades.
TL;DR it's fallacious to compare manufacturing to programming. The latter is a design process even if it feels superficially like an assembly line sometimes.
One final point. I would have very much liked to have read an analysis regarding the different approaches of the engineers vs the craftsman. This is an interesting topic and does have crossover between disciplines. Oh well...
Working in an easier field doesn't mean your field is more mature. Especially when there is such dramatic uncertainty in the schedule and the output is so flaky.
The computing systems used in a car itself (what do you think triggers the airbag on an accident?) can be trusted because of maturity in Software Engineering.
Startups don't choose to follow these methods to keep costs low and stay lean. I say that in the broadest sense, because a banking startup will be different from your average SaaS startup.
We (as software developers) don't need to adopt their models and processes, but we need to figure out how to work with them to ensure higher reliability across the board.
The alternative will be formal certifications/licensing pushed on us by legislators who don't know how to send an email and think mathematics is a weapon to be regulated.
That said, I've worked in teams who were developing software that would be boxed, shipped, and installed on millions of remote devices, and I've worked on teams whose main focus was "keep the service running at all costs". The former definitely emphasized root cause analysis, because screw-ups were insanely expensive and difficult to fix once they were shipped. The latter - not so much.
I couldn't disagree more strongly.
> Our design processes are far in advance of the automotive industry. Continuous design is normal practice. It's only the most careful (or most backward) shops that'll take months to update a design.
I'd say the latest trends in software engineering processes are actually adaptations of what has been practiced in the automotive sector (and other manufacturing sectors) for a long time, but outside of regulated industries, practiced with far less rigor. I don't think the claim that development processes in software are "more advanced" holds water (what does that even mean?).
It's impractical to practice continuous design on a physical product like a car so I think that's a poor point, but kaizen (which I will argue serves an analogous role) is certainly used throughout manufacturing and has origins in the TPS.
> The automotive industry, by comparison, regularly takes years to update a design even by a minor version. The best shops at this i.e. the F1 race teams can fix bugs in days, minor releases in months and, if this season is anything to go by, struggle horribly with major version upgrades.
Ignoring regulatory and documentation concerns, part of the reason it takes years to update a design is that the costs associated with producing a physical product are completely different from a pure software product. It's more cost effective to do good design work and heavy V&V of your design before you order tooling, rather than trying to modify your tooling (if it's even possible) after an error is discovered.
What does that word even mean?
> Software design is much better and more advanced, because our industry allows us to choose better practices.
I don't know what "much better and more advanced" is specifically referring to, but in the general sense, I see SDLC processes used in the stereotypical tech startup to be cargo cult appropriation from other more mature industries.
* Several of the component programs read from multiple config files
* Manual registry settings
* Config files have True/False, On/Off, and Yes/No settings.
* The version is reported to a db on program run, but the version is read from an xml file, not the program.
It is true that a new car program can take 5 years or more to roll out, but at least it actually happens. This system has been in use for at least 15 years, and it just keeps getting kludged.
It's comparing the build stage vs the literal produce-the-thing stage.
I also think the stoppage of the production line happens in software.
If my code breaks while in production, I grab as many log files and stats as I can, roll back the code, deploy, and then study exactly why everything failed, and if we recovered post rollback.
That seems like the similar case on the car production line - haul the faulty model aside (get the logs/metrics), figure out exactly what happened, fix it.
Who knows how the actual assembly line was put together. Maybe they iterated hundreds of times.
I don't expect my code to fail in production, otherwise I wouldn't have put it in production, but I understand it _may_ fail. That's why I put in logging and metrics collection, the same way those engineers hauled aside the car.
Executing a program is analogous to running the production line.
It's intellectually dishonest to compare design with production.
[...] a whole bunch of engineers tried to figure out not how to
fix the problem but why the problem happened in the first
place.
This is exactly what debugging is. My friend was quite obviously an accomplished hacker.
His hacker mentality would have fared quite well in
software production, where debugging is part of the
normal workflow [...]
Finding a workaround in production is not debugging even if it's useful.Yes and no. Debugging (in the sense that programmers do it) stops when you learn "X wasn't initialized" or "that buffer isn't long enough for the trailing zero".
This is root cause analysis (https://en.wikipedia.org/wiki/Root_cause_analysis). It goes way deeper into "why the problem happened". The buffer may be undersized because the programmer isn't educated well enough, because (s)he was too tired to think properly, because the spec specifying the maximum size was wrong, etc.
Each reason leads to a different fix, often through more iterations of asking "why?"
The programmer's lack of education can be caused by a problem in hiring practices, by low quality of internal courses, etc.
The programmer may be tired because of long working hours, because of troubles at home, because (s)he goes out too much, etc.
That buffer size might be a typo, an OCR error, somebody declaring the maximum size to be N because, otherwise, the data couldn't be fit in the current hardware, etc.
Each explanation comes with a different way to prevent the problem from occurring again (something that traditional debugging does nothing about. Initializing that variable makes the problem go away, but in itself does nothing from you writing code with initialized variables in the future)
I still have memories of my pass by Uni and how a Software Engineering professor constantly refered to programmers/developers as "coders" or "monkeys". He never stopped of repeating to us, day after day, that we shouldn't program ever. Our job as engineers was to plan, design, assess and manage those "monkeys".
More than 15 years of professional career later, my initial suspicion have been confirmed one time after another: those professors from Uni need a good deal of fresh air and to get in touch with the reality, instead of spend the day pontificating like celibate catholic priests about what you can and cannot do in your bedroom.
I know that many people here love Dijkstra and people like him, and even enjoy to feel their intellectual whipe in their mind, reading their essays and feeling bad about themselves as the essays repeat constantly how that everything is wrong and broken in the software profession.
For those people, I suggest to read "To Engineer is Human" by Petrosky and learn about how real engineers really work.
Unfortunately this view of people that work in manufacturing is pervasive, but it is also incredibly naive. That view is essentially the reason why American car companies (along with many other industries) got their clocks cleaned by the Japanese car companies over the last 3 decades. I've been involved in a lot of companies where we worked really hard to correct that mentality and the results were always obviously positive, but incredibly difficult to achieve. Engineers just love to look down their noses at everyone else.
There is nothing worse than some new kid that just graduated with a BS in engineering pretending like he's smarter than a guy that has been building the product we're designing for 30 years. Unfortunately some people never grow beyond the 'new kid' stage.
My comment tried (possibly not very clearly) to talk about the differences between the other engineerings and software. People keep thinking that programmers are like factory employees. And that's so wrong. Each engineering endeavour is different. In software, our factory workers, are the tools that do the job. What in other places is designing in software is just programming. In my view (and I'm not alone in that) people keep applying the manufacturing metaphor to software for the wrong reasons.
In construction, the distance between design and implementation is so huge, that the construction workers save or doom construction projects all the time. Spend some time in a big construction site, in their meetings and you will discover why those projects take so long to finish. Hence, they are important because what the designer design is just an specification for the project, it's a map not the territory. An ideal brick in a CAD program is not a brick. Your perfect design for the AirCon machine in that room got destroyed by an in place decision by a worker of changing the wiring in a wall, ignoring the plan. Through tremendous effort during centuries construction has accumulated more or less accurate models for bricks, but the workers do so many things that the architects doesn't know how to do that the distance persist, and my times just ignore them.
In manufacturing, you have the Ford model of car manufacturing. It has being employed in nearly every single factory in the world. Yes, the Japanese empowered their employees more than in other countries but hey, they took note and introduced some of the practices. As of today American companies are as competitive as the Japanese ones without all the mumbo jumbo and black zen magic.
My point here is that each engineering is different. Different factors forced us to discover what work for them and what doesn't. That knowledge was accumulated in the industry and depending on the engineering, the feedback loop from the industry into the University could vary a lot. In software, the distance between the two is huge, due to the lack of transparency and secretism that most software companies exhibit.
I'd like to work in a flat(ish) meritocracy. Maybe I have to start my own company...
Here's a good article explaining how Toyota make use of workers to develop new ways of building and improving construction:
http://www.japantimes.co.jp/news/2014/04/07/business/gods-ed...
If you are ready to pay the same amount of money - you can get the same level of quality in software too.
The problem is most people do not even think software should cost anything.
The problem is, for most things, it's too expensive. What we have is good enough.
Do we? Let's say, starting tomorrow, you have to write rocket-guiding software. If your program has bugs, you will be summarily executed. What are the odds that you will survive the process? I wouldn't estimate mine at anything significantly above zero.
This observation is often used to point out that "software is different" from cars and other hardware. Cars and hardware have lots of chances for defects, but software copies are essentially perfect. Assembly line products are not copies like a package or disk is a copy; each finished unit has lots of effort in it.
But if you still want to use the hardware example, find what the equivalent products are. A lot of effort goes into physically assembling a car, and then the car cranks out transportation.
A lot of effort goes into physically assembling a software team, and then the team cranks out software.
And while of course there are profound differences between a car an a human team, it might be useful to put engineering resources into designing and producing a team.
When an issue comes up, the first thing to do is to put together a workaround, a way for business to continue in absence of a proper fix. I then, without the need to rush, dig into the issue, figuring out a way to reproduce it, then isolate it, then fix it in a way that makes the whole system more robust and not less.
Having worked on the codebase / infrastructure this way for two years, my workflow is pretty-well refined. For my current issue, I'm devising a command-line tool to download and parse the log files to give me exactly the info I need to run down web issues. I can yak-shave to my heart's content.
However the biggest obvious difference that prevents dev becoming like an engineering assembly line is that the engineers on the line are not surrounded by car customers changing how the car should or shouldn't work/look/cost every two weeks.
Not that cleaning up the actual process of development so its more easily and naturally verifiable wouldn't be a good thing.
That's a good start
His friends job is to fit one small body panel on the chassis, over and over again. So when he gets one that doesn't fit his instinct is to make it fit as opposed to asking why it doesn't. His job isn't questions, his job is to fit it on.
When he was working in his own body shop this hacker wouldn't have stood for his tools putting out misprints and mistakes, but then his job was managing the whole shop.
Static monotony gets static monotonous thinking.
Civil engineering is real engineering but jobs are usually unique. Problems in civil engineering projects are usually 'debugged' and 'patched' on site without going all the way back and revisiting the original design.
Diane Hartley got no credit for her rôle in saving the building until nearly twenty years later.
So there's this topic of failure analysis, and we don't always do this well in the world of software engineering. We don't like to analyze our failures probably as much as we should, possibly because we find them embarassing, or perhaps they might be bad PR if they were known. So we just want to get the fix done and move on.
But the thing is for most software it doesn't matter how we deal with bugs, and that's because they cause minor problems. But when it comes to loss of life or money, bugs do matter. Doctors and lawyers have professional boards that oversee actions and disputes. Airplanes and Trains have the NTSB that oversee accidents and engineering failures. Civil engineers study the collapse of bridges as part of an effort to make bridges safer.
Failure in those cases are bad.
If my software solution is off by a few bugs, meh, so what? It doesn't cost the company money, usually. And it probably didn't kill anyone, and it probably just fixed some small css issue that looked great on firefox, but bad on chrome...
This happens all the time in software engineering. Less because we're shocked that bugs could happen, and more because when you are working from the assumption that what has happened should not be possible, when it happens, that means your mental model of what is happening within the system is wrong in some way.
There are the bugs you understand (or think you understand), and those you can safely allow workarounds for in most cases, because you understand the scope of the problems. Then there are the bugs that you don't understand, and it's important to be very careful and provide them the attention they deserve, because while they may be simple fixes, they could just as likely be an insidious case of data corruption or loss that you weren't ware of.
If you take the same code and put it through a compiler over and over, you will get roughly the same result. If you run the same code in production day after day, some days you will get different results because of the interplay between various systems or conditions.
If you take the same design and put it through a computer controlled machine you will likely get the same results unless something breaks. If you take the same design and put it through an assembly line, you will get different results on different days because of the subtle interplay between various variations in processes.
The distinction is critical, because while you generally qualify a single operation in an assembly line much like we would run tests on a piece of software, you need to monitor an assembly line just as we need to monitor our software in production.
I see this as indicative that debugging production issues happens in non software engineering contexts as well, in spite of the fact that we're always told that their processes and procedures are so much more mature and robust than ours.
Software engineering is so far removed from other facets of engineering that comparisons to poetry make for better discussion.
The software written for mission critical systems, such as nuclear reactors and even cars (think direct injection, ABS, AWD and now auto-cruise control systems) goes through an extensive series of evaluations. Same goes for the software that helped New Horizons reach Pluto and beyond. It's just not possible to achieve such feats without a thorough software development process.
Agreed that the same cannot be said for startups, but comparing a startup to a multi-billion dollar car company with tonnes of regulatory and compliance checks is not a fair comparison to say the least.
But here is our dirty little secret - none of the software 99% of us write is mission critical. And probably 75% of it is never used.
Still, software systems are both more fragile and can easily be more complex than real-world systems, so the naïvity is actually trying to apply engineering processes with rigid "design" phases and handoffs to complex software processes.
(As seen in the software, in just about every car.)
Software debugging has 2 parts:
a) find the bug - lucky for you a machine does this step automatically
b) fix the bug - once the bug is found this is self explained
Industrial debugging has 2 parts:
a) find the bug - requires going on the production floor and measuring things
b) fix the bug - once the bug is found this is self explained
How are they different?
Neither is true.
I've answered you in the other thread.
The reason why there's a lot of bugs in software is that software is soft. People know that changes can be tested with relative ease, and they also have a good idea of how wrong things will go if they do go wrong, and quite often an error is tolerable. This means we can often play a bit more loosely with correctness and reliability in order to gain some speed.
There's plenty of software problems where this isn't the case, and you end up with a very different process.
I've worked for a software company where it would take almost a week to add one new property to a JSON object because there was 3 different sets of tests I had to write/update across multiple code repos and it involved the coordination of multiple teams across multiple cities.
I've also worked for startups were it would take me an hour to make a similar change on my own but with lighter (albeit just as effective) test coverage.
It all comes down to company processes and approach to risk.
I think the lean approach is much more efficient if you have a good/experienced engineering team which you trust and who actually care about the product.
If Google tried to figure out why the hard drives fail in every server, how many years before we could type McLaren in the search box ?
Instead they embraced a different truth - things fail, so we'll build our systems to handle that.
And I think this is the reason why we can get away with bugs in software - because we can isolate and prepare for components failure or imperfections.
So instead of trying to absolutely NOT introduce any bugs, we can instead focus on making our systems fault-tolerant and introspective and handle all errors in a graceful way.
In a startup environment you are first and foremost trying to figure out what your customers want. Formal proof is not going to help solve this problem. I'm sure McLaren is also not spending days trying to understand why a panel is off by a few millimeters when they assemble a prototype.
[1] Paris subway (line 14) is controlled by a software written using a formal proof method called the B-Method: https://en.wikipedia.org/wiki/B-Method
It is easier to achieve string AI than to write bug-free code, and even that won't help: a super-intelligence will still have bugs in the code it writes.
At the end of the day it's about money and risk.
But, the interesting difference between software and mech eng isn't the attitude to faults, it's the fact that software isn't designed to work within tolerances. We don't build software around dimensions or stresses, we expect it to work always like maths, but it isn't, quite.
You see all. You see people hacking a bugfix in one day on the F1 site, but this fix will not survive. Some shops restore from their backups nightly, you will certainly face a lot of meetings and reports to get your one-line fix back into production. Even on R&D problems, which mostly describes F1.
This case reported here was on the assembly line, there are different rules.
If a strange misproduction happens you certainly want to be able to evaluate the riscs properly. The analogy is a Heisenbug in the production compiler. How often does it happen, how expensive is the problem, how expensive is the fix. There's not one car on the risc, there's the whole line in question. The new robot? The SW or the HW? Wrong measurement, wrong planning, broken sensor, broken motor? Or a real Heisenbug, which appears in HW much more often than in SW. Magneto/electric cabling problem. Or sometimes even light. The cable guy bending an overlong fibre cable <30cm radius. A fibre with a bad cut at one end causing all sorts of crazy effects, and crazy hackish fixes. Missing CAN terminators (analogy to missing SCSI terminators on harddrive chains) in certain F1 cars. It might work, but I wouldn't want to sit in this thing. This new 140.000 rpm high-speed rotor behind my back, making the driver uncomfortable, who is starting making bad mistakes out of fear, and rather pulls this thing out. Relations to completely crazy effects from outside, such as a cleaning lady turning on a bad fan, causing a jitter or spike, having some effects elsewhere. A bug, a real bug, in the system, esp. in Asia or Australia.
I see not much differences from SW and HW engineering. But the biggest difference I know about is the methology.
Agile in car manifacturing or HW engineering at all only happens at Toyota, nowhere else. Properly planned waterfall, with excellent and high powered middle management is the norm. If in Asia, Europe or the US.
There do exist some very small HW shops, with 1-4 engineers and 2 managers/sales people, and 2 support. There the development is of course engineering driven. Or there are some exceptions, such as Toyota or VW with crazy assembly optimizations, or e.g. Honda where their boss is one of the best engineers (they are lucky). But they eventually they ran out of luck with their funding and had to sell their very best test site, which was then taken over Daimler. That's why their are leading the F1 currently.
Another difference is proper training. These good engineering companies train their people a lot. Maybe 10x more than in normal SW engineering.
Having done 10 years of OO before I began coding in a functional style and realized all the benefits, this is my experience talking.