The Black Hole of Software Engineering Research
blogs.uw.edu
blogs.uw.edu
If you major in "mechanical engineering", you will take courses in calculus, linear algebra, probability, computer programming, physics (thermodynamics, mechanics, electrodynamics, strength of materials, maybe aerodynamics and viscous flow), control theory (what used to be called "cybernetics"), manufacturing techniques, CAD, and maybe a management course or two.
The corresponding curriculum for software involves things like programming languages, algorithms, data structures, networking, operating systems, software architecture, and formal methods, and it is called "computer science". What is currently called "software engineering" is the software equivalent of "art appreciation". It's a desperate attempt by the cult of managerialism that dominated that 20th century to figure out how to run software projects without understanding software itself; it's an attempt to substitute credentials and command structures for intellectual mastery of the problems to be solved and understanding of the techniques for solving them.
The consequences, generally, are what you'd expect from a book publisher run by illiterates or a record label run by Deaf people. Experian and CGI Federal wasted hundreds of dollars building a "healthcare.gov" site that didn't work; SAIC wasted hundreds of millions of dollars on building a Virtual Case File system that was ultimately scrapped. These are the companies that represent the approach known as "software engineering", that get CMM certifications from the SEI.
By contrast, the companies that make software you actually use — companies like Canonical, Google, Facebook, and Apple — put a great deal of emphasis on programming, motherfucker, and very little on things like SDLC, requirements engineering, and "formal" specifications.
This is not to say that the issues that "software engineering" attempts to tackle — like testing, requirements analysis, and software architecture — are unimportant. They are very important. But the field that identifies itself as "software engineering" is in fact failing to tackle those issues. It is not where the insights are coming from. They're coming from programmers and computer scientists.
The worst code you'll ever see written is by computer science PhDs who just have to get something done for their research paper experiments. It'll be 6,000 lines of matlab with no loops, no functions, and entire 30 line sections just get copy/pasted when they need to repeat an action.
That's code without software engineering. That's "coding" and not "developing."
In a way, software engineering is a moving goalpost in the same way as "if a computer can do it now, it's not AI." Many of our standard practices these days used to be "must be considered separately" software engineering practices 10 years ago (revision control, proper documentation, creating libraries, etc).
Open Source projects run by communities with a focus towards testing, documentation, and maintainability are the best thing to happen to software engineering. But, there are still crazy people running around writing code just for themselves and refusing to realize software is more than code code code these days.
We pretty much need two divisions here: "software engineering (social)" and "software engineering (technical)". Social software engineering is defined by well maintained, long-lived, multi-contributor software projects. Technical software engineering is the implementation of the testing systems and contribution systems behind those long lived sustainable projects (e.g. are you calculating the branch coverage of your unit tests? do you even have unit tests? (no, testing your live servers aren't "unit" tests, those are integration tests) are you testing all combinations of acceptable and invalid inputs (null, proper value, invalid value, out of range value positive, out of range value negative, ...)? are your systems even testable and not bound to side-effect laden global states? do you have proper mock objects to compensate for your side-effect driven development? do you allow outside contributions so you don't have to maintain all 60 of these detailed testing and failure scenarios yourself?).
(there's also the entire other branch of "how to be a better software writer through requirements gathering and agile/lean/XP/monkeypants practices," but that's out of scope here.)
No, that's software development. It comes in many forms. Your academic example is a good one of a failed approach.
Software engineering is applying engineering practices to software to produce The Right Thing. Examples of it are Cleanroom, Praxis' Correct-by-Construction, Green Hill's PHASE, and so on. These connect requirements to clear specs (semi-formal or mathematical) to code. The systems are usually abstracted, modular, and layered in a way to allow testing of individual components and how they connect. Both successful and failure states are documented with fail-safe methodology. Design/code reviews, testing, static analysis, automated build processes, and so on are common here. The results are consistent, effective, and with low defects. Much like engineering.
Then, there's software development and programming in general. They're selectively applying some of this stuff but no way looks like engineering. The results are likewise all over the place.
"Open Source projects run by communities with a focus towards testing, documentation, and maintainability are the best thing to happen to software engineering."
All of that was happening in software engineering back in the 1960's-1980's for companies that tried to do the real thing. Just formal, code reviews and testing got incredible reliability gains (see OpenVMS). High assurance systems did much more. An article, about both security & software engineering, described the very best of security in 1970's with specific practices and results they had. Most software development hasn't caught up to that level despite the fact that every aspect of it was exceeded in academia and professional communities focused on quality/security.
What I do agree with is that OSS projects using such methods have gotten them mainstream adoption in many OSS and proprietary projects. The benefits were real. However, the stuff already existed for decades, was even mandated by some standards, and the mainstream is a pale imitation of all that. Improving, for sure, but it's actually them that still have much to learn. If only the will.
"We pretty much need two divisions here: "software engineering (social)" and "software engineering (technical)". "
This I'm less sure about as now we're on an open-ended topic. :) There's definitely social and technical factors. I'd say the technical side is designing, implementing, and verifying any aspect of it. You need to understand the tech to varying degrees to do each. At same time, there's a social component because everyone is working together and one can affect another. seL4 teams, development and formal methods, were an example where constant communication solved issues where one side's abstraction/implementation was causing severe issues on the other. So, it's all technical and social minus maybe project manager or owner(s) making the money off it. ;)
I agree with you that those things are how you take raw code and turn it into reliable systems, and that you don't get there by copying and pasting code and writing 6000-line functions. But this article is not promoting those things. It's promoting formalized process engineering.
(However, in my book, using very-high-level languages like Matlab (or, better, Octave or Numpy), and high-level operations like matrix-multiply instead of explicit loops, are best practices.)
I do agree that social practices around effectively collaborating with other people are very important and require a lot of effort. You can have the world's most effective hackers on your project, but if they're wasting their time because you don't have a budget for a test server or because they're not allowed to fix the bugs they find, your productivity is still going to be low.
"What we usually call "unit tests" these days, and certainly "mock objects""
That's not true. Like other assurance activities, Extreme Programming just did what was already done, gave it a new name, mixed it with bullshit (esp pair programming), and got it some needed attention. The Orange Book testing criteria, in 1983, required as much testing as resources allowed of primitives, their compositions, all interfaces, and common failure conditions. That's on top of specifications, covert channel analysis, pen testing, using only proven components, build system automation, etc that far exceed XP's quality regiment.
Funny you mention Cleanroom because it was first thing I thought of when I heard the "new" Agile approach. Iterative development, code reviews, usage-based testing for acceptance, verifying each module... Cleanroom had all that. At least they eventually figured it out and it seems competitors in Agile identified XP's problems as well. Although old, I'm glad the concepts are mainstreaming because it gave us something important: solid, well-maintained, constantly-updated TOOLS to support the efforts. :)
XP "unit testing" isn't what's called "unit testing" in the software engineering literature, and it isn't designed to reduce bugs; it's an executable functional spec, and the reason you write the tests first is so that you don't end up designing some monstrosity of an interface whose clumsiness becomes contagious to everything that calls it. "Mock objects" are a particular way of writing those tests in which your testing stub objects actively validate that they're being invoked the way you expect, rather than just passively simulating the subsystem being stubbed out.
Orange Book testing, by contrast, is aimed at reducing bugs.
XP-style test-first unit testing is not the best way to develop all software, but it works pretty well for some things. Similarly for pair programming: it can help a lot with intra-team communication, code reviewing, and not getting stuck, but I'm pretty sure there are projects where working alone and doing the code review later is a better practice.
Yes, Cleanroom has a lot in common with XP.
I think that a lot of the "competitors in Agile" actually suffer from the same disease I'm blasting here: they try to solve technical problems with management practices, and consequently you have a lot of Scrum teams with really shitty code who can't make much progress because they're beginning programmers, or experienced programmers who have fallen into bad practices.
Yes, I read on unit testing and XP with its other goal. Most of what the crowd pushes it for is the benefit of catching defects usually with requirements, dependencies, or breaks from later changes. Old projects used testing for those reasons, too, although not as consistently.
I'd say the real difference are greater focus on mock-ups and documentation focus as you said. Other than that, the techniques are similar and more ad hoc than before. Just getting more attention and use. Still A Good Thing.
"I think that a lot of the "competitors in Agile" actually suffer from the same disease I'm blasting here: they try to solve technical problems with management practices, "
Totally agree. This is a bad sign any time we see it. The focus should always be on getting the right people in there, making sure they have good tools/methods, and letting them get shit done. That's development or engineering. Documentation, etc really just becomes evidence of what was done and an aid for further development/maintenance.
I'll add that I think there's a zealot kind of thing going on. They blasted me years ago with all of the stuff like it's better while not providing any empirical evidence showing it was. Some evidence appeared around testing, etc and Spiral already proved incremental approach would help. Much of it had nothing but some people's word to back it. Pair Programming even countered years of evidence about effects of interruptions on flow.
So, I think it would help if they or the next group focuses on controlled studies pitting methods against each other on similar projects with objective and subjective accounts of results. Then, we'll have a better handle on what works and where. Right now, I have to rely on little empirical data and mostly just take the word of the experienced people delivering best results for given domains. Gotta get programming a bit more scientific in industry.
Same for any "end user programmer," they just want the result and not good quality code, even if it would benefit them more in the long run to write good code (this is known as the "paradox of the active user").
https://cdn.ncees.org/wp-content/uploads/2012/11/SWE-Apr-201...
And specifications for (a version) of the Mechanical Engineering PE exam:
https://cdn.ncees.org/wp-content/uploads/2012/11/PE-Mec-Syst...
It sounds like you could get a perfect score on the "software engineering" exam without being able to write so much as a for loop, let alone design a web server component architecture or debug a priority-inversion problem in a real-time system. And yet you have people who propose that such exams should be required for writing safety-critical systems.
This is intellectual fraud on the level of substituting homeopathy for medicine. It makes my blood boil, and it puts lives at risk.
Yes....theoretically these materials work this way, but practically no one can manufacture this thing for you. It's definitely more apparent with SEng (especially because I'm familiar with the topics), but the real incentive of having licensure is the ability to revoke it and put liability on the licensed individual.
That was very much a management problem; from what I posted to HN back when it was happening (https://news.ycombinator.com/item?id=6586705):
The government took way too long to get started "bending metal", it had the inexperienced/no experience on this scale HHS Centers for Medicare and Medicaid Services government bureaucrats be the integrator for this 50+? contractor effort, per the NYT "In the last 10 months alone, government documents show, officials modified hardware and software requirements for the exchange seven times." (https://news.ycombinator.com/item?id=6583327), those changes continued through the last week before the launch, and full testing obviously was delayed until that last week. Underlying their inexperience, evidently CMMS didn't see a need for bottleneck monitors, given that we've heard of them being put in place post-launch.
And we learned they completely underprovisioned the database, and the specific one CMMS forced the contractors to use was unfamiliar to them, in paradigm (not a RDBMS), let alone the specific one.
And refused to listen to the people waving red flags started no later than spring of this year, when one of these bureaucrats changed his goal/desire from a "First World" website to one that wasn't "Third World"....
That this front end was a disaster turned out to be a major blessing: https://news.ycombinator.com/item?id=6701748 (more details as well).
As you say:
It's a desperate attempt by the cult of managerialism that dominated that 20th century to figure out how to run software projects without understanding software itself
The key to me being "without understanding software itself".
I've personally observed that many of the worst managers I worked under were failed software developers....
I didn't mean to claim that management wasn't important. I meant to claim that the proper way to manage software projects is not by "software engineering", but rather by giving the reins to somebody who knows what the fuck they are doing. "Software engineering" is a smokescreen to provide fake credentials to charlatans so they can get put in charge of projects like this.
Any ABET-Accredited program will have you taking the majority of those classes as a Software Engineer. I went for ECE with a focus in Software and this was certainly the case for me. The problem is that "Software Engineer" isn't a protected term, so anyone whose ever compiled a hello world can title themselves with it. With Software becoming a serious infrastructure to daily life, it's time to add rigor and standards to the people calling themselves such, just like a Civil would. Of course, this wouldn't apply to most software jobs. But if you're in charge of making, say, emissions control software, you should be properly vetted and held accountable.
Second, I don't agree with your prescription for emissions control software. The problem is not that too many people are able to write emissions control software; the problem is that too few people are able to write or read it, which makes it impossible to hold the authors accountable. If every car engine control unit shipped came with the fully-commented, compilable source code on a CD-ROM, and it was commonplace for people to patch it to reduce emissions or deal with unusual circumstances, Volkswagen would never have had a chance at sneaking its cheat in.
I believe that I find those problems once in one or two years. Through all my career.
According to Wikipedia, Software Engineer is protected to some degree in Texas and Florida as well.
First was probably Dijkstra's THE operating system in 1968. Careful use of abstraction, interface checks, systematic testing of individual components, and careful composition led to a system that performed nearly flawlessly. All high assurance systems used a similar model in safety- and security-oriented systems with near flawless results despite much analysis, pentesting, field use, etc. Still same concepts used to this day same results. QED already. However, high assurance is just too limiting, slow-moving, and heavyweight in skill set. Best to focus it on premium components (eg TCB, key protocols). So, let's keep going to see what software engineering might have given average programmers.
By early 70's, Fagan at IBM implemented Inspection Process that formalized software review to try to counter all these bugs they caught. That it was systematic, regular, measured, and adaptive meant it caught tons of errors in the studies. Code reviews have only become popular among "programming, motherfucker" types in the recent decade with companies still trying to lure more in with HN articles despite it being empirically proven back in the 70's. Software engineering R&D also developed tons of tools to help but most programmers don't use them outside maybe testing or build systems.
Note: Around this time period, Wirth's languages and work built on them provide strong type safety, memory safety, interface checks, and one was immune to concurrency errors w/ language + architecture choices. Only a small niche adopted them and programmers were mostly not using them. All the buffer overflows, race conditions, etc followed logically.
By the 80's, Harlin Mills tried to stick it to hardware people by inventing the first, true engineering practice for software: Cleanroom methodology. Combined a formal notation, structured programming, decomposition, safe subset of language, by-eye verification conditions, usage-based testing, and statistical certification of quality. IBM tested it to find their defect rate was at shockingly low levels on first run with it improving from that point on. CASE tools were developed to semi-automate it. Small businesses differentiating on quality issued warranties (WTF!?) for their software at specific defect rates. Got mentioned on Dr Dobbs with a chance to reach programmers but "professional programmers" don't engineer software regardless of evidence. Cleanroom faded, typical programming continued, and basic defects are all over the place in software which is barely comprehensible.
From the 80's to now, research continued on building tools to automate as much as possible. Some prevented issues, some detected them during development, some caught them at runtime, some just recovered. In some situations, the whole problem could be solved via a synthesis and with reliable results. Programmers totally ignored this line of software engineering research. Unfamiliarity with languages like Wirth's, Ada, ML's, and Eiffel meant that would happen to a degree. However, programmers continued to ignore results even when applied to C, Java, etc that they used. A niche market of companies recognized the value and started applying such tools. CPU vendors, in EDA market, poured tons of money into it due to the cost of defects and got a series of incredible results going to this day. It's why hardware has a tiny defect rate despite tens of millions to ten billion logic functions running concurrenty.
Now, in recent events, some lessons of software engineering research and innovations have ended up in some mainstream stuff. Go, Rust, Scala, and Clojure come to mind immediately. The articles on them are pushing the stuff software engineers, especially in academia, had been pushing all along. Imagine how much better our code & tools would be had programmers listened to software engineers idk maybe 40 years ago.
Meanwhile, there's still companies and R&D leading the way. Altran/Praxis's Correct by Construction process has been in the lead with mathematical specs of behavior, proper error handling, SPARK implementations proven free of common bugs, and field results showing among lowest defect rate in existence. Galois is applying software engineering to low-level stuff combining Haskell, DSL's like Ivory, and C with good results. The Design-by-Contract crowd has been catching errors others missed. DO-178B scene is applying every tech and trick they can to reduce defect count. Embedded scene, like Esterel's SCADE, is generating whole apps from specs that work the first time. One academic combined Cleanroom and Python to get an amazing productivity-to-defect ratio. And so on and so forth with every independent evaluation and case study showing these methods deliver results "programming, motherfucker" types still believe don't exist. And predictably. And repeatedly. And even mathematically if you want. Kind of like engineering.
So, software engineering exists. The term got abused by the wrong types of people pushing the wrong stuff. The good stuff got ignored by them and most programmers. However, there's a string of tools, papers, companies, and so on applying that stuff. Their results are better than programmers almost every time sometimes with less work. Programmers are even adopting some of their stuff after decades of mocking or ignoring it. Those programmers own stuff is getting better. So, software engineering is a thing, it's contributions to our capabilities are enormous, and we'd all benefit if programmers used more of it.
My slogan: "Engineering, motherfuckers, do you speak it!?"
"No? We'll, that explains the results..."
You don't get high-assurance systems by improving your metrics collection and reducing the standard deviation of your effort estimates; you get them with, as you say, careful code review (by people who know how to program) as practiced by "programming, motherfucker" types like Gosper and Beck and Fowler for decades, effective abstraction, and code that is correct by construction and formal proof rather than merely by testing — which you also need.
Dijkstra spilled gallons of ink in the 1970s and 80s railing against the hucksters who appropriated the "software engineering" term, trying to redirect people's attention back towards writing software that could be logically proven correct. It's taken decades of effort to start to realize his vision (and Hoare's, et al.), but that research isn't happening in "software engineering"; it's happening in computer science.
Hardware has been able to take advantage of formal methods for much longer than software, because hardware has never been more than a thousandth as complex as software since at least the 1950s.
Cleanroom (btw, you misspelled "Harlan" as "Harlin"; the poor guy deserves better) is great as far as it goes, but it doesn't focus on proper abstraction design, and its testing methodology basically assumes you're programming a finite state machine. But if you're doing that, there are much better ways to verify it, including rigorous formal verification.
Formal verification comes from logicians and computer scientists, not the kind of "software engineers" who write System Definition Documents. Golang is from the Bell Labs tema that gave us Unix. Rust is an application of affine typing, a variant of Girard's linear typing. Clojure is an application of research into FP-persistent data structures, using the syntax of Lisp, the classic hacker's language.
These are not the kinds of things the original article is directing your attention to. It's not talking about affine type systems and Hindley-Milner type inference. It's talking about management research masquerading as research on how to engineer software.
And that's a fraud.
I totally agree with you there. I think a lot of the dispute in my post and yours centers on the definition of software engineering. You're applying it as the approach strictly about process, paperwork, drowning in numbers, etc with nothing produced. I'm applying it to mean any software development activities that apply engineering principles to software with engineering-like results. Computer Science is where a lot of what I'm citing came from, although CompSci != engineering software in general case. So, I have to have another term and "software engineer" literally means a person that engineers software.
Think I need a solution to this problem as it's recurring. I've clearly illustrated most programmers, even good ones, don't apply engineering fully and leverage what's proven. It's always a subset, if at all, often with informal approach driving the whole thing. There are groups doing the whole thing like engineering but they're rare & not paperwork pushers. If the phrase is irreparably tarnished, I might need to ditch "software engineer" for a phrase that describes the difference between how a group like Praxis does things and how a decent programming shop does it. The difference is how systematic it's designed, assured, and documented with every available, proven method put to use plus careful tradeoffs. Need a new terms I'm thinking because old one caused a huge chunk of your counter-points which apply to the crowd we both can't stand but not really the niche I'm promoting.
"Hardware has been able to take advantage of formal methods for much longer than software, because hardware has never been more than a thousandth as complex as software since at least the 1950s."
Partly and often developers are reason. Getting from hand-wired hardware to synthesis, formal methods, simulation, and testing took hundreds of millions in investment with many bright minds. The reason it worked was they noticed common patterns in their designs/implementations, started reusing them in new ones, built verification techniques for them, focused heavily on integration/reuse, and just went from there. Cleanroom was one of many methods that did something similar (but semi-formal) for software in structuring, verification, and integration. Formal verification teams did it again in so many isolated cases covering so many properties and types of software. The Cleanroom, FV, and model-based synthesis results show software development could get as much verification as hardware for sure. Just would take a sizeable labor investment in tooling and libraries plus agreement to simplify implementation and interfaces to facilitate analysis. Not saying that will mainstream but it's achievable and can be incremental. At least we're getting type systems and tooling, though, from the niche (esp CompSci) that cares about the stuff.
"btw, you misspelled "Harlan" as "Harlin"; the poor guy deserves better"
Damn haha. He really does. Yes, Harlan's methods could be improved. No argument here. An example was how the graph-based models were doing better than hierarchical. Another is executing your code to test assumptions on its dependencies. Point about Cleanroom was that it worked, its style/results were much like engineering, demonstrated many good practices, had empirical evidence backing it, and was a logical starting point that was entirely ignored by developers. And it wasn't the only one.
Far as FSM's: you nailed that one. Almost all high assurance systems used an Interacting State Machine model where FSM primitives were analysed and composed functionally (often CSP style). It was the only proven model in terms of getting results that survived all analysis and pen testing. You can do rigorous formal verification but that's high assurance: something you and other readers reject automatically in these discussions. So, I left that off.
Plenty of research & products working on different models but they're still getting proven out over time. Gotta use what's proven for now if I want get the best results.
"Formal verification comes from logicians and computer scientists"
All true and good examples: I listed some in another comment. Then they or the engineers I refer to put them into practice. Most developers don't. Still need a term for the former that won't tie into the other crowd you've been describing that stole and tarnished the phrase.
1. "Programming is easy; what's hard is working together with other people." No, they're both sometimes really easy and sometimes really hard. For a lot of objectives you might have, you have to learn to do the hard parts of both of them. The hard parts of programming include design, analysis, testing, debugging, and verification. You aren't going to learn how to write a compiler in 21 days. (Most programmers don't learn how to write a compiler in 2100 days, although that's probably more because they aren't trying because they're intimidated.)
2. "Software is in a crisis because we don't know how to program, and the solution is to manage software projects as if they were construction sites, using misguided metaphors about bridges invented by people who have no idea what happens when you actually try to build a bridge." No, in fact you have to understand both management and programming (motherfucker!) to manage a programming project effectively. When programming is being done effectively, you automate all the repetitive parts, which dramatically reduces costs and timeframe, but dramatically inflates the unpredictability.
Instead of managing software projects like construction sites, we're going to be managing construction sites like software projects, using proven libraries, mathematical models, repeated simulations, iterative exploration of unknowns to reduce risks, and pervasive automation. This is all in Engelbart's papers from the 1960s. Doing things that way (applying software techniques to hardware) is how we became able to build hardware with billions of transistors in the first place. It might be 2020 or 2035 when most of the effort in building a house is done by people using computers to manipulate models of the house, and only the last step involves (fully automated) machines rolling onto the construction site to materialize your carefully debugged plans, but it's going to happen. (I don't know if being able to deploy a new iteration of your house every week in order to fix bugs and roll out new features is going to happen before or after that.)
Now, it's true that we don't know how to program (although developments like seL4, coqasm, Excel, NaCl, and PHP seem to be making some progress on that front). But the "software crisis" isn't about that. Rather, it's about how Moore's Law has reduced the cost of computers so much that programming is suddenly the limiting reagent in nearly everything in the economy. You can get a 48MHz ARM processor with 64kB of program Flash for US$1.76 and burn ten thousand lines of C into it, then use it to run a string of Christmas lights. For US$6.46 you can get a 48MHz ARM processor with 1MB of program Flash and burn a hundred thousand lines of C into it. Or you can put a Lua interpreter into 20% of that memory and fill the other 80% with eighty thousand lines of Lua. The "crisis" is that it costs literally a million times more to write the code than it does to make the processor. (Prices from Digi-Key, unit price for quantity 1000.) The "crisis" started in 1968 when the price of hardware at last fell below the cost of writing the software to take advantage of it, and it's been "worsening" ever since. And that's the origin of "software engineering".
3. "Programming is super hard and you have to be a wizard to do it at all." No, anybody can learn how to program. We have Excel, JS, and PHP. If you have a computer and already know how to read, you can learn how to program a little bit in a few days, and usefully in a month. There are still going to be a lot of things that are beyond your capacity, but you'll be able to do some things. It might take you years of work to learn how to write like Dean R. Koontz, but you don't have to be a Dean R. Koontz to write a shopping list, a YouTube comment, or a love letter. Programming's the same way.
http://www.dev9.com/article/2015/1/the-myth-of-developer-pro...
Considering making it my current reference for explaining this stuff to managers I meet in a short time and with a quick link. I also agree the software crisis largely came from the cheaper unit prices of hardware, complexity exploding in software to take advantage of it, and many management issues. Also inspired Wirth's Law where he made the same observation while trying to avoid it himself.
Only inaccuracies I see are these:
"When programming is being done effectively, you automate all the repetitive parts, which dramatically reduces costs and timeframe, but dramatically inflates the unpredictability."
It doesn't just do that. The 4GL's, the good ones, did that with great productivity benefits in the common case and ability to call custom code for the rest. Many programming aids do today, too. So do the languages offering good macros. However, the methods I advocate here do more than that: support designer's understanding of what system does; find high-level issues in requirements or specs; find interface errors (critical); prevent all kinds of things by construction; detect others; enable easier portability with native efficiency; ensure executing code matches source. Just a few recent examples and with each requiring human effort. No amount of programmer skill in C with ad hoc requirements/design techniques replicates these without considerably more effort or the most experienced, elite programmers.
So, the tactics and tools from Comp Sci + High assurance/integrity field certainly have value beyond automating tedium. These also tend to increase predictability once people are used to them because they reduce risk and the debugging that follows. The trick is to choose just the right ones that boost the developer's abilities without bogging them down. I don't have a single recommendation there as (a) one must use best tools for the problem and (b) they experiment too damn much in Comp Sci to the point that I can't say whats better for most methods or tools. Only what's worked and for what. So, the specifics I bring up change depending on topic & audience. Eventually, we'll see some convergence with more uptake. Like how Amazon's using TLA specifications now with results I predicted when encouraging that sort of thing for years.
" The "crisis" is that it costs literally a million times more to write the code than it does to make the processor. "
Here I believe your overall point is hardware, unit cost vs developer cost along with what leads to. I agree with that. Just wanting to make sure you know that line isn't true in general as many programmers think it is. Hardware development is expensive, esp CPU's. I've been pushing on Schneier's blog and elsewhere a number of CPU modifications that eliminate entire classes of errors (eg code injection, leaks). Most of them have prototypes to draw on or are specific enough for opencores crowd to handle. Yet, you see almost no attempts to put them into an ASIC even when people know about it. That goes for opencores stuff, too, for most part. That's because hardware development is ridiculously expensive even for adding a block or two to a RISC CPU with prototypes.
The reason is mainly the mask costs, tooling, and labor. The masks, which are necessary for printing chips, for 350nm-180nm (minimum useful) are $100+k per design test. Wafer and packaging might be $5-10k. Every screw-up requiring a change means new masks and maybe 4 months waiting. You minimum staff of 3-5 pros, as rookies make too many mistakes, will set you back same amount if not Asian and with tools that start around $50k a seat on low end. For hardware that's smaller or faster, all those variables go up dramatically with tools for deep submicron (esp synthesis) runnning $1+ mil a seat plus needing four or five different tools. Sometimes more than 10 on most advanced nodes. Your CPU was probably made on a node where one prototyping run (think unit test) of your HDL code cost $1+ million dollars.
There's tricks to keep costs down, mostly software-inspired as you guessed. Adapteva probably the best at it using pre-proven I.P., simplicity in custom stuff, re-use of it in future iterations, HW's version of structured programming, cheapest tools, utter pro's for team, and customers willing to pay significant, unit prices. Still cost them $2 million. Most SOC's costing $15-50 mil. I managed to come up with flows, open and proprietary EDA, for anti-subversion of open, secure hardware. Found shortcuts like S-ASIC's (eg eASIC, Triad) got it down to like $400,000 but with volume requirements. Still high. RISC-V Rocket processor, etc are happening because academics get paid in grants, get EDA tools for low five digits, and get discounts on ASIC prototypes. They're our best hope for OSS, secure chips. But I hope I've illustrated HW is in no way cheaper than software except in unit price of something where demand spreads out development costs to make it low.
Note: Microcontrollers you cite are a special case not like the rest. Developed by pro's, very limited functions, cheap tools, much reuse, on fully-depreciated fabs often owned by the vendors that reduce itereation costs, and in such ridiculous volume that further reduce costs. Even the cheapest fabs with your tiniest design will charge you $3,000+ for the minimum order. When has a compile-and-test cost you that much? ;)
With that foundation laid, I can then say your CPU is going to cost way more than the software. Especially if it's Intel, AMD, or IBM: full-custom for top performance, power efficiency, density. That requires more of the best engineers, more time, and more masks ($1+mil per set) on the best nodes. The inevitable logic screwups and their respin cost require extra tooling, both commercial EDA and custom, to counter the problems: Intel uses formal verification [1], IBM mostly falsification [2]. I'm not sure what your CPU cost vs the software on it as numbers are hard to come by. I do know that Solaris 10, a whole UNIX rewrite, cost around $200-300 million while the Cell processor cost IBM $2 billion to develop. Intel spends $10.6 billion a year on their product lines with most on hardware end. So, Standard Cell hardware is more expensive than software with full-custom plus high-end specs being insanely expensive.
So, it's probably best for that old meme to die since the world has changed and shit much more expensive now. Your points about effects of unit price vs developer cost still stand and all. I just wanted to make sure you knew just how hard hardware was today & at what prices. Plus, there's always the tiny chance that a smart guy like you reading these posts will say "Challenge Accepted!," then build a secure, surprisingly-fast processor for us on an older node. :)
Note: +1 for the reference to Engelbart. His videos prophetically talking about (and illustrating) remote collaboration, etc were some of the most visionary things I've ever watched from the field's history. Mind-altering to look back at them trying to look forward and getting a lot right for the time period. Very specific stuff, too.
" No, anybody can learn how to program. "
VERY TRUE! I try to remind everyone that doubts it. More tools and ease of use today than ever to get started at whatever level they're comfortable with. Worst case, they think it's too hard, say that yes they played with Legos, I hand them Scratch [3], and they on the way to programming haha. The basics are really as easy as breaking things down into steps and building them like Legos. A kid can do it... :)
[1] https://www.cl.cam.ac.uk/~jrh13/slides/nasa-14apr10/slides.p...
[2] http://fm.csl.sri.com/SSFT11/FV_Summer_School2011__HW_Verif....
Seems a bit misleading, UML is consistently the best way I've seen to communicate ideas to other devs (especially in an OO environment, but some diagrams could be useful from a functional viewpoint). The worst is spending an hour trying to work together with someone to figure out what they mean because they use their own invented notation or some pidgin-UML.
But maybe this is just a reflection of the limits of my experience, and if I had worked with the programmers you've worked with, I'd be more sanguine about UML? I might be one of those guys you hate because of his pidgin-UML!
As far as pidgin UML, as long as you keep to the basics you know and don't try to shoehorn something else in, that's fine. What grinds my gears is a poor reimplementation of something UML has already defined. I'd say most people only go so far as learning how to do class diagrams and don't realize the other diagrams are useful outside of upper management waterfall jackhattery.
The best part of UML (I believe this is the view even from the creators but I may be wrong) is that you can take only bits and pieces and it's still damn useful.
The immersion in industry was absolutely critical for me to see 1. what was actually needed in the real world, 2. what types of assumptions can be made about the predictability of large economic systems.
If I had gone to grad school right away, I think I would have followed several ill-advised models down way too deep. So, I'd say at least one year in a real role in the industry you hope to contribute is really essential for someone planning to do applied/translational research.
That doesn't seem like a great deal of experience to me.
Off the top of my head I can only think of one or two professors with distinguished professional careers who are principally focused on engineering, rather than computer science. Maybe there's a list somewhere?
I think what some people miss about having a formal methodology is that a lot of time can end up being wasted addressing the same types of issues (design problems, bugs, re-implementing things, multiple requirements changes, etc...). The same can actually be true with applying methodologies too rigorously though.
The most effective software development I've seen tends to be where the culture values both quality and time spent. If something can improve quality and reduce the amount of time spent on boring tasks, it generally doesn't take much convincing there.
I think in a lot of ways, a lack of focus on software engineering isn't necessarily an academic problem. It can end in some fairly bad results though, like the Knight Capital incident (https://en.wikipedia.org/wiki/Knight_Capital_Group#2012_stoc...), so obviously there can be a very tangible impact.
Of course there are exceptions, and I don't mean this as a criticism of the researchers. It's just that there's a lot of work involved in "tech transfer", and it's often hard to justify the investment -- especially for a relatively-small team. Microsoft and Google and Facebook all have plenty of people doing various forms of "applied research" to bridging that gap; but they've got the resources to do it.
That's a lot of stuff. For example, probabilistic programming is concerned with all computable distributions. This is the most general class of sampleable distributions.
Furthermore, I don't think you need to expand or distort the definition of mathematics into "math-ish". For instance, many branches of engineering or even physics can be intensely mathematical without being mathematics, strictly speaking, whereas code for algorithms and data structures are highly abstract and more or less form a kind of proof. Sorry to sound pedantic, but some fields use math at a very high level, whereas CS is math.
There are some courses in CS, like operating systems, that I wouldn't quite consider to be math, that are still a fundamental part of CS... I suppose you could call those "computer engineering" but to me, that would be a dodge. Computer architecture and operating systems are pretty much computer science, and I wouldn't call them a specialized branch of math.
They've been introduced a huge variety of tools: vim, Git, valgrind, tmux and so on. A variety of languages: HTML/CSS/js, C, Java, oCaml. Platforms: Linux, windows, Android.
A lot of the stuff seems incidental. If you're going to mess with the Linux kernel, you need to understand C, you need some Linux CLI commands, and you'll end up learning a bunch of other things as well. You end up spending a lot of time debugging, profiling, etc, basically what anyone does when they're programming.
They also seem to learn at least the basics of how to write programs as part of a team. Code merges, planning things together, bits of scrum.
Shy of Universities hiring prominent open source software developers to serve as part-time teaching faculty, I don't know how we could push the transfer of these skills into the university setting and expose students to this knowledge at an earlier age.
I try to get this info to people by pulling batches of papers, skimming them, finding the worthwhile stuff, and mentioning it in online forums or to specific people in projects. I have over 10,000 papers in my collection from ACM/IEEE that were organized at some point (sighs). Anyway, when a topic comes up, I can just search them and (surprise!) many modern problems have already been solved. People just aren't applying the solutions when they know about them and more often just don't know about them.
So, the second one was my idea. A resource, crowd or taxpayer-funded, that's essentially an organized encyclopedia of papers and software developed in academia pertitent to aspects of software development. It would start as a manual effort that pulls in the links and paper that represent the best of most topics. As it got buy-in, the professors might have their students categorize and submit their own stuff. I thought an invite-only forum for serendipidous discussion among people that know their shit (esp researchers, professors, industry vets) could also get ideas across. Might be a fee for all of this charged to users, universities, someone that's way cheaper than ACM/IEEE. Basically, enough to cover the servers, bandwidth, and at least one employee.
Ran it by Epstein at NSF. He thought it was a nice idea but wisely warned it might not take off at all without buy-in ahead of time. The buy-in would have to come from a ton of places, too. So, bootstrapping that might be tricky. Even worse, many people I talked to in academia indicate the people and organizations on that side don't care about practical things at all. The ones that do are an exception and the norm actually filters out that stuff in favor of publishing/citations.
Well, we could hit it from the other side. Closest thing I've seen from industry was Dr Dobbs with articles on Cleanroom, Design-by-Contract, etc. Many professionals read that and could get exposed to cutting-edge via practical treatments from a trusted source. Additionally, the likes of Microsoft, IBM, Google, and Facebook do the cutting-edge academia and practical stuff simultaneously while publishing some and deploying others. This brings attention to it. Some results at these companies got widely deployed and are standard fare today.
So, I'd say the easiest route is a hybrid between companies and the nonprofit concept. Anyone wanting to try it can look at the kind of stuff companies are working on. Package up the cutting-edge stuff for those things to make it the first deliverable of the site. Add in plenty of IT and INFOSEC stuff on top of that which has immediate value for them. Get those companies as members to sponsor the initial development. Simultaneously, try to get University groups on the bandwagon who operate in key areas of IT or INFOSEC where the research results need attention. For big ones, maybe let them pay to have a dedicate person doing nothing but pulling in and curating their research as a brand boost for them. The site/service should grow in both usage and material as businesses pull on it for info while Universities push theirs into it.
Outside an effort like this, I'm not sure how else we're going to bridge the gap between academia and industry in a major way.
Balking at the $200/year ACM cost wouldn't signal to me that you're a professional (in the First World, then again, they have Developing World discounts). IEEE, though, is $40/month for 25 articles with a max of 10 unused carrying over to the next month, $17/month for a "Basic" level of only 3. Plus membership; that's prohibitive.
However, even limited explorations found great stuff that solved or might solve problems people talk about. An example we see on HN a lot is storing data safely in the cloud. There's so many ways to do that in academia with prototypes and detailed analysis it isn't even funny to see someone ask that question. Yet, they won't see it unless they knew their University had ACM/IEEE access or they paid for it themselves. I think this is a huge problem given how much of a gold mine it's been for me.
Good news its that CiteseerX and university web sites have many of the papers, too. Just gotta type keywords of the title, their name, and pdf (or PS) into Google to see if it's there. Sometimes Wayback Machine is only source. Great stuff disappears every year without getting mainstream attention. It's sad.
Anyway, might help to get affiliated with a University or enlist the help of a student. ;)
It takes a super-genius to make is simpler.
Sounds like an inherently complicated situation given all the parties, differing incentives, and legalities involved. I made an attempt in another post but not sure what will work.
For example, software developed for commercial avionics systems in the United States follows the guidelines in DO-178C:
https://en.wikipedia.org/wiki/DO-178C
Every stage of development from writing requirements to coding to verification to selection and usage of tools (including compilers and operating systems) is reviewed across a mixture of internal reviews and external audits. There's an extensive paper trail with each subsystem (i.e., software + hardware unit) being tested and enduring a certification process before being released to the public.
So while the engineers on these projects aren't licensed, the development processes and the resulting software are examined and signed off on.
Does that count for software engineering? I think it does. The next interesting question, in my opinion, is, if that counts as software engineering, then what aspect could we take away or relax, and still have software engineering? Is there a graduated spectrum of software engineering? Or is there a hard cut-off point, where we can say, "yes, this is engineering, but that isn't"?
I suspect that reason that such engineering process isn't applied to more software is that we don't really care that much. Are we willing to accept a 10x, or 100x, increase in development time and cost in order to have a more robust iPhone game? Word processor? To-do list?
Which is why I'm interested in the question, what can we take away or relax from avionics-style software development to bring better "engineering" practices to other areas of software? Maybe a 2x increase in time and cost would be acceptable...
The definition of engineer is NOT someone licensed and liable. I'm not sure where you got that idea.
The definition of engineering is the application of science and mathematics to solving everyday problems.
Could you be more specific then about what part of following DO-178C guidelines and FAA certification process you find inadequate, such that, "no such thing exists for software"?
I do not mean to belabor the point, but am genuinely curious what you might find to be an acceptable level of rigor for software. [EDIT: or perhaps you literally mean that no such level could possibly exist?]