Why Programming is Difficult (2014)
joearms.github.io
joearms.github.io
Any concrete coding task can be dealt with straightforwardly enough, but projects as a whole tend to deteriorate by arbitrary fixes, maintenance patches, and failure to uphold a conceptual schema.
The idea of composability is a good approximation to what it means for a program to make sense. If any part of the program can be understood as a sensible composition of parts, then the program probably makes sense.
But so many parts of real world code bases seem to end up incoherent. You look at a function and instead of seeing code that obviously, say, maps the render function over the list of widgets, you see three obscure conditionals, an apologetic comment, a bug tracker issue number, two boolean parameters named "force" and "dontRenderLast" (for some reason), and so on.
This is why I think the ideas from "domain-driven design" are so important, including the notion of a ubiquitous domain vocabulary. Also the FP idea of denotational semantics. And that well-known benefit of unit test driven design.
There's not enough recognition and clear understanding of this problem, in my opinion. We all know the problem, but somehow we don't talk about it enough, and don't acknowledge it as an enormous problem for development speed, programmer happiness, agility, and so on.
It really slows you down though (good thing?) when trying to dump your ideas into code, especially if you end up writing exploratory code that you later delete (I do this a lot).
I think literate code could be really great however for refactors. Once you've figured out the correct implementation and have the time & hopefully budget to revisit your fast code. Maybe it won't feel like the chore of straight up documentation, and more like the enjoyment of polishing your art.
I'd always rewrite what is left of such an effort from scratch, but it is easy to do so with all the literate brain-dump surrounding my code.
I pair quite a bit these days, and this may compliment our process nicely.
A contributing factor is organizational: There can be a managerial or "coordinator" layer composed of people who have neither the domain-knowledge of end-users nor the technical-knowledge of programmers.
Unless this layer is composed of very good people, what you get are specs that resemble a game of "telephone": Even if they're detailed with mockups etc, what the specs actually encode is what Alice thinks she needs to ask for in order to satisfy the request of Bob which is based on him thinking about what design might solve his problem, etc.
The actual way the data is used isn't particularly high-performance or massively-parallel, so it's all about fencing-in crappy logic and avoiding zillion-phase commits.
There are a whole heap of "human nature" problems that are the same like getting into debt, being over weight, protecting the environment. The short term pay-off is fairly big, the contribution to the long term problem is very small.
Also the person who is often responsible for making the decision may not be technical, so finds it difficult to imagine the eventual impact. Plus in the fast moving field of software development, they'll be unlikely to even be around when it happens.
Business people are very familiar with terms like "infrastructure", "extensibility", "maintenance" and the like. They're not idiots. May be you should try explaining why refactoring and good architecture is good in business terms instead of programming?
But also, I think there has to be some conceptual basis for us to be able to think about factoring in a clear way, and that kind of intellectual labor is hard on another level.
Eric Evans' book on domain-driven design has some examples of programming tasks that generate confusing and unclear code, until the coder together with the "domain expert" reaches some kind of epiphany—like "ohh, if we think of these three things as being parts of the same kind of thing, then this part of the program makes more sense."
A program is a kind of philosophical theory about a certain topic, and if the theory isn't very clear, then you'll end up with a lot of special cases and obscureness even if you have unit tests and refactoring tools. They say naming is the hardest thing in programming and that becomes more true when you start to introduce any level of abstraction.
"Design patterns" were an attempt to create more general small theoretical pieces, above the syntactic level. In the FP world, a similar vision is that algebraic and categorical abstractions can provide some of those pieces. We still haven't seen much of what algebraic abstractions can do for ordinary programming...
But these abstract pieces, I think, are pretty crucial if we want programs to be understandable as anything but arbitrary collections of procedures and types. So yes to refactoring tools, and yes to more discussion about design patterns—it didn't end with the Gang of Four! read Christopher Alexander and Richard P. Gabriel!—and yes to more inspiration from mathematics and engineering and philosophy and all kinds of stuff, to provide inspiration for factoring—yes to study of logic and language and ... I'm getting carried away.
1. https://en.wikipedia.org/wiki/Lisp_%28programming_language%2...
Only if those refactoring tools could refactor the domain itself (and the environment), not just the code :-)
Often these special cases aren't intrinsic to the code, but caused by inconsistent business rules, or weird issues with the environment (some browsers not sending some events, workarounds for compiler bugs on ancient platforms that must still be supported, socket closing might or might not cause a flush depending on whether you're on Windows or *nix, ...)
Getting to a deeper understanding of what the old code is doing takes a lot of work. When we get to that point, refactoring to make what we learned apparent is a good idea. Assuming there is a way to test the changes. (Now I get to spend 2 days making "mocks" in environment X -- learning the mockup tool, and determining the content appropriate for the use cases)
It's a good point you bring up, but it's not a freebie. I am in favor of doing what you said, when you can.
Most systems I've worked on have some overriding architecture and assumptions. The worst quality erosions happen when changes are shoehorned into the code without a complete understanding of those factors: it's likely there'll be semantic mismatch between the new feature and existing components, unexpected bugs, and harder to modify code (because the overall architecture is no longer coherent).
You can mitigate the disintegration by having a maintainer with a strong understanding of the architecture and the problem(s) being solved by the original system. Talking with them can give you a roadmap on how to modify it properly, and if it's even possible to do so while maintaining architectural and design coherence.
Note: my comment is an opinion based on personal experience. I'm not trying to play it off as fact.
Somewhat the same arguments apply to approaches which put a lot of emphasis on small-scale/unit testing -- they offer reassurance when doing small-scale refactoring at the expense of making substantive refactoring harder.
Personally, I think this isn't inherent, but is just an effect of how all the current programming paradigms (procedural, functional, even logic) tend to compose features together by simply sticking more code into each node of the function graph. So code that was originally a single stream of simple definitions becomes a complex mess where each line could have come from a consideration a week ago or a decade ago, and you can't "follow the train of thought" that led to the current state any given function is in.
I think we need to think very differently. A truly aspect-oriented programming language, like Inform7 but not limited to the domain of Interactive Fiction, could allow for "literate programming" that treats a codebase as a linear narrative of feature additions, rather than a mutable graph representing current behavior. You'd actually be able to read a codebase from "beginning to end" and get a sense of what thoughts went into constructing it, in order.
Just reading git commit history won't get you there. Commits aren't the right granularity, for one thing—a log of merged PRs might work better. More importantly, git commits to a codebase are a mix of feature-commits and bug-fixing commits, which horribly muddies your vision of how each feature is defined. What you really want is for features to each be files, and bug-fixes to be commits to those files, such that you can read "the up-to-date edition" of each of the features.
Do note that we already have this in one particular place. This is how database migrations work. We just need code migrations! :)
That might be a good way of understanding the evolution of the system if you're already familiar with what the system does. But if you were a newly hired developer trying to understand the system from the ground up, wouldn't you want the features grouped together by logical function (systems and subsystems) rather than by their order of addition?
So, rather than staring at a function's git-blame trace together with commit history to try to deduce a reason for each of the (likely uncommented) lines in it to exist, you would instead just see that the function started simple, and then there was a feature-patch for feature Foo that added lines 1, 5, and 7 to the function, which all reference a global variable $Foo that was also defined elsewhere in the same feature-patch. Etc.
One of the big problems with "only having the diffs" is being unable to know what they add up to: usually, the only way to tell what "flat schema" a set of DB migrations add up to is to run them against a real RDBMS and then dump the resultant schema. But again, with a LightTable-like editor, I could easily imagine each feature-patch being presented together with the "generated context" of what the resultant function looks like at that point in the code's history—or what it looks like now—with that feature-patch in play. You could tweak the feature-patch, save, and the code in the other pane would auto-update.
---
† One interesting thing is that there would definitely be times where a particular subsystem got a feature added, and then removed when it became unnecessary or was obviated by another feature. The feature might persist in other places (so you wouldn't git-rm the entire feature-patch that introduced the feature) but rather than seeing a complete feature history for some subsystem Foo (+A, +B, -A), it would be useful to be able to get an "abridged" summary of Foo that just introduces feature B and pretends feature A never happened.
Both views would have their purpose; the abridged view would be more useful for first learning the subsystem, while the unabridged view would be more useful for understanding the codebase as a whole, since you could see why, for example, there are two logging systems in play—it would be clear in unabridged view that one is in the process of replacing the other, but there are some subsystems where the newer one hasn't yet been introduced.
This sounds like it shouldn't be an issue, but in fact it's a terrible state of affairs, because abstraction design - i.e. domain modelling - is in no way the same as process design.
It's why you get feature patches for Foo that add a couple of lines to handle some tiny little edge case. The patch works procedurally - as a little blob of disconnected logic for one specific case - but it isn't really part of the main abstraction model.
Repeat that a few times and you don't have an abstraction model any more - you have a library of epicycles and special cases, and the conceptual equivalent of what used to be called spaghetti code. Or a machine with a lot of spare cogs all over the floor: they're all needed, but they're not connected how they should be.
A system like the one you're suggesting specifically highlights intent and context, which can bring you back to the abstraction level and help minimise diff confusion.
It's strange that devs often have to build systems that create and handle metadata, but there's been hardly any research into making code more legible with supporting metadata and context.
The 'easy' approach is to split the codebase up into modules that can be versioned independently.
A more extreme approach would be to store the previous history away as a separate branch and squash related changes with the log referencing the commits that make up the aggregate.
Feature branches accomplish this to some degree and it's not unusual for OSS maintainers to request that added features get squashed but they also tend to dispose of the branches that contain the more granular history of changes.
It seems like building only the things that _really_ matter would be a solution. Then time would be more easily allocated to test driven design, composable architecture, state management, and data modeling.
The question then becomes: "what truly matters to build". That's a whole other discussion.
Or, it's epicycles within epicycles until a paradigm shift to a simpler theory. Though most real-world programming goals are not intrinsically simple, but a mish-mash of mess.
Deleted comment
Can you expand on this? What is "structured programming"? Why is it misguided to make it visual?
I am working on a tool that could be described as 2D functional programming, so these are not idle questions.
I think these structures are inappropriate in 2D because they assume a linear, 1D context for understanding.
>Structured programming gave us the conventional control structures present in most PLs: for-loops, while-loops, if-then-else statements, switch-case statements, etc. None of them are appropriate in 2D (but are very appropriate in a 1D program, usually written in text). Historically, structured programming was developed to enable the elimination of the GOTO statement.
However I don't understand how conventional control structures are necessarily inappropriate for 2D. It seems to me that the 1D nature of text is used to map to execution point; even if that point jumps around a lot, it starts at the beginning, moves down by one line unless told otherwise, and ends at the bottom.
In my tool, execution order is determined by dependencies, which is fine because there's no mutation (not by code anyway) so order doesn't really matter. So some constructs---`for`, `while` and `until` for instance---do become obsolete/meaningless.
If and switch, however, remain very relevant. For that matter, switch draws its name from real-life switches in real-life machines, which run on ultra-concurrent physics. But we still find a use for them.
Have I mis-stated or misunderstood you?
As for the 2D similarity of 1D languages, that certainly sounds interesting. By "presenting them in 2D," do you mean something like Blockly? Blockly kind of leaves me cold because it seems like normal code, plus extreme syntax help. I mean, cool and stuff, but we can dream bigger, you know? Anyway, do you mean that imperative, functional, and concatenative all have the same basic made-from-text look when put into something like Blockly, or is there a deeper point I'm missing?
Thanks a ton, btw. I'm always interested to hear what people think on this subject.
That's just my opinion and I don't have a good justification for it other than to point to a project like Google's Blockly as an example of how it seems to work poorly in practice.
Good luck with your tool! I'm happy people are working on projects like that. :)
However, programmers more than other professionals love blogging about their work, which is great - because it helps spread ideas (more so than any other field I've seen), but also creates this illusion that programming is this super hard thing. I mean people complain about going to interviews and being asked about algorithms and stuff that you could "just google" when, if you were to ask an average scientist, the amount of active working knowledge they have at any given time is simply astounding.
There are lots of jobs that are harder than mine in some sense, but if you look at the objective of building high-quality complex software in the minimum amount of time, keeping it running correctly, and then being able to change that software as effectively as possible, then that in itself is pretty much an infinite problem. Maybe there is lots of scientific work to be done to help us struggling coders, and maybe we need to get our shit together more as a profession... It's not all easy.
The same can be said of almost any "high-quality complex" project in any other profession.
As an example, I started doing Angular development about a year ago. At first, it's really easy - little more than templating output expressions in double curly braces; binding input fields with ng-model attributes; and adding ng-click callbacks to a few buttons.
Then you start making tag libraries, er, excuse me, "directives". Now you start to find out where the monster lives. The complexity of the definition object you make for a directive mushrooms out from the other work you have been doing. Is the directive just an attribute (for existing HTML tags), or does it need to be a custom element with multiple attributes? Is the template part just a string literal, or a function? Do I need to make a "link" function (what does that name even mean, anyway?), and how did all these magic parameters get passed to it? Does this directive get its own controller function? (why do we need "link" and a controller?) Do I give the directive its own "scope" (data model and event-handler/callback function object) instance, or just reuse the scope from the app? How do I pass data between 2 directives?
I'm not faulting Angular (other than maybe for better docs) - I think every framework has that issue of making 80% (if it's a good framework) of some problem easy, but sweeping a real mess into some corner for other work.
I'm waiting for all these custom syntax and conceptually burdened frameworks to implode :)
If you like pure js, but declarative templates, check out domvm [2]. I wrote it to improve on Mithril and domchanger. Feedback welcome!
It's funny how some fields get the difficult label, and some fields get the easy label. I've always felt a good mechanic, whom can work on any aspect of pre obd2 vechicles, and has a good grasp on today's computers on wheels, don't get enough respect.
I belive it comes down to how a subject is taught, and complete honesty. I'll pass this along. Watch repair is not that difficult. In about a year's work of time, and the right tools; the average person could repair watches. Yes, there are odd balls out there, but most high end watches are still using the basic escapement to regulate that train of gears. You open one, and the next one looks familiar.
I do find programming difficult though. My front end work, I don't find that difficult. I've noticed each year Programming tutorials just get better. The books get better. I guess I gave the Internet to thank, and honest/benevolent folks who don't mind spreading information?
I think there's actually some speculation it's the opposite. I realize there is a lot that's not yet understood, but it seems to point in that direction. See: https://hbr.org/2015/10/its-harder-to-empathize-with-people-...
" First, people generally have difficulty accurately recalling just how difficult a past aversive experience was. Though we may remember that a past experience was painful, stressful, or emotionally trying, we tend to underestimate just how painful that experience felt in the moment. This phenomenon is called an “empathy gap.” "
There are low-level programmers who do things similar to other sciences involving math, and my brain wouldn't be able to do that.
So if you're talking about scripting, which I think most people do, then I agree it isn't difficult.
Mathematics and algorithms aside, there are vague requirements and tight deadlines. There is a need to estimate how long will it take (know of anything else harder than this?). There are maintenance considerations which complicate everything, even a one-line script: documenting, packaging, future-proofing against platform updates, etc.
Many complex phenomena in nature is explained by very simple mathematical equations, likewise great engineering design is achieved when there is nothing else to remove, and in order to achieve those results, a great deal of effort and thinking has to happen first.
The thing is, as the author put out, the environment is just too hostile in order to design beautiful programs, from the insane expectations of what can be done, to incompetent peers, politics, distractions and what not.
If you are skilled, you can define a problem in a precise way, be able to dissect it in smaller problems, solve them and prove the solution correct.
What is ultimately hard is how to handle the hostile environment you're usually stuck with.
Here is what I mean. Implementing new functionality where nothing existed before is easy: just figure out what you want the machine to do, and tell it to do that. It isn't trivial, but on a relative scale, it's easy.
Precisely because it's easy to create functionality, we do a lot of it. Our programs grow. As they grow, they get complicated. As they get complicated, it gets to be more and more difficult to understand what they're already doing so that we can make them do something different. So, the less tongue-in-cheek version of my thesis here is that programming gets to be difficult because it starts out being easy.
This is why I always say that the essence of software engineering is the management of complexity. And this is the part that people who have not actually worked on large programs will probably never understand. (As someone once said to me, "what you need to get is that management views programming as primarily a clerical function".) Viewed one-by-one, the tasks seem simple. And yet, somehow, when you put them all together, they're not simple anymore.
Your point is valid, and you can delete the word "software" and make a more general point that is also true.
I love it when I go to bed mulling over a problem, and when I get up I have the solution. The brain is truly a wondrous machine!
At my last job, I think the only time that I had ever seen our programming team actually happy to work on a project was when by chance and via a minor mutiny, we were able to break them away from the minutia of constant maintenance and repair and let them actually just work on a project with no expected outcome. It was an alternative path for our identity management solution and the entire project was a challenge to the managers for the Enterprise team. Our programmers were down-right chipper at the prospect of just being able to flex their creativity and try something just to see if they could do it.
Begin tangent:
I work for a university doing grant projects for a few state and federal agencies, though. Some might like it, some might consider it hell for an entrepreneur. It works for me :-)
We are in the process of re-writing our "flagship" app from 10+ year old java libraries. We were able to dodge the bullet of "upgrading" (???) from Struts to JSF. The new app is going to be a services based (REST + JSON) Java (mostly) back end, with an Angular based front end. We just did a demo of a subset of the new app at a trade show where a 3rd party app imports a partial record (via web service) into our app, then launches a nested browser window to run the rest of the relevant data entry. (the rest of the workflow will be handled by the legacy app until the complete new version is finished)
No, I couldn't persuade my coworkers and supervisor to discard Java entirely and go with Node.js :-) I'm not entirely sure I would be comfortable with that, either. We are doing some server side Javascript lately, though, by using the Java scripting engine interface.
We will likely soon be redoing another (30 or so year old) app from a sister university that the state uses, as well. Interesting times!
For almost 5 years before this, I worked at a financial services company. I got paid a bit more (bonuses), but I spent all my time, really, keeping 2 legacy applications going without ever getting to spend much time REALLY updating some sore spots.
Also I find open concept offices terrible to work in. I'm not the type of person who can put on a pair of headphones, blast music and pump out code. I like solitude as I think though the problem and code. Unfortunately for me, the open concept is here to stay it seems.
Services like parse cloud code and AWS lambda are another step closer.
Context: We (me + couple of collaborators) are writing a similar platform for ourselves, although the purpose is dogfood more than $.
Lambdas tend to be < 100 lines of Javascript which means maintenance and rewriting is trivial. And because the jobs are so simple people don't need to be Javascript experts or aware of the JS ecosystem to write them.
Lambda is also very cheap if you have bursty or just very low-volume workloads.
Perhaps it's even a part of the (now) open source parse-server?
I am not sure, but as I am building a similar thing I will definitely be having a closer look at their source code in the coming weeks.
But there is a lot of game changing practical advances which totally overturned the way we design systems.
Microcontrollers are powerful and dirt cheap now. Low end FPGAs are more accessible than ever. GPGPU is a dramatic game changer, and it is dead easy to use it these days. RAM is nearly infinite. SSDs are cheap and fast. 1gbps networks are everywhere. CPUs have dozens of cores. So, yes, on a practical side everything changed.
Ten years ago, if you said you were into functional programming, most people wondered if you were crazy and why you didn't just program in a real, normal language.
Fifteen years ago, garbage collection was considered slow and mostly only suitable for "scripting."
Twenty-five years ago, there was Haskell, Erlang, Common Lisp, Smalltalk, and so much cool technology and visionary ideas. For most of the time since, it's all been considered weird, fringe, academic stuff.
The past decades of industrial talk about best practices and methodology doesn't seem to have been wildly successful, and now we're looking at other ideas in a diverse way. So it's an interesting time.
"Why would you want to trace the dataflow through the routines, as a pipeline, rather than just whacking on the DATA DIVISION, er, beans / instance variables, as needed???"
The 70s brought us computing, the 80s brought us networking, the 90s brought us eCommerce, the 00s brought us social and mobile. I don't think anyone has a clue what the next broad stoke is.
So I'm talking about the landscape of development, if you will... Paul Graham's "Beating the Averages" is about how he and Robert Morris knew back in '95 that in a web startup you could use whatever language you wanted, and I think that came into play in a big way in the '00s. I speculate that Ruby + Rails had a pretty big impact, and then also there was just this generational wave of internet nerds, and a big upswing in open source infrastructure, and some other factors...
Hey, while I'm speculating! I think the basic shapes of collaboration-type apps are kind of settling to the point where technologies will come out that eliminate much of the coding drudgery.
This is an old dream, of course, and it comes in cycles—like, maybe a truly excellent framework for modern real-time web apps will stabilize, and the year after everybody will be in virtual reality and the whole paradigm changes again.
But if you look at the structures of "social and mobile" apps, there are some common denominators that we still reimplement tediously all the time. Users who have relations with each other and other entities, via various permissions. Distributed resources modified with different consistency properties. Timelines, searching, embeds, comments, notifications, and some other things.
So what if you took a bunch of clever people with experience from projects like reddit, SoundCloud, Instagram, Airbnb, Slack, or pretty much any social thing, and gave them a couple of years to dream up an architecture that would make all their jobs easier, and that they could maintain as an open source framework? Maybe throw in some theoretical experts to help tease out elegant abstractions and semantics.
I think something like that would make sense for YC Research, even if it's not as radical as, say, extending life spans. Because how much of the YC-backed developer workforce is right now sitting around hacking their own thing for real-time notifications? How much more efficient would these startups be if they could configure a mature system declaratively and get the basics, cross-platform, like you could get the basics for a CRUD app with an hour of Rails configuration?
Anyway, that turned into a long semi off-topic rant. It's just one incremental step I see on the horizon.
Disclaimer: Java programmers get jobs, so I am one. My feeling about it is "The elegance of C++ at the speed of a Lisp". Fortunately, we are starting to get better language front ends which can reuse the old JVM binaries, as well as having a GC-ing VM that has had decades and $millions thrown at it.
It's been a problem longer than that, and the culprit (ironically) is Fred Brooks. Yes, the guy we hand the "I didn't get a Turing award, all I got was this lousy Fred Brooks award" T-shirt to people.
If you don't know, Fred Brooks was the manager of the first large scale programming project at IBM that actually tried to use a programming language, as opposed to machine code/assembler and the toolset and paradigms people used with that approach.
Anyway, long story short, the project was deemed a "failure", and this Fred Brooks guy, presumably in a bit of ego-damage control and to save his professional reputation, did something interesting: he claimed that anyone would have failed, not just him, and that there is no possible solution to the problem that fundamentally fixes it. (Literally: "there is no silver bullet.")
Well, Brook's knee-jerk proclamation about killed all research into the "programming is hard" problem at the academic level. People now strive for x% improvement but firmly believe--after literally one bad experience--that fundamental improvement is impossible. But maybe, just maybe, we're "doing it wrong" as it were and we were wrong to give up after literally one bad experience by a pointy-headed boss in CYA mode on a major project.
I firmly believe that we're doing programming wrong, that Fred Brooks was (and is) wrong, and that 50 years from now, we'll look back at a million monkeys (referring to myself) typing out text to the computer and laugh at how primitive we all are today.
A major why waiting to the end to test document and clean up code is a failure.
Write beautiful, tested code from the start. Even throwaway code, because half time prototypes, etc end up production.
What makes sense in testing and documentation is really project dependent. What's right for flight control software is very different from what's right for Yet Another Twitter Clone.
User docs are completely outside my realm as a software developer.
My point (and strongly held belief) is NO, when you do testing and documentation is not really project dependent. All projects need to do it from the start. Some, flight control software, need more rigorous process. Everything needs to be viewed through ROI. My contention is too many devs underestimate the return of "doing it right" from the start. and overestimate the effort and effectiveness of doing it later/at the end.
Even the simplest problem, like a Purchaser placing a Purchase Order for a Product from a Supplier can't be easily modeled and programmed by the majority of software developers - with any sort of predictability based on prior art or their own experience.
At best, we approach a problem with a combination of modular procedural code (modularity via subroutines)(from the 1950's) and structured programming with functional decomposition (from the 1960's).
There is no shortage of technical prowess, but a warehouse system is not based on design patterns, libraries, frameworks, or technologies. Instead, it is based on people, places, things and transactions interacting together to solve a problem.
How to represent real-world concepts and systems (real or abstract) into code, remains the challenge it always has been. We just don't seem to learn any lessons from the past.
Great article though- I think it applies to most technical roles.
whaaaa
Sadly, what too much code tells us is "Design? That's what gui people do, right? I just kept typing stuff until it worked enough for non-technical person to think it was OK or we ran out of money/time."
Those are all basically the same language; if you think only syntax differs, you haven't tried enough languages that are actually different. Try a Lisp, a Smalltalk, and a Haskell and you'll learn they're semantically different, not just syntactic.
if you're trolling, that's a good one.