Not that I want to discourage you, but my view is that anything that makes software engineering simpler just leads us to tackling more complex problems until the complexity reaches the limit that people can handle.
So in that view, you can't succeed at making software engineering always simple. Instead, you can make previously intractable problems tractable.
Frameworks approximately do this, but for specific domains: they let you organize complexity into well-defined areas and put a pin in them, easing the cognitive load. Then you can handle those abstractions more easily. But yes, to your point, because frameworks typically only tackle a few levels of complexity, you still get that complexity back when you use the framework to advance the problems to the edge of what your framework is designed to address.
A recursive/fractal management of complexity will allow all levels of the hierarchy to feel similar, so you are never increasing complexity, only looking at a different resolution. I think the key to this is mapping out the fundamental organizational problems and how they relate to each other at different resolutions.
This is the 'induced demand' argument. As perfection seems impossible, and mistakes inevitable, this seems likely.
However, the more I look into issues, the more I realize that a very small number of errors introduced early on is what ultimately causes a plethora of them. Just fixing a very small amount of these mistakes would have untold effects on computing over the long-term. That is, if we can get over the initial switching costs.
> small team of engineers
It took 2 men to over-take thousands creating Unix over Multix. It took Linus Torvalds - just about alone - a week to create git. We have already seen this prophecy come true.
As teams get bigger, communication sales factorially, and the more people you have the more mistakes you make, which increases the size exponentially with each further mistake. What Unix and git showed is that when you put everything into a small team of engineers heads, they can work through the complexity enough until they can do it themselves.
> Frameworks
One of the things I realized a while ago is that pure/impure and library/framework have a decent mapping between the two. If you have a framework, you give it code and it acts for you, just like impure code. And so the problem we keep running into, is instead of inverting the flow like 'hexagonal architecture' says, we keep piling impure onto impure onto impure. Hexagonal says to not do anything of substance, keep the adapter clean, but each framework sure is doing something. Each layer on the stack we go up, the harder it is to get down. So now we have OS and applications and containers and k8s running micro-services that ends up being run in a web browser, when all we really needed was microkernals.
That sounds a bit drastic. Surely at worst the scaling is quadratic?
The ability to suss-out meaning behind those channels and come to a shared understanding pushes it to factorial. If A is speaking to B and C, A needs to also think about what communication is happening between B and C, and this is different to what B knows about C and C knows about B.
In a carefully balanced classroom you might manage to get everyone knowing the same things, and this will allow knowledge to scale. Think about how much better TV got once they knew that viewers would watch every episode. What happens in software is that specialization quickly comes into play such that this is impossible in any organization that doesn't use mob programming. Other people might as well be speaking a different language.
Next time you are in a Sprint meeting, think about how much you don't understand. Even as the team lead or architect who designs the entire system and whose job it is to understand everything it will be a shocking amount - it's black boxes all the way down. You'll claim that you can't know everything and that anyone who tries will fail.
And this is made worse as the only true way to learn something is to do it, and if you aren't actually challenging yourself on something, anything you learn without doing will just fade away - as Spaced Repetition shows.
The bigger the team, the bigger the software, the more inevitable the collapse.
Metcalfe's law is quadratic rather than exponential -- O(n^2), not O(2^n). And if n people all need to be aware of communication happening between every pair of them (which is a worst-case), then that should just add a factor of n, bringing it up to cubic.
(On an unrelated and less pedantically nitpicky note, one of the most valuable professional skills I've developed is an ability and willingness to dive into black boxes. It's remarkable how many bugs arise from the interaction of two or more pieces of software with no overlap between their developers. Debugging those can involve a long, agonizing back-and-forth between people who know how X works and people who know how Y works -- or it can be over in a jiffy if you can quickly familiarize yourself with the basic internal workings of both. Past a certain point nobody can know everything, it's true, but this doesn't need to be as crippling for productivity as a lot of people allow it to be.)
But that doesn't stop black-boxes from being created faster than I can fix them, unfortunately.
Multics did not have thousands of developers. Linus was just more practical about software development, was maybe better at building a community, and more importantly focused on x86 commodity HW which was a bet that only Microsoft also happened to make.
Thanks.
Actually, "combinatorial" or "exponential" would be better words to describe complexity explosion. The number of possible cases/scenarios/logic flows/etc grows exponentially as a system's complexity increases. It is a curse and a blessing at the same time keeping so many SWEs gainfully employed.
In a lot of human work (not just software, but any endeavor involving technology even as "simple" as building a tunnel or simply involving many people coordinating together), complexity emerges this "unknown unknowns" property. This has happened since time immemorial. If you don't believe me, go master how to live out a subsistence life in an agrarian area of $nation during the 1500's with their tools but modern knowledge, then try to convince yourself complexity did not exist back then.
All still manageable by humans but at continually higher and higher levels of complexity.
you seem to have thought about this a bit did you have any ideas of how things could be rearranged to a more recursive solution?
Similarly using an operating system/assembly is much easier than trying to control pulses of electricity yourself!
Just an example: around here, most people have a first (given) and a last (family) name. If I don't model that as separate, I have trouble interfacing with other software. If I do, I have trouble with people from other cultures that don't follow that convention. Storing both risks the data going out of sync. What's the "right" way to store person names? There doesn't seem to be a simple solution.
Another example: we model physical cables (both for power grid and for data transmission) in our CMDB. All works fine, until you suddenly have a Y-shaped cable with three connectors that doesn't fit into your data model. The real world always has these 1% of cases that don't fit the general pattern; if you focus on the 99%, the 1% make trouble. If you focus on modeling every case, you have 10x the complexity, even for the simple case.
And then there are things that are moderately complex and security critical, like password recovery workflows. We haven't really found a way to reuse these among different technologies. Like, if you once figured out the perfect password reset workflow with Ruby on Rails, and your next job uses Python + Django, you're back to square one.
If somebody has a good idea for how to tackle these problems, please let me know!
But where software does fail in my humble opinion is making it easy to pull in tried and true tested solutions to the problems that we do face, even if they are not as common. Because even though they may not seem common, I'm absolutely certain many face the same scenario.
The amount of duplication solving the same problems is insane. But this is not an easy problem to solve and I don't intend to trivialize it.
I think we need to come up with better tools. code sharing through libraries/repositories (like npm) is great, but it can't be the final solution.
Back in the day in Haskell I dreamed of a system where you could type out a type signature and a fully tested rated implementation would be imported from an "open source" service. You could import modules, functions, data structures, anything. But that vision is still a long ways off.
I think of that as optimizing the entire stack for what the customer wants of rapid/high-speed analytics.
The type signature of GPT-3 is `string -> string`
But if you look at like co-pilot for example, if that was given the ability to have type signatures serve as input you might get a lot more powerful results than what it does with raw text (which is very impressive).
But this comes down to type signature design. You can encode any function using simple types like int -> int which aren't very useful. Where Haskell shines is when using types to limit the scope of inputs & outputs. What I am getting at is that you can still write uninformative type signatures in Haskell, but it also gives you the power to write more informative ones.
I don't think Haskell is the answer, so please don't take that as what I am saying. I do think however using richer type systems could be a stepping stone towards a solution to this problem.
I feel like it wouldn't work if you could actually mathematically constrain the outputs to be syntactically correct either. That's probably one of those Gödel things.
Type signatures don’t tell you which should be the “then” clause vs the “else” clause in any conditional.
Don't try to force schemas onto schema-less data. Store the "name" as a JSON string/blob representing the various possible attributes (given, middle, family, title, etc) and provide a variety of functions for representing that data. IF you really need to do this at all (for an internal app, you probably don't).
> Physical links
Include an Hardware Asset FK in your Link M2M table. Model each binary link explicitly, so a Y cable = 2-3 different Links that point to the same cable Asset. Or you can have single Links with M2M inputs/outputs. But definitely don't model Y cables explicitly.
> Password recovery
What is so complicated about this, specifically? Python and Ruby are completely different languages with no guarantees for interop.
If you can't do that, you have more pressing issues.
Where are you getting "a dozen assumptions" from?
Our cables have stickers with cable IDs, which are stored in the CMDB (and can be used to trace the endpoints through patch fields, for example).
So if you model a Y cable as two physical cables, you have to both allow duplicate cable IDs AND multiple cables per port. Both of these have significant downsides in preventing data entry errors.
No. Model it as two abstract "Physical Links" or "Connections" which FK to the same physical cable Asset, which is uniquely determined by its inventory ID and/or manufacturer model+serial.
If a Link needs to know whether its part of a Y cable Asset, it can join the Asset table.
If a Y cable Asset needs to know what it's connected to, it can join the Link table.
Additionally, you have a nice clean Link table for doing graph traversal (get me everything connected to this impacted Asset by degree <= 2), and a nice clean Asset table for inventory tracking.
C and C++ libraries with light language-specific wrappers largely serve this purpose, in practice. It's plausible that, say, a PHP postgres client lib and a Node postgres client lib will share nearly all their code—probably as C or C++.
This doesn't get leveraged much aside from interfacing with daemons and sometimes for extremely complex e.g. media or crypto libraries, though. No-one's doing this for high-level software workflow building-blocks, like a password reset flow.
[EDIT] as for the "why", I suspect it's because the things it's used for are far simpler than the things it isn't. A password reset flow can potentially need to interface with lots of different things, some of which may be custom to the project, some of which may vary with run-time input, et c. If you cut it down to only the parts that could truly be re-used anywhere, with all kinds of points where you can hook in as needed... you've only made about 5% of the work re-usable, so it's just about pointless.
The Common Language Environment on VMS, or its lesser(but infinitely more available hence virtuous by effect) imitation, dot net CLR . FFIs, if you bought a good environment. Sounds like if open source implementation of good foreign interfaces existed, the OP's problem wouldn't exist, let alone system level interoperability.
Store both a unified name field and separated given name / surname fields and let the user manage both. Yeah, this risks going out of sync, but that's the user's problem, not yours. Yeah, three fields are technically more complex than one or two, but it produces less complexity down the line.
Falsehoods programmers believe about name: https://www.kalzumeus.com/2010/06/17/falsehoods-programmers-...
Then they leave that field empty.
https://www.kalzumeus.com/2010/06/17/falsehoods-programmers-...
As a practical matter, the HL7 FHIR data model for names can work well enough for the vast majority of applications in almost any country.
struct Screenname {
Matches<String,regex"\w+"> _;
// TODO: ensure it is illegal to ask for users' real names
// for now we'll just do the right thing unilaterally
}For how to model physical objects like those cables, go see how Grainger does it as a working practical model.
To really handle names properly, you need more context than the name in the presentation layer that many schemas take their modeling from can obtain. Government health care or similar widely-adopted encoding is sometimes Good Enough. If you want non-lossy exactitude however, then that's a much bigger scope (I'd be investigating a first-pass classifier with contextual hints taken from various geolocations, age, etc., that implies soliciting the name comes after what you normally solicit for input, and refining from there).
Password reset; this is why vendors like Okta exist to abstract authN away for us, and auth0 for authZ. Then there is the rabbit hole of what this abstraction leads to...
[1] https://www.kalzumeus.com/2010/06/17/falsehoods-programmers-...
The way past bureaucracies that didn't have the luxury of unlimited cheap complexity did: you set some rules and anyone who doesn't like them can either suck it up and follow them, or deal with the consequences themselves. Allowing users to specify arbitrarily complex requirements and never saying "no" is what gets us into this problem.
He said that it used to possible, practically speaking, to treat software like discrete components so that you can reason about how to combine the units and predict the result, etc. By contrast, much programming today is bodging together a bunch of libraries whose behaviors are often poorly defined or otherwise opaque so that one often enough has to treat them like black boxes against which you have to apply scientific reasoning to understand how to use and integrate it. Hence Python and robots.
Presumably the new course teaches skills around managing this complexity, but the shift in model advocated -- from the metaphor of well-behaved discrete components in composition to opaque units whose inputs and outputs (possibly depending on hidden state) must be discovered was interesting to me, and ultimately adds to the accidental complexity of doing software now.
Do you think this sort of problem may have a distinct visual representation that would fundamentally reduce the complexity of library-dominated development? Or are you imagining a whole programming ecosystem built around your idea, so that the fundamental problem Sussmen described is somehow obviated? Or are you really trying to address essential, not accidental complexity?
1. becomes a post from which clever and ambitious people can build complex things. You invent docker to simplify application deployments and then somebody builds a n-dimensional microservice cloud on one side and starts commercializing new hardware architectures on the other. You don't remove the complexity, you just let it move around into new domains. (not a bad thing!)
2. is more temporal than you expect. The "finite and fundamental" problems of yesterday, today, and tomorrow are of different sets -- partly because (1) opens up new problem classes and partly because most problems are inescapably cultural and therefore subject to fashion cycles.
To tie it to something concrete, just about everyone thinks C++ template metaprogramming is unreasonably difficult to reason about even aside from the syntax, but it exists to be extremely powerful at precisely expressing behavioral contracts of a software component that are difficult to do in other languages. Even then, it barely scratches the surface of what would be useful to express for the interaction of those components to be "simple". The number of design parameters of a component that really matter in various contexts are astronomical. No one can deal with reasoning about that many component traits in practice, so the software engineers simply hide most of them -- the complexity is still there and will manifest in unexpected ways because it is not visible at the interface.
All systems engineering is complex for this reason. Any non-trivial software system has to be reasoned about as a monolith at some level to correctly address complexity, which creates an unavoidable cognitive load. This isn't a software engineering problem, it is a systems-thinking problem. In chemical engineering, for example, there are often complex system design problems that cannot be adequately addressed by decomposing them into a sequence of sub-problems, all that does is hide major complexity at the interface of the sub-problems.
Could you give real world examples of:
* "In reality, components have complex interactions far outside of what is captured in the component interfaces"
* How template metaprogramming "precisely express[es] behavioral contracts... that are difficult to do in other languages"
I think you contradict your poetry with "Any non-trivial software system has to be reasoned about as a monolith at some level to correctly address complexity". If _any_ non-trivial software system _can_ be reasoned about as a monolith then it is true there exists some model with sufficient simplicity that there are well-defined discrete components.Traditional engineering disciplines that routinely have this problem of non-decomposability actually treat it as a systems problem (or a supercomputing problem), they don't shy away from it simply because the cognitive complexity is difficult. They convert the entire system into a set of simultaneous equations that needs to be solved for. In software, we would call this designing a monolith, but the reason it is done in traditional engineering disciplines is that you can't wish away the fundamental nature of the problem. In software, because it is not a physical engineering discipline, you can pretend that this issue doesn't exist, for a while at least.
Systems complexity is intrinsic, there is no trivial reduction. If there was engineers of all types would serve little purpose. Not coincidentally, the engineers that command the highest salaries are those that have the cognitive capacity to reason about the most complex systems dynamics, the dynamics that can't be decomposed into independent sub-problems before understanding the behavior of the entire system.
The fundamental problem of simplifying software is humans. Just consider date formats, time zones, and tax codes. Humans love to make complicated things.
My philosophy on this has been to take the complex human stuff and stick it in a black box. A professional feather in cap with this approach is called BladeRunner ( https://dl.acm.org/doi/10.1145/3477132.3483572 ) which radically simplified the distributed system aspect by putting all the gnarly glue and business logic in a V8 VM (JavaScript).
My next thing is a focus on board games where I have invented a programming language and have started to evolve a platform. It's called Adama ( https://www.adama-platform.com/ ), and I think it is pretty cool. The interesting thing is that the complexity of board games is exceptional.
I have a few clues to share. The first is that reactivity, which is found in excel, is a key to simplifying software as this makes the glue more automatic.
Another clue is figuring out bidirectional communication which relates to reactivity as a two-way street. However, this is primarily hard because we don't have great things off the shelf to deal with this beyond TCP. For more of a deep dive, check out https://www.adama-platform.com/2021/12/22/woe.html which talks about WebSocket.
My final clue is that you can't run away from state. So many people offload state because state is hard, and you have to contend with it. I'm building yet another database.
If you look from a person or user perspective - time zones or date formats are used in one location or in some context. People living in that context have it a lot easier because most of the time they don't care about other time zones.
I would say people simplify things but on local scale. If you want to go global that is your problem not humans that live in one place and use single time zone all their life.
To your first point, in my experience, the best software, including complex platforms that handle massive amounts of traffic and data, can be built and maintained by small teams. I mean less than 20 developers. The larger the number of developers working on an application or platform has an almost inverse correlation to the speed in which new features are built or bugs fixed.
Software architecture has been done visually since perhaps its inception (tools like UML). At most places I've worked, every new project or large feature is diagrammed visually to use as a guide in breaking down the project into component parts.
In my experience complexity arises, not from lack of tools or industry knowledge, but from 3 main causes:
1. Inexperienced developer asked to create project, who just starts building without planning beforehand.
2. The main one - business demands features built that were never expected or planned for, and built as fast as possible. This causes developers to take shortcuts, make inelegant and difficult to maintain design choices. And leads to often inscrutable code that becomes technical debt especially after the original developer leaves the company. This will always happen as long as software is used to make a business money.
3. High team turnover - I've seen places where developers came and went so often that there was a myriad of things half started and never finished.
How I've solved or helped alleviate these issues: Make the business case to company owners or management that technical debt will be an ever increasing impediment to development velocity and the dev team will need a percentage of work in any given sprint to tackle tech debt issues (as opposed to having everybody work 100% on new features and bug fixes all the time).
I have successfully taken a platform that was bug-ridden, difficult to maintain, and where new feature development had slowed to a crawl due to the over-complexity of the software, to a place of stability and ease of development, simply by allowing our team to chip away at tech debt over the course of several years. Tech debt issues were rated by level of complexity, risk to the business in change (regression bugs), and impact on team velocity. We worked on the highest impact, lowest risk items first and kept going until there wasn't much left on that list.
“When great thinkers think about problems, they start to see patterns. They look at the problem of people sending each other word-processor files, and then they look at the problem of people sending each other spreadsheets, and they realize that there’s a general pattern: sending files. That’s one level of abstraction already. Then they go up one more level: people send files, but web browsers also “send” requests for web pages. And when you think about it, calling a method on an object is like sending a message to an object! It’s the same thing again! Those are all sending operations, so our clever thinker invents a new, higher, broader abstraction called messaging, but now it’s getting really vague and nobody really knows what they’re talking about any more. Blah.“
https://www.joelonsoftware.com/2001/04/21/dont-let-architect...
If one thinks of software programs as stored knowledge, then it makes sense why they are so buggy, error prone, etc. It's that the knowledge was never complete to begin with. Throw in the difficulties of computation, and you have the mess that is the software landscape.
A huge part of the problem is the messiness of the real world and costs.
While there may be a finite set of fundamental problems, the set of possible hardware and software platforms is ever increasing and each has it's own constraints and strengths.
A "universal programming environment" would need something like a "universal hardware interface". That's why things like Java and the Web became so popular despite their poor design.
Also, visual programming is a hard thing to make work and human language skills seem to be far greater than human visual skills.
Perhaps the best that can be done is "starting from zero" and making sure everything, from the ICs in your hardware to the memory models behind your software is thoroughly tested and formally verified.
This is insanely expensive with current technology. Perhaps some fancy math and AI advancements can make formal verification powerful enough for ubiquitous use. Until then, I see little hope for a "universal programming environment".
For example, I look at our software running on .net MVC and think as an Engineer, it simple and knowable but the Front-End Team are less worried about simplicity and more on flashy front-end stuff since that is their job. We end up bolting on a Front-end JS framework and complexity immediately ramps up by like 300%.
Are they wrong for wanting a better and more flexible front-end? Not necessarily, I mean all the other companies have cool stuff and if we don't maybe our company dies.
Ditto for lots of other examples...
There’s an awful lot of cultural baggage in coding. Many of the concepts that seem essential - powerful text editors, devils tooling - can be completely removed. It requires rethinking from first principles and being willing to upset the Apple cart, but I believe it is possible.
Let's say your software is in a financial company. Their software has to enable them to follow all the government financial regulations. Well, the government is following the larger societal "complexify to the point of collapse", and the financial regulations are certainly doing so. That means that the external behavior (the "business logic") of the software is insanely complex. You can't make that go away just by visual programming.
But maybe you're not in the financial world. Maybe you're just writing programs for internal corporate processes at some generic company. Well, your software is still subject to the complexity that builds up in the company processes. Again, the programmers can't eliminate that complexity.
Or maybe you're writing a customer-facing app - an external-facing web app, or an application that people actually install on their machines. Here you're at the mercy of the product or project manager trying to find new things for the app to do, and they still complexify the app to the point of collapse.
The problem isn't that programming is too complicated. The problem is that what we want programs to do is too complicated. Visual programming can't save us from that.
What, you think "microservices" is a new invention? We called it distributed systems, and we really knew not to go there unless we wanted to decimate our productivity and sleep.
The other reason complexity happens is dependencies and DRY thinking, which is often good, but dependencies, centralization of code, and sharing of code are also a risk. Avoiding repetition is what causes abstractions to grow. Using other people’s libraries means that there are complex interactions you don’t understand. Most of the time it’s all fine, but occasionally sometimes it’s not. This trade is made consciously, for the reason that it’s much faster to develop your app using existing libraries, and not reinvent all wheels, and nearly everyone is doing it. The reason that tooling is unlikely to solve this part of the problem is that failures are so often caused by incorrect expectations - someone using a library hoping that it will do X when it only does Y. Even in theory, better tools could only tell us low-level information about what a library does, but I don’t think it’s possible for tools to tell you at a high level what a library can’t do.
We already have had, for a long time, a 'universal' model for a specific slice of information processing that consists of a very small set of components: Graphic User Interfaces.
One problem is the diversity of domain semantics and all the associated nuances. In GUI programming, this is the tedious bit of naming the components and mapping them to processing elements.
Another problem is engineering culture (du jour). A reductionist approach that attempts to leverage structural commonalities requires what is poo poo'd as "boiler plate" and all the associated warts of component oriented programming (including factory-factories). A revisionist shift in mindset is required.
I think one thing missing from this line of thinking is that these are people. They have their own motivations and egos. They get sick, have families, take vacations, quit.
I think part of the way large teams and large companies are structured is a hedge against changing teams with highly variable capacity. It's hard to overstate how much it can hurt a project to lose a someone who has a ton of experience locked up inside their head.
The alternative is that I have worked for places that try to treat thier engineers like fungible tokens that can be shuffled and replaced at will. That environment feels extremely dehumanizing and demotivating.
I believe you are wrong, I'd wish you are right ;)
But FP seems a bit stuck nowadays (to my untrained eye) with the breakthroughs from the greybeards of the last decade.
Do you have any progress thus far? Anything people can contribute to?
I don't know why programmers think this way. We don't expect everyone to be a doctor, a botanist, a novelist, or a musician but for some reason we think anyone can be a programmer. That's just not the case -- programming is a skill like any other -- it takes some natural inclination, some training, and a lot of practice. Just like any other skill. A programming language is to a programmer what an instrument is to a musician.
Anyways, there's a reason programming should be made more widely accessible - it develops thinking skills and rationality when you learn it, as Mindstorms pointed out. Between the anecdotes in the book and things like [0], I think it's a travesty that we aren't pursuing this to its fullest extent, and instead try to teach children Python or JavaScript, two decent languages for software development but not exactly forgiving with beginners.
[0] https://medium.com/@stevekrouse/goodbye-seymour-cb712757264f
And like with these other things, you can make getting into them easier. With music, children start with simple instruments that no professional ever uses. And it's the same with programming, there are plenty of easier environments for children. But if you want to make programmers and you want to make musicians, eventually they have to use the real thing. My own son jumped straight into the deep end of Unity development knowing nothing because he wants to build something real. I neither encouraged or discouraged that environment and it's pretty unforgiving.
I don't think the solvable problems you speak of as are as solvable as you think they are. Also making software development out to be special both in terms of it's benefit to thinking skills and rationality and how it's merely some tools away from being professionally approachable to the masses is totally unfounded.
I try to work on the side in a spiritual successor of Foxpro/DBase (https://tablam.org).
I consider a mix of relational/array model fill a lot of ergonomics for this (and it was proved to be right by the family of DBase langs).
But what makes this much more complex today is the explosion on OS targets (Windows, Linux, MacOs, Android, iOS, Web), and the requirement to integrate with many other stuff, dealing with many formats (json, xml, ...), is harder to do UIs now and the base support is more inconsistent and mixed as ever...
So, is possible to make a simpler tool, I certain of it, but then the developer/user will say "ah, ok, so how this connect to Redis, GraphQL and Amazon Web Services, run this on Android, Windows, parse CSV, ..."
and that is what make this very hard at the end...
BPMN and all of the UI pallet, drag drop coding applications try to do this very thing, this is super successful at smaller scales but breaks the first real world application.
> Nobody can subtract from the system; everyone just adds.