Simple Systems Have Less Downtime
gkogan.co
gkogan.co
Don't solve problems you don't have.
That is insane!
We make a release once a week at work and things still go wrong sometimes. I am in awe how they are able to pull this off, especially at their scale.
The faster development cycle also helps people invest in testing infrastructure more effectively.
Let's say 10-20 commits per day per dev. Over 5~10 hours that's what, on the order of 1 commit-test-release cycle every 15 to 60 minutes? (subjectively for each dev)
What do we actually write in that timeframe on average (thus including the ~90% of time we don't type code but think or read or test)? What's the "unit commit" here?
So I'm thinking... let's take an example: today I'll refactor a few functions to update our model handling; I wish to reflect our latest custom types in the code. So it's a lot of in-place changes, e.g. from some list to tuple or dict; and the syntax that goes with it. No external logic change, but new methods mean slight variations in the details of implementation.
- refactor one function: commit every testable change, like list to tuple? At least, I'm sure I'm not breaking other stuff by running the whole test suite every time it "works for me" in my isolated bubble. So I might commit every 5-10 minutes in that case.
- Now I'm touching the API so I can't break promises to clients: I actually need to test more rigorously anyway. I'm probably taking closer to 20-40 minutes per commit, it's more tedious. Assuming I commit every update of the model, even insignificant, I get immediate feedback (e.g. performance dump), so I know when to stop, backtrack, try again? And it's always just one "tiny" step?
- Later I review some code and have to go through all these changes. I assume it's easier to spot elementary mistakes; but what of the big picture? Sure I can show a diff over the whole process — I assume you'd learn to play with git with such an "extreme" approach.
Am I on the right track, here? I totally get your comment but I'm trying to get a feel for how it works. I typically commit-test-release prod 3-4 times a day at most (on simple projects), and typically more like once every 2-3 days, 2-3 times a week. Which is "agile" enough I reckon... So I'm genuinely interested here. I feel there's untapped power in the method I'm just beginning to grasp.
The article talks about static analysis, I wonder if they do human code reviews at all?
Any which way we slice this, this is incredible! Sure instagram is not healthcare, transport or banking application - nobody is going to die if the website goes down, it is still an awesome achievement.
Maybe you'll find this video interesting: https://youtu.be/2mevf60qm60
The engs they had were world-class, so really, I think, saying microservices (or latest-fad) get in the way etc is disingenuous since you also require world-class talent to begin with (if you're going to keep the team-size small and yet be able to manage crazy scale) and a competitor willing to pay through the nose for the acquisition.
[0] https://www.sequoiacap.com/article/four-numbers-that-explain...
[0] https://www.youtube-nocookie.com/embed/c12cYAUTXXs
[1] https://www.youtube-nocookie.com/embed/wDk6l3tPBuw
Microservices are a design philosophy that people confuse as a deployment strategy.
I'm going to print this and show everyone, great words, thank you.
As a rule of thumb based on my own experiences and the opinions of more experienced engineers I've had the good fortune to work with, language choice is far less important than the quality of the team using it.
While I have no doubt that trying to build WhatsApp in a language that would be the wrong tool for the job (say... PHP) would have been fatal, I have many more doubts that choosing Erlang was the key enabling decision.
Success is more to do with user experience..
If you hit gold and scale you can almost always optimize the implementation later.
Sunk Cost Fallacy tends to fight any broad-stroke improvements. Until a competitor starts eating your lunch.
I've worked at companies that have monoliths that are 50x more difficult to work on because of the size. Some of them millions of lines of code. Nobody really knows how they work anymore.
However, if the organization is not capable of structuring the monolith, why should it be successful with microservices?
Such organization may lead to sharing data stores and custom libraries between microservices and that's when the real fun begins. Maybe even trying to deploy all of the services atomically to not worry about API compatibility.
But 1 semi-complex system > 10 simple systems. Especially when you consider the points of integration between those systems increases geometrically with the number of systems.
A core skill problems on a team that would prevent building a maintainable monolith do not go away because you've added more things for them to manage.
Microservice usually means distributed system. A distributed system is more complex than a non-distributed system since it has to do everything the non-distributed system has to do and, additionally, handle all the distributed problems. Microservices just hide the complexity in places where people don't see them if they take a cursory look over the code, e.g. what is a function call in a monolith can be a call to a completely different machine in a microservice architecture. Often they look the same on the outside, but behave very differently.
The hierarchy of simplicity is: Monolith > multithreaded[1] monolith > distributed system. If you can get away with a simpler one it will save you from many headaches.
> I've worked at companies that have monoliths that are 50x more difficult to work on because of the size. Some of them millions of lines of code. Nobody really knows how they work anymore.
That is a bad architecture, not something inherent to a "monolith". There's probably also a wording problem here. A monolith can be build out of many components. Libraries were a thing long before microservices reared their ugly heads. What you describe sounds more like a spaghetti architecture where all millions of lines are in one big repository and every part can call every other part. Unfortunately, microservices are not immune from this problem.
[1] or whatever you want to call "uses more than one core/cpu"
This is a false assumption. Some problems are distributed. Sometimes you'll have an external data store or you'll need to deal with distribution across instances of the monolith. You really run into pain when you build a distributed system in your single monolithic code base and your monolithic abstractions start falling apart.
In my experience you end up solving these problems eventually, monolith or not. You might as well embrace the idea that you're deploying multiple services and some form of cross service communication. You don't need to go crazy with it though.
Anyway, if your problem requires a distributed system, congratulations, you'll have to go to the top of that complexity hierarchy, and will have to solve all the problems that come with it.
That doesn't change anything about there being more problems. You just don't have any other option.
There are now almost fifty developers and we already have a third (micro-)service. ;-)
I agree completely that starting with a monolith is a win for feature development, performance and operations. But if your team grows, your release cycle will slow down unless you decentralize it.
My experience is quite the opposite. To deliver a feature, you still need to integrate changes in multiple services. To do it in a monolith is IMHO easier. You have a single artifact you can test, the probability you have the right automation is higher. Also, the devs will run bigger part of the system in the development version.
What I was talking about was the latency of smaller changes. Think of bug fixes. Somebody has to notice it, triage, fix, (wait for automated tests) and deploy. Often, a bug fix affects just one service. With smaller services, this chain is simpler.
Monolith architectures often degenerate into the proverbial "big ball of mud" over time. But if the team has the discipline to maintain proper design through continuous refactoring then they can retain the ability to deliver features with short cycle times.
Microservices are like modules but for SaaS rather than shrinkwrap, and are where you end up when you follow SLAs, encapsulation of resource consumption, etc. to their logical conclusion.
Indeed, https://en.wikipedia.org/wiki/The_Nature_of_the_Firm
Far from it. They are the most onerous, least tractable way to get what you are going for.
Lacking any better skills to identify and avoid those problems when they begin, they try to stop it from happening in the first place. Fences get erected everywhere in case they might be needed, and they frequently turn out to be in not quite the right spot or shape. The code becomes coupled to the bad interface instead of to other code, and the fixes are just as bad.
YAGNI in theory is about trying to develop those other skills, but gets twisted into an excuse for bad tech debt loads.
But what do I know? I would not have expected Instagram to run things like user login on the same instance as photo upload and processing.
what you're talking about (isolation of responsibilities) can be done and still be considered monolithic. you can also use the exact same instance type for everything and still be considered microservice.
I think what we're really talking about is containers vs machine images. And personally I think containers right now are suffering from the same abuse/hype that datastores suffered like redis/mongo/couch etc. Sure they have an application and solve problems, but they're being over used to the point of causing technical debt.
I agree wholeheartedly that simple systems have less downtime.
I would like to add a line from the talk that I give:
Simple systems fail in boring ways. Complex systems fail in fascinating, unexpected ways.
rsync.net storage arrays typically have multi-hundred day uptimes. But across our entire network, our actual aggregate uptime is unimpressive - perhaps 99.99% ? Our SLA dictates 99.95.
But when they do fail they fail in very, very boring ways. There are usually zero decisions to make in response to a failure.
there a broken link that I was interested to read on. its on https://www.rsync.net/resources/howto/rsync.html ctrl+f "rsync snapshots are detailed here"
I think given the topic, someone is expected to figure it out on their own (remember linux cake?), but I feel a bit lazy after a work day.. sorry
The definitive page for rsync snapshots has been, and always will be, here:
http://www.mikerubel.org/computers/rsync_snapshots/
... we actually don't want people to do rsync snapshots anymore because ZFS snapshots are much more efficient and use up less of their rsync.net account.
If you change one bit of a file, the rsync snapshot method will cost the entire size of that file, since it's a hard link and is either broken or not broken.
But if you just do a "dumb" sync to us and change one bit in the file, your ZFS snapshot will take up just one more bit.
So not only do you get a much simpler backup script - you can really just do a dumb sync to us and let us rotate your snapshots - but you get more efficient space usage for changed files.
I've read that page (i linked) inattentively, re-read 3 times that bold text next to the broken link and still asked that question... :facepalm:
I still miss some of the features that are hard (if not impossible) to replicate in layered storage systems, like Merkle tree checksumming, but near-nightmare that low-level debugging of ZFS was actually turned me away from that filesystem.
Obviously I have a few things to say about this ...
First, zfs was considered production worthy on FreeBSD and put into fairly widespread use as early as ... 2008 ? 2009 ? We could have made very good use of it as we were running into all kinds of limits and corner cases with UFS2 and very large filesystems with hundreds of millions of inodes.
But we waited until late 2012 to do our first deployment and it took us about six years to finally deprecate the last UFS2 systems. We did this out of an abundance of caution and a desire to see things shake themselves out over several major releases of FreeBSD.
As for the complexity of ZFS, our previous architecture had 3ware RAID cards and all of their firmware and complexity sitting between the drives and the OS. In a way there was a beautiful elegance (in my opinion) in giving the OS a single drive:
newfs /dev/da0
.... and that single drive just happened to be 40 TB in size and the OS has no idea what's going on underneath ... but there is a lot of complexity inside a full-blown RAID card and a lot of ways for drives to interact weirdly with it - especially when the drive is X years newer than the latest firmware for the card.I find that on the very deepest, hardware level, handing over raw disks to ZFS and letting it manage them is simpler. We remove a fairly complex piece of hardware and accompanying firmware (since our HBAs are "degraded" to dumb "IT" mode).
Further, all of the "bolt-ons" of UFS2 that we absolutely relied on, such as quotas and snapshots, are elegantly built into the filesystem from the very lowest levels.
We ran rsync.net (and its predecessor) on UFS2 for 11-12 years and I can tell you from the perspective of both day to day management and middle of night firefighting, ZFS has made our lives much simpler.
That is not much of an insight. It is as insightful as saying "water is kinda wet.". Well... sure it is.
What we need to deal with, is not "make a simplest system".
Rather, we need to deal with: "build a system that does A, B, C, ... and so on". Now, if you can do all of the above and make it simple... awesome. But if you cannot do all of the above, but the system is simple.... that is useless.
I agree with this. It is a good point.
Some systems and protocols have very difficult requirements to meet and different constituencies driving those requirements - it's not always possible to implement simple and elegant solutions.
Mesa/Cedar, Oberon System were also like that by the way.
Java and .NET also share a bit of it, after all Java is half-way to Lisp (as per Guy Steele), and .NET follows up on it.
This feels like an excuse to me. I’ve worked on a lot of rather simple web apps that all more or less do the same stuff. A few of them have managed to have delightfully simple codebases, most of them haven’t. There’s no reason that couldn’t be true for all of them. You usually end up having at least some complexity. But small complexity trade offs don’t necessarily require you to undermine the simplicity of the entire system.
This sounds weird. Why burn the whole thing to the ground? You've got frames from all threads available. Why would it take days to resolve the issue? Why do you think it's easier to resolve it in common lisp?
Worst case for CL means that at the very least I don't have to wonder if gdb is installed on the system. It provides a level of assurance and certainty that vastly simplifies the decision making around what to do when something goes wrong.
To be entirely fair, the introduction of breakpoint in 3.7 has simplified my life immensely -- unless I run into a system still on 3.6. Oops! I use pudb with that, and the number of uncovered, insane, and broken edge cases when using it on random systems running in different contexts is one of the reasons I am starting no new projects in Python. When I want to debug a problem that occurred in a subprocess (because the gil actually is a good thing) there is a certain perverse absurdity of watching your keyboard inputs go to a random stdin so that you cant even C-d out of your situation. Should I ever be in this situation? Well the analogy is trying to use a hammer to pound in nailgun nails and discovering that doing such a thing opens a portal to the realm of eternal screaming -- a + b = pick you favorite extremely nonlinear unexpected process that is most definitely not addition. You can do lots of amazing things in Python, but you do them at your own peril. (Disclosure: see some of my old posts for similar rants.)
Second, you don't even need that. Gdb with some tooling (see pyrasite) lets you attach to an arbitrary Python process and evaluate Python code in its context.
[...] in the real world simple is _always_ a lie [...]
I think "incidental" (non-essential, secondary, happenstance), and "accidental" (non-intentional, happenstance) are more or less synonyms here, not contrasts. I think there, uh, is basically no significant difference between "incidental complexity" and "accidental complexity".
Essential complexity is inherent to problem being solved and nothing can remove it.
Accidental complexity is introduced by programmers as they build solutions to the problem.
Lisp is nice because eliminates a lot of the accidental complexity through minimal syntax and lists as a near-universal data structure.
As a result, the simplest system is obtained, but the design process is a systematic project.It is difficult to design a complex system into a simple and smooth pipeline system.
As a programming language, Common Lisp is large enough and multi-paradigm enough to allow for elegant solutions to problems. It does require some experience with the language and some wisdom and discipline to know what pieces to select and how to best use them.
However, like all large, multi-paradigm programming languages that have been around for a while (I'm looking at you C++), programmers tend to carve out their own subsets of the language which are not always as well understood by those who come after them, particularly as the language continues to evolve and grow.
There is also the problem where programmers try to be too clever and push the language to its limits or use too many language features when a simpler solution would do. All too often we are the creators of our own problems by over-thinking, over-designing, or misusing the tools at hand.
if the language is not powerful enough to allow for that people will inevitably add preprocessors, code generators, etc... to do the things they want.
This is definitely true and it adds to the accidental complexity of the system, usually to save programmer time or implement layers of abstraction for convenience.
Common Lisp and C++ have both incorporated preprocessors and code generators through Common Lisp macros and C++ template metaprogramming and C++ preprocessor / macros. These features give the programmer metaprogramming powers, enable domain-specific language creation, implement sophisticated generics, etc.
They are powerful language facilities that need to be used wisely and judiciously or they can add exponential complexity and make the system much more difficult to understand, troubleshoot, and maintain.
We know somewhat how to control accidental complexity in systems that are highly constrained and don't let you deal well with essential complexity. And we know somewhat how to give you powerful tools to deal with essential complexity.
We haven't figured out in general how to make systems powerful enough to deal with the range of essential complexity, but constrained enough that this sort of accidental complexity isn't common.
Of course a lot of real world systems don't do a great job of either.
Some languages attempt to discourage bad taste by limiting the power given to users. Common Lisp embraces the expression of good taste by giving users considerable power.
All those complicated lists? Forth is even further down the path to ultimate minimalism. Really only one data structure, a stack with binary values on it...
Just pulling this out because it's a good sentence.
> A complex system that works is invariably found to have evolved from a simple system that worked. A complex system designed from scratch never works and cannot be patched up to make it work. You have to start over with a working simple system.[9]
* https://en.wikipedia.org/wiki/John_Gall_(author)#Gall's_law
More seriously, I think the word OP is looking for is theory. Nuggets aren't theory, they're really empirical by essence — observations.
Nuggets of knowledge are not laws. Referring to them as such annoys Klysm.
I don't have experience with Looker, but the unseen complexity of no-code tools often leads to very complex systems with interactions that are hard to understand. A dashboard may not have that many interactions with other tools, but the black box nature of these kinds of tools usually leads to significant downtimes when there's a small detail not working as expected, or where you need a small customization that wasn't foreseen by the no-code API creators.
P.S: I'm affiliated with the company Rakam.
I think a better point would be to make that that’s not a bad thing, sometimes abstractions provided in “no code” systems can greatly simplify a solution. But even the most well-designed solution will fail to some essential complexity of the root-problem the original authors didn’t understand. Without debugging tools (or privileged access to the backend of the implementing system) it’s difficult or impossible to understand what went wrong.
Incidentally I'm all ears if there's a good open-source LookML style project out there.
This was presented as the rationale quite plainly, so I don't understand why people are missing it.
A startup is a new destroyer that has been floated, did not have sea trials and went to war with a crew that might have a couple of people that used to a drive container ship but mostly staffed with kids that thought it was cool to play with a destroyer. Oh, and 3/4 of the systems are still at best have been drawn on a napkin and most of the rest came from a salvage yard. Oh and if you complete your next mission you may get money you can spend on some of the systems but if you do get that money you would be expected to do a lot more runs, quicker.
That's why it is possible to have a container ship with a crew of 5-6 people pilot the load. Containers aren't sort of identical. They are completely identical and completely standard ( several standard sizes ). They have the weight distributed in a specific way and they are stacked on a ship in a specific way depending on the weight of every container.
The startup equivalent of a container ship in simplicity is a startup sorting a pile containing A4 paper, standard business envelopes and hang folders at a daily rate of 500 items.
No, the containership is not a simple system. It's very sophisticated one. It takes massive engineering effort (literally historical effort and knowledge) to build, massive resources for the material and outsourcing, very complex (like the pic) infrastructure in case of repair, satellites to guide the ship and check the weather, operators from all around the world to track the containers, massive ports to load-unload the shipments, etc...
The operation of guiding the containership through the sea is a simple one but only once you have everything of the mentioned above in order.
If you are building software (a full SaaS solution), you are not in the business of guiding the containership. You are in the business of building either the containership alone; or the whole related infrastructure (or parts of it). That's a huge undertaking that requires massive engineering efforts. It's not a simple task.
If your task is a simple team that guide the container, you might make money for a while until everyone else figures out that you have no edge there and they can do your job for much less.
The Instagram/Tinder examples are a bad one. It ignores the fact that to find the right (addictive/highly viral) app, you'll need (again) massive psychological knowledge and expertise into how primates brains work. We don't have that, so engineers are brute-forcing by making too many apps until one hits the jackpot.
My point is that simple-to-understand systems--not simple as in primitive--have less downtime, not that we shouldn't ever have complex systems. A container ship like the one in my article can be drydocked ("in the shop") for repairs and back on the water in less than two weeks. A nuclear-powered aircraft carrier cannot. So be mindful when you're building an aircraft carrier when a container ship would've sufficed.
By the way, most of the "complex" stuff in the photo--which isn't a container ship--is scaffolding and piping.
Source: Got a naval architecture degree and a marine engineering license, operated a steamship for six months at sea, and designed ships for the US Navy for three years before abandoning ship to work with startups.
The result was a system that worked fine until you encountered bugs in the software, at which point there was no helping that several hundred people were about to die.
My point here is that the way you measure reliability is subjective and there are likely trade offs associated with “simplicity”
For example, having a global signal point of failure, like keeping all of your api servers in one cloud region, or having a single database with no replicas is much simpler and easy to understand than a more distributed alternative. However, global resources make your system more likely to fail catastrophically when there are outages.
There are trade offs here and _oftentimes_ in complexity arises in distributed systems because of reliability issues, not the other way around
Also, I think this is straight up wrong:
> If you are building software (a full SaaS solution), [...] That's a huge undertaking that requires massive engineering efforts. It's not a simple task.
If you believe it will require "massive" engineering efforts, you will certainly prove yourself right. But quite a lot of people do it differently. Look at Amazon's rule about two-pizza teams, for example. Or just this week I visited a ~30-person company that's been going 8 years with a successful SaaS business. They have a team of 4 developers, and it works just fine because they work to keep things simple.
But simplicity does not mean easy. It is actually systematic engineering. It is difficult to design a complex system into a simple and smooth Warehouse(database, pool)/Workshop(pipeline) Model system.
https://github.com/linpengcheng/PurefunctionPipelineDataflow
At this point it's literally spam.
EDIT: Literally figuratively speaking, of course.
I think you should enhance your ability to appreciate technology, Too poor technology should learn more, instead of criticizing others.
https://en.wikipedia.org/wiki/BOKA_Vanguard
I think this story is about the same liftee:
https://www.projectcargojournal.com/shipping/2019/11/20/new-...
Your system should be testable, debuggable and it should allow a human to step in and take control if needed. An open system is generally better than a closed one.
The latter part of the article talks about simplicity but doesn't seem tightly tied to the earlier part.
A good example of how to do this would be London's Docklands Light Railway. A DLR train is capable of autonomously going from one station to another. Trained operators are aboard every DLR train, and they have a key to a panel in the front of the train which reveals simple controls for manual operation. But very deliberately the maximum speed of the train with an operator at the controls is significantly lower than maximum speed under automation (there are two modes in fact, driving with the train still enforcing rules about where it can safely go which is a bit slower, and driving entirely without automation helping which is much slower). This underscores that manually driving the train is the way to sidestep a temporary problem, not a good idea in itself.
The DLR has had several incidents in which trains collided, it will come as no surprise that they involved manually controlled trains. Humans are very flexible but mostly worse as part of your safety system, don't use humans if you can help it.
But you example works nicely as well, you should avoid having to work at lowers abstractions, and limit the work there as well. "changing the core to adapt to new features rather than adding to the core"
I think a key element is that this system is composed of parts that have one input. That allows a human to replace a component that failed and provide that single input to the next component in line.
Any line of code, especially one that introduces a new concept, like a class or a method, is in “addition” to the root problem we’re trying to solve.
The closer the code/data/org-processes/whatever matches the fundamental problems we’re trying to solve, the less complexity we introduce.
However, you cannot “remove” complexity! A common mistake I see in software companies is using a single Jira ticket for one or more reports, one or more conclusions, and one or more action items.
For example, if we get reports A, B and C, they may all go into a single Jira. Oh, but C is a totally different thing! It’s not a “duplicate”! Ok now someone needs to extract the info from a comment into a new Jira (which no one will... and it’d be a mess if they tried).
Now a bunch of discussions are happening in the comments of this mega-Jira. 3 conclusions come out of these discussions — perhaps what the cause is and what should be done about it.
Then let’s say there’s two action items, one is a temp fix and one is a long term fix. Unfortunately, I usually see people is the same Jira for both tasks (“Assign this back to me when the temp fix is in so I can do the real fix”).
But splitting Jiras up is frustrating. The actual splitting is frustrating to set up, and it’s a pain to browse them.
So most Jira workflows remove fundamental complexity (reports vs discussions/conclusions vs tasks) and introduce extraneous complexity (tonnes of fields and workflows that don’t necessarily apply).
One should “embrace” the fundamental complexity of the problem(s) they’re dealing with.
This applies to programming languages, too. Python is easy but when you do advanced stuff, often without realizing, you start to grok many complex details to make it work or make it work in a performant way. On the other hand, Go (and Rust supposedly, but I don't have experience) is simple. It doesn't let you do many things and that results in fewer ways of doing a specific task but when you do it, you feel safer because in the end, it does look and feel simple although you spent 1+ hour than you'd normally spend compared to coding in Python (this assumes you're doing some advanced stuff, not quick start).
I'd say Go is able to put so much meaning to those fundamentals because 1. they have to, there isn't many other mechanisms 2. they can because it's still early in the evolution of the language. After some adoption and wildly different use-cases that users wants to be addressed, those patterns start to disappear and people end up with lowest common denominators, pointers being just pointers in this case.
It does frequently feel like this could've been solved by making basic slice assignment always go as a := b[:], and adding some type of syntax and checker where you needed to confirm that you were doing a not-completely immutable copy.
I was reading most if this yesterday which seems to disagree: https://fasterthanli.me/blog/2020/i-want-off-mr-golangs-wild...
It’s interesting to see the relative naïveté in some of the implementations of what seem to be relatively key builtin libraries like file pathing. I also have never had to deal with cross platform support in Go, so I’d never seen the magic compilation comments/suffixes.
Which is fine until the built system needs to grow. Then the pain hits.
This isn't necessarily a knock on Go. I've enjoyed working in it full time for several years now. I don't think there's a such thing as a simple programming language. Go just hides the complexity from initial inspection.
If you (parent, other readers) haven't seen it, Rob Pike's "Simplicity is complicated" is a good discussion of just this point.
video: https://youtu.be/rFejpH_tAHM
slides: https://talks.golang.org/2015/simplicity-is-complicated.slid...
Edit: if you find yourself spending most of your time trying to come up with the perfect abstraction, Go may really piss you off, I won't deny that, it's not a strength. You have to be satisfied with 'good enough' and move on, solving edge if/when they arise. Go often encourages moving toward the concrete, and you can solve generic problems in really basic ways sometimes. As an example, I was writing a DAG server last year, and coming up with ways to move data between vertices in the graph in a general way. Rather than getting out my abstractions, I just pass around []byte, and leave the interpretation of those bytes up to each vertex (often just a type cast). I personally find this refreshing, and while there are costs to doing it this way, with a few basic helper funcs, you can get 95% of what you want from a generic server like this without doing a lot of modeling.
The first half of Rich Hickey's "Simple Made Easy" presentation does a great job of defining easy/hard and simple/complex axes and distinguishing them.
video: https://www.infoq.com/presentations/Simple-Made-Easy/
It has been discussed before on Hacker News:
Or just tell them to copy it to a USB stick. Yep, same machine, it's still the best answer.
It's easier to tell someone what keys to press than to find and manipulate UI elements.
on the same PC
On a PC, everyone already has access by default, so "sharing" doesn't involve anything additional.
(That's the reason why first the Windows explorer was suddenly sufficent enough and then almost completely disregarded)
A public folder on the same PC is way harder to explain than just telling your photo app to share it with your resident spying company/cloud provider, despite how wasteful the round trip is.
Not that that has anything to do with a discussion about the "unix way" as a special case or even superset of IDEs.
Christ, it's such a fucking mess that one of the most compatible ways to distribute software is to write it for Windows and rely on WINE.
And I split my time between Arch, CentOS & OpenBSD. The vast majority of things these days are packaged in a useful way. There's also flatpak and similar now.
Good for you. Some of us have to install up to date software more often than every 5 or 10 years.
Way simpler to do on Linux than M$ Windows.
> [no way to ] to directly distribute binaries without a gigantic fucking headache.
I mean, flatpak, appimage, docker, etc ...
Completely unnecessary on Windows since it is an operating system, not a kludge of random source code from the internet.
> I mean, flatpak, appimage, docker, etc ...
Which of those is ubiquitous enough to ensure availability of what you're looking for in that format? In my experience: none of them. AppImage is easily the best, but the community seems to hate it because it makes things too flexible and simple or something.
Where, pray tell, do people who write the software for this "operating system" store their source code if not somewhere that is connected to the internet? Do they use punch cards?
> Which of those is ubiquitous enough to ensure availability of what you're looking for in that format?
Have not met someone who hates AppImage. flatpak is also rather ubiquitous. I use both every day.
Talk to Drew DeVault, to name one.
And it is 2020, I guess nowadays everybody understand how unix-like system works and how simple they are.
"Choose tools that are simple to operate over those that promise the most features."
Doing this very much depends on your requirements, doesn't it? What good is a simple-to-operate tool if it doesn't do what you need it to do? Sure, maybe you can simplify your requirements; then again, maybe not.
The example of troubleshooting a whitepaper form very much depends on the design of the system. Maybe there's a reason for having multiple forms. If they all shared common architecture and the problem lies in what they share, it's not necessarily more difficult to troubleshoot many as opposed to just one.
Moving from Marketo to HubSpot is good and well - for now. What about years from now? How did they end up with such complexity with the Marketo solution? Both solutions and requirements evolve. Down the line, you could end up with the same difficulties with HubSpot. I think it depends in part on how the organisation handles change.
Lastly, I agree with the sentiment of Mateusz Górski, who commented on the page: is there any hard proof beyond anecdote to back up the article? If you expand the article's concept of a system to beyond just software, all you need is a hardware fault in a third-party hosting company with lousy support for your simple piece of software to be down for weeks.
I suspect the author would reply that you can just replace the hosting provider. Since the interface and division of responsibilities is clear, one host is completely replaceable with another.
One non-software system was a rather expensive UPS/generator. It was meant to trip on a power loss and provide X minutes of stable power. In reality, it was incredibly sensitive to minute power fluctuations and would trip and then immediately fail, dropping all supplied power in an instant. The system was dramatically more reliable with this unit simply disabled.
I've seen similar things with rabbitmq, in situation where the load & number of events to process is tiny but HA is some checkbox item to deliver without a serious QA process to measure if the setup is actually working effectively. Number of events in production we genuinely needed multiple rabbitmq nodes in the cluster to cope with failure of one or more nodes: 0 . Number of times when rabbitmq got too excited processing large messages that it delayed heart beating & subsequently decided that there was a network partition when there was in fact no such thing, leading the prod support team to step in and manually recover the cluster: > 0
Perhaps I've just had bad luck, but my impression is that getting this right, both internally and in the field, is a lot harder than it looks. And when it breaks, it can be a real s___show compared to the simple, non-HA version.
Due to my lack of web development knowledge and massive time crunch, I built it with no framework, zero best practices and the most basic code I could build. The code is atrocious, and would not scale at all but it does what it is supposed to do and works because I did everything in the most basic way, there are no gotchas. Straight html, basic javascript and php on the backend. It interacts with google docs as well. Its almost to dumb to break.
Sometimes we really do complicate things with all of our new fancy frameworks, microservices, states, etc.
With that said, it would be impossible for anyone else to maintain and if anyone posted the code on the web I would probably be unemployed. But it works, very difficult to insert new features though.
I'm not sure if the simpler solution - hubspot existed when the company started hacking solutions together.
Also, the move to hobspot was relatively easy because the business processes and workflows were already well known and defined.
Starting from scratch, fighting the fires as they appear, a bit of patchwork seems unavoidable.
Everything is obvious in hindsight.
I guess the real lesson here:
Simplify once in a while. Like Facebook just did with messenger.
Author's value proposition is "I will help you outsource your complicated marketing tools". If you reject the gospel that your marketing tools are complicated, author's service would be not needed.
There's an ELK stack + plugins for logs and alerting and saltstack to trigger some basic automated sysadmin tasks.
I left 3 years ago and the entire infrastructure has run on autopilot with zero downtime. Content editors and web developers still add to the content constantly. There is some dynamic content on the site handled by JavaScript and a few light APIs that I built to support the site. There's over 5000 pages of content. This simple infrastructure powers a website with a low-5 digit alexa rank for a business bringing in about a billion dollars a year.
The entire infrastructure bill comes in under $100/mo and they no longer need to employ any backend engineers. We negotiated a contract agreement for me to support the site if they need it. They haven't.
Are you saying that you use s3 as your only data store?
* Yeah, but often they do less / are less effective during uptime.
* "simple" is an inspecific concept. Is the implementation simple? The API? The hardware? The requirements?
* "simple" is a vague concept. Is a 1000-line C program more complex than a one-liner script in a fancy higher-level scripting language? Yes, if you ignore the 1,000,000 line abstract virtual machine, JIT compiler, interpreter, large standard library and what-not. Otherwise - maybe so, maybe no.
> Modifications before additions
I actually agree, but:
* Managers don't like it.
* Sometime this downtime happens before there's ever been any uptime...
* Tomorrow you need something else, and the changes for that require re-engineering the parts you already re-engineered to best accommodate the previous addition.
Need to handle machine failures? Push a bugfix without causing downtime? Maintain a single state across the redundant machines? Handle load spikes? Ones created on purpose to DOS you?
All of this requires additional mechanisms, which can themselves cause failures.
A lot of Software Engineers / Developers in my experience tend to over engineer projects or they are mandated they do so by so their dev leads. It is easy to get caught up into abstracting almost everything away almost for the sake of it while it provides no real benefit.
When the shit hits the fan, there's an 80% chance that your most senior members will find the problem. They may find it quicker in the simple system, even, and others might get there first. But that last 20% is a very long tail, and even longer in a complex system.
In a complex system, many of your members cannot participate in the triage process, because they don't know enough to know what's relevant. They slow down the people trying to work the problem trying to learn new things (good) or offering low-probability scenarios (bad).
There is no 'All Hands on Deck' scenario for the complex system. To be responsive, you have to kick some or even a bunch of people out of the room, and once they leave they can't really participate.
The first phase of a triage is getting a tight repro case. Many avenues are blocked until that happens. And some people have a knack for repro cases but aren't so great at debugging. With more people you get a higher quality repro, which cuts a lot of time off the rest of the process.
Part of debugging is the cost of the verification versus the likelihood. Simple Systems afford the opportunity for people to test out unlikely but plausible scenarios that are in the long tail, without distracting from the more 'boring' checks already being done.
And when bugs are identified prior to deployment, a simple system means you can hand the repro case to the responsible party and expect/demand that they get their fix on the first try.
> When new requirements come up, the tendency is to add layers on top of the existing system—by way of additional steps or integrations. Instead, see if the system’s core can be modified to meet the new requirements.
I like the mention of this as it is - at least to me - really counter-intuitive.
- Thomas Paine, Common Sense, 1776
This is a well written article, but the metaphor is a little stretched.
Marketo and Salesforce are the "ship builders". Marketers are the crew operating the ship.
The declarative tools in enterprise apps provide Marketers with conditional branch and execution capabilities without code structure, so the average Marketer just keeps adding layers of complexity. Whereas a Developer would continually refactor and make old code obsolete.
The crew operating actual ships do not have anywhere near this level of access to the ships configuration, hence more reliability in the system.
Such a good line. I've seen companies killed by hard-to-follow arguments advocating for technical complexity which serves feature goals which serve business goals.
'Good strategy is simple' from some strategy book is sometimes right; arguments that hold business goals hostage to technical complexity are trouble.
> In the end, the system I put in place had 97% fewer processes (from 629 to 20) while providing all the same capabilities. A bug that was found a few days later got resolved in four minutes.
It‘s called refactoring. It‘s a constant process, that none likes to allocate time to, because it does not move metric immediately. Especially in marketing with a very fast changing landscape you go from 20 to 600+ very fast.
What are alternatives to commonly used complex systems, like Docker/Kubernetes, Active Directory, Remote Access or just shared folder structures? How to keep the functionality of centralized management while making it simple? Would be happy to hear your thoughts on this.
But after a certain point, it's no longer enough to make it simpler. You have to get a lot more involved again, and this time the complexity has to be intelligently focused around the end goal (be it performance or uptime).
Good luck attracting good developers with that. We use something over engineered but we have fun. The customers don't care but we do.
5 pages of pure awesomeness
is the complexity appropriate for the goal (i.e. Everything should be made as simple as possible, but no simpler. )
Hence, often a required step to fix complexity is to change the goal, in such a way that it can be implemented with less complexity.
I am suspicious of this statement. Is the author now supporting the company and defacto replacing the guy that left?
https://github.com/linpengcheng/PurefunctionPipelineDataflow
It is simple and it just works for its intended purpose.
I am working hard to apply this to the design of Webase [1] which is a #nocode platform that is inherently complex. But I really believe that if we get the design right it will support a large number of sophisticated use-cases.
Thank you for sharing!
It's fine on HN to post your own work (1) in places where it's relevant, as long as (2) you only do it occasionally and (3) are also participating in the community in the intended ways, submitting interesting articles and having curious conversation. But when users break these rules of thumb (and you've been breaking all three of them!) it crosses into spamming. The community is very aware of that and really doesn't like it, so it's not in your interest.
Doing this by hijacking top comments and threads about other people's projects is particularly not ok.
I have been trying to add to the overall value of HN with insight in general and only including a link to my site if it is relevant.
But I hear you and will adjust accordingly.
I do sincerely appreciate the feedback as I have been steadily gaining karma and then out of no-where it when backwards today.