“Clean” code, horrible performance
computerenhance.com
computerenhance.com
> So by violating the first rule of clean code — which is one of its central tenants — we are able to drop from 35 cycles per shape to 24 cycles per shape
Look, most modern software is spending 99.9% of the time waiting for user input, and 0.1% of the time actually calculating something. If you're writing a AAA video game, or high performance calculation software then sure, go crazy, get those improvements.
But most of us aren't doing that. Most developers are doing work where the biggest problem is adding the next umpteenth features that Product has planned (but hasn't told us about yet). Clean code optimizes for improving time-to-market for those features, and not for the CPU doing less work.
If you're scaling to 1000s of users then yes. If you have a GUI for a monthly task that two administrators use, then no.
The less something gets used the longer the payback time on the initial development.
Fine, but be honest with yourself and admit that you are contributing a lot to making the lives of those two admins miserable.
It doesn't matter if I'm using your software once a month, or once a day. If it's anything like typical modern software, it will make me hate the task I'm doing, and hate you for making it painful. In fact, shitty performance may be the very reason I'm using it monthly instead of daily - because I reorganized my whole workflow around minimizing the frequency and amount of time I have to spend using your software.
The funning thing is I'm thinking of a specific case and I work closely with those admins. I even have filled in for them when they're sick. Yes I know it's a pain, they know it's a pain, and they let me know it's horrible. The only reason this one is monthly is that it's a stock take. They forgive the crappy performance as it still saves hours of work when compared to the previous manual option of entering things into multiple systems.
In non-performance-critical areas, it's pretty important that when the original dev team leaves, new hires can still fix bugs and add features without breaking things.
... And maybe we want set-operations in the future...
We spend most of our time reading code. If the Casey's code snippets are easier to reason about (which they are, especially as the codebase get larger), that's a big win. I'd imagine you want to optimize for code that is easy to (re)write, rather than minimize the number of key strokes while increasing the time spent understanding the code.
What is adding a polygon going to do? Either back to the 'bad' interface to hide variable object size, extend to a tagged union with a size of the biggest datastructure - hurting cache performance each time it grows -, or an involved Object of Arrays structure - not good for clarity but great for performance.
All while having to remember which field, width or height, to use for the circles radius.
As demonstrated in the example, extensibility costs in performance, and also sometimes in comprehension. If we apply the “clean code” rules as a matter of course, we pay this price always.
In my opinion, we should use interfaces etc at module boundaries only.
if he simply kept the original "unoptimized" switch case method, what you say wouldn't apply. it couldn't. from a pure feature standpoint, a switch case is functionally identical to polymorphism except that you can't add types that are unknown at compile time (like loaded at runtime as an extension from a library or something). and that version is already faster.
what the blog post does after that point is merely point out that by having everything in one place, you see opportunities to optimize. this is a separate thing where you still have to consider whether that actually makes sense.
like, if you can guesstimate that an entirely different type is likely to enter the picture at some point, you may skip this optimization. if production realities mean that code needs to be faster, you can still apply it and add some comments about how to change it back, or just keep the original version commented out with a reference for why its there. and so on.
If that's true, why does it take forever to load and frequently fail to keep up with my input?
Modern computers are ludicrously fast, and modern developers have somehow managed to make them slow.
Regardless, these programs should be spending 99% of their time waiting for user input, but instead they're working data through a mountain of abstraction layers on the faulty assumption that it is a) saving developers' time, and 2) that that is worth much more than user time.
Emotionally we have no idea how bad it is to waste as little as a few seconds per day for millions of users. It's just a few seconds, right? We're just forgetting to multiply those seconds by the number of users.
[1] https://github.com/rails/jbuilder, maintained by DHH and the Rails team. AFAIR, the "official" JSON serialization DSL.
This is related, though. While the final leaps of performance often come from using more memory to make CPU work less, in my experience, most of the performance-problematic code wastes both CPU and memory, and meaningfully faster alternatives end up using less of both.
Either way, from the end-user perspective, I tentatively agree with the article I mentioned (but don't have a link handy, sorry) - most of the software I end up using, or see people around me using, is clearly CPU-bound, briefly network I/O-bound (mostly web apps), and rarely local I/O-bound.
EDIT: I guess the caveat is that end-user software is often spending 99% time doing nothing at all, just waiting for user input. But that bit doesn't matter - what matters is how fast it reacts to the input once the user starts providing it. This is where a lot of software suddenly gets CPU-bound (or net IO-bound, if doing something stupid).
If we had perf issues that showed up outside they were higher level design issues like 1) trying to take a thumbnail of the page at full resolution for a tab thumbnail while loading another tab, not because the thumbnailing code itself was slow, or 2) running slow O(tabs) JS teardown during shutdown when we could run a O(~1) clean up step instead.
That's really not even close to true. Loading random websites frequently costs multiple seconds worth of local processing time, and indeed, that's often because of the exact kind of overabstraction that this article criticizes (e.g. people use React and then design React component hierarchies that seem "conceptually clean" instead of ones that perform a rendering strategy that makes sense.)
Unless we’re talking about specific compute-intensive websites, this is almost certainly network loading latency.
Modern web browsers are very fast. Moderns CPUs are very fast. Common, random websites aren’t churning through “multiple seconds” of CPU time just to render.
So if you were to manually run a for loop over an array instead of iterating through it I'm not surprised you got an order of magnitude faster performance.
In JavaScript on the other hand it handles the abstraction by constructing a separate function for each pattern you use. In those functions it allocates an array the same size as the working set, iterates through each element, then returns the new array. If you're doing multiple operations at once this means that you have to have multiple allocations and multiple iterations through the entire working set.
Because it's doing an allocations it's basically doing an extra memcpy for each additional pattern past one which is a giant slowdown. Then if the working set is too big for the L2 cache it needs to be reloaded from L3 each loop. If the working set is too big for L3 then it needs to reload from main memory EACH TIME.
If you wanted to implement patterns as slow as molasses I can think of no better way than to sugar it out like JavaScript did.
I use Jira every day and, no, it does not take 30-60 seconds to load a page.
The hyperbole in this comment section is something else. Either that or people are using 15 year old computers to browse the web.
- Loading: 29ms with uBlock vs 42ms without
- Scripting: 548ms vs 1850ms (!!)
- Rendering: 42ms vs 105ms
- Painting: 8ms vs 29ms
- System: 216ms vs 295ms
- Idle: 460ms vs 8707ms
Holy shit, advertising and trackers are absolute resource hogs.
I’ve personally spent many hours doing performance analysis, triage and remediation on websites built using modern tech stacks that had inadvertently exchanged UX for DX. Too much JS sent over the wire can definitely tie up the browser’s main thread for whole seconds even on desktop, though in my experience it’s much more common on mobile. This situation can be difficult to correct depending on the abstractions, organization and overall architecture you chose early on, and code-spitting and dead code elimination won’t always fix what’s broken.
I know React tends to lack in both dev UX and performance (at least in my exp). Personally I've taken a look at Svelte and Solid, and liked them both. I haven't had the chance to build anything larger than a toy app, though.
When react introduced hooks, it was fun for a while, but then we discovered the frequent re-render issues and had to change the way we think when building components and refactor old components. We have to manually wrap components in useMemo. This is the kind of things Vue has avoided me so far and React let me down. React relies on running all render methods instead of doing granular updates, I think this hurts performance. Vue listens to property changes and can perform granular updates.
(It was not. It was why I stopped doing frontend work.)
I'm curious to give WebComponents a try as well.
But overall, I think you overestimate how much time you spend loading the website and how much time it's just sitting there, mostly idle.
And in the end, as long as it's fast enough that users don't stop using the site/webapp/program/whatever, then it's fine, imho. When it becomes too slow, the developers will be asked to improve performance. Because in the end, economics is the driver, not performance.
reading the website is productive. Waiting on it to load is not.
I'd rather have app burn a whole core whole time if it cut 1s off load time when clicking something.
2s you notice as a user. 50ms you won't. In fact even 500ms you won't notice too much but we are getting close to where optimization will be noticeable to a user.
This is separate from the discussion about tradeoffs between flexible design patterns and low-level performance.
https://www.marketingdive.com/news/google-53-of-mobile-users...
My most recent misadventure: trying to write hand-coded ARM Neon assembler, and/or C++ code with neon intrinsics to optimize a piece of real-time audio effect code. The clear performance winner: plain C++ with no intrinsics, but tweaked to allow auto-vectorization (plus judicious use of the __restrict modifier for a small but significant boost). GCC produced code that had better instruction scheduling than I could (but not for the NEON intrinsics, oddly). And as an added bonus, the same plain-old C++ code generates AVX vectorization on MSVC without modifications! (MSVC also supports __restrict until the C++ standards committee gets their act together to adopt the eminently necessary C99 restrict keyword).
That is wrong. Just because you didn’t succeed in writing faster code in one case, doesn’t mean it’s impossible. See e.g. [1] on why the Lua vm is written in assembly. It’s from 2011 but not much has changed in the meantime.
Oh it definitely is. But no single framework (or frameworks) are to blame. It's the CMS. The thing that allows every Tom, Dick, and Harry at the company to drop their little snippet for this or that which adds up to a mountain of garbage over time, half of which is disused or forgotten.
It could be that most of the traffic is served by websites using heavyweight frameworks, but the long tail of low traffic websites use them rarely.
Do enlighten us then, what is it due to?
You attempted to present an argument that's a textbook example of an hyperbolic fallacy.
There are worlds of difference between "this code does not sit in a hot path" and "let's spend multiple seconds of local processing time".
This blend of specious reasoning is the reason why the first rule of software optimization is "don't". Proponents of mindlessly going about the 1% edge cases fail to understand that the whole world is comprised of the 99% of cases where shaving off that millisecond buys you absolutely nothing, with a tradeoff of producing unmaintainable code.
The truth of the matter is that in 99% of the cases there is absolutely no good reason to run after these relatively large performance improvements if in the end the user notices absolutely nothing. Moreso if you're writing async code that stays far away from any sort of hot path.
If you're lucky enough to eat your own dogfood, run unit tests against a copy(!) of your in-house database. A utility to anonymize the data was a fairly significant investment, but absolutely worth it in the long run. Being able to run and monitor benchmark unit tests for selected critical operations on enterprise-scale test data as part of the continuous-build process: fabulous!
How this typically works with performance is very similar, with a small team working to identify problematic areas where optimization would drive the highest impact, and the rest of the organization keeping performance in mind but not otherwise concerned with it in their day-to-day work.
If you can't recognize the percent where it matters a lot, clearly that percent does not exist and you're bothering about nothing.
I see a lot of talk about the importance of tuning Formula 1 cars in a world where everyone drives ford fiestas.
And when you do identify the 1%, you need to be testing optimizations with a profiler constantly while optimizing. Profile. Do some optimization. Profile again. Roll back if not successful. Repeat until done. It's impossible to optimize well if you're not doing profiling.
The ultimate tools would be either Intel's profiling tools, or ARM's profiling suite (both very expensive). But MSVC and GCC do a fantastic job of scheduling instructions to avoid pipeline stalls these days, so these deep profiling tools are unlikely to gain more than 2 or 3% performance increases these days. (Worthwhile pretty much only if you're writing GPU drivers for NVidia or AMD).
Taken as a given that 100% of your code is at least algorithmically correct in the first place. (Appropriately better than O(N^2) whenever possible).
- Former writer of graphics drivers, currently audio DSP engineer.
VTune allows you to determine where pipeline stalls are occurring at the instruction level (for that last 2 or 3% gain in performance). I haven't worked with ARM profilers (way out of my price range), but I assume, given the exorbitant price, they provide the same sort of in-depth analysis. Probably a handful of people on the planet that need that kind of in-depth analysis.
My point being, if most of the program is slowed down by a slew of wasted CPU cycles (costly abstractions, slow interpreted language…), there's a good chance what should have been obvious bottlenecks get drowned in a see of underperformance. They're harder to spot, and fixing them doesn't change much.
So before you even get to actual optimisation, your program should be fast enough that actual optimisations have a real impact. And yes, actual optimisation should be done quite rarely. But first, we need to make sure our programs aren't as slow as molasses. See https://www.youtube.com/watch?v=pgoetgxecw8
I feel you're missing the whole point.
It's immaterial whether anyone can get to optimizations that compound multiplicatively. The whole point is that halving something that costs nothing earns you nothing. That's the whole point. Go ahead and shave off that millisecond. Will anyone actually notice whether you add or remove that penalty? Odds are, not at all.
Monocypher's speed was actually an important component in its success in the embedded market, even though I didn't explicitly target it initially (I was lucky my portability driven decisions made it a good fit there).
There definitely are niches where there are quite a few performance optimization opportunities that users do care about.
In your example making something a user is actively waiting for go from 3 seconds to less than one is a great optimization target. What is not a great optimization target is making something the user is actively waiting for and that takes 30ms take 25ms instead. That's wasted money on developer time.
If your "user" is a developer of embedded software with memory constraints and using your library leaves them more room that's awesome. If your user was someone using the library on a general purpose computing device with loads of memory then the 2 vs. 5 does nothing.
I decided not put it off for later, and keep things simple for now.
For example, in the case of web development, if you build a medium-sized website with React (i.e. pretty normal behavior nowadays), then if you make default decisions that don't consider performance at all, your website will end up noticeably slow, because you will:
1. Write code that re-renders components all the time during loading and UI interactions,
2. Which depend on tons of third-party dependencies that perform poorly,
3. So you end up spending a ton of time in re-renders while the site loads and while someone is using it.
Dealing with this isn't literally the same performance work that Casey put in his article, because it's at a slightly higher level of abstraction, but it requires the same mindset. It requires writing most of your code (and taking on dependencies) with performance in mind, not just 1% of it. You can't avoid it without your notion of "clean code" including some amount of mechanical sympathy, rather than just being about abstract extensibility and generalization concerns.
This is a mistake that I see far too many government and utility sites making.
If you're really interested in the impact of performance issues on everyday life, you need to provide concrete examples instead of putting up unverifiable strawmen.
The truth of the matter is that 99% of real world applications run just fine, and it doesn't pay off to invest in shaving milliseconds here or there. Would it be desirable to have a magic wand to improve some edge cases? Yeah, why not? Is it worth to pay people to spend time with a stopwatch at hand to shave off these milliseconds? Not really. It's all about tradeoffs, and there is no real world payoff in wasting developers' time to shave off that millisecond here or there.
Not to mention that web apps redrawing everything whenever they feel like it can almost give motion sickness, lead to clicks on the wrong things etc.
Yeah, but users' time is worthless.
There's so much hardware out there that can run native applications just fine, that can play back HD video, that can run complex 3D real time video games, but crawl like molasses when loading your average webpage. Facebook and YouTube are terrible offenders, but so are your average blogs. Many banking websites are terrible (yet they don't have to be; my local credit union has a zippy website that looks attractive to boot, has modern design elements, etc.).
Maybe the hardware you're running is eye-wateringly fast, or maybe it's just barely fast enough and you don't need the cycles for anything else. But we're not talking milliseconds. We're talking order(s) of magnitude. I can't bring myself to believe you don't see at least some of it, if you just open your eyes and look around.
"Unverifiable!", he wrote from within a web browser.
Another app that I'm not sure if it's React or not is New Reddit. It is significantly slower on my computer and on a lot of people's computer and sometimes you have to refresh the page because it consumes too much memory.
I can come up with other local examples, from Germany. Vatenfall's website seems to "traditional" page navigation, but the content is loaded via a framework. Due to having to reinitialize everything, text takes up to 5 seconds to appear when using the back button or when navigating for page to page. Similar things happen in the German Agentür fur Arbeit.
Let's not even get into low end phones in developing countries or bargain basement android tablets, let's stick with something straightforward - an ordinary PC.
I took a quick look online, sorting by cheapest first I found something with an AMD 3015e in. Based on cpubenchmark.net that gets a benchmark of 2691. Taking a look at the big list of CPUs I see that's equivalent to a powerful desktop CPU from 2008, or a decent laptop from 2012. (The Apple M2 in the current Macbook Air gets a score of 15369, just for comparison.)
So, if you're writing PC software or making a fancy web app and you want everyone to have a good experience with it then you should see how it runs on a terrible new laptop, or a 12 year old good laptop, or a 15 year old powerful desktop.
(And yeah we all have SSDs now which is much better than in the old days, and JS is generally single thread, and single threaded CPU performance has not improved so much - but I think my point still stands.)
You're taking deep quaffs of the Kool-Aid and so are most of the people commenting on this story. General software responsiveness and reliability (i.e., usefulness) has been in decline for decades. This is an objective fact.
Writing objective reality off as mere "milliseconds", "edge-cases", or only relevant for "toy problems" exemplifies the arrogance and severe incompetence of most programmers. People are seriously trying to talk down to Casey Muratori when in all likelihood they haven't accomplished even 1% as much as him as programmers.
I get it—no one wants to leave fantasy-land as long as the easy money is flowing. But sooner or later the glittering carriage turns back into a pumpkin.
It's certainly a huge problem for real world applications
Who was talking about a single millisecond here?
I notice, broadly, two types of people who engage in these arguments.
1> OMG, computers are thousands of times faster than they were a decade ago, why is everything not lightning fast? Why are so many things slower than they were back then? Why is my chat program eating 2GB(!) of RAM?
2> Because we're busy writing six billion features on our Nth iteration of this problem space, we can't be bothered to shave a few millis bro!
And they just talk past each other.
It's because of this mentality that almost all desktop software nowadays is bloated garbage that needs 2GB of RAM and a 5Ghz CPU to perform even the most basic task that could be done with 1/100th of the resources 20 years ago.
And other lies you can tell yourself to sleep at night.
Most people notice. Very few have the capacity, power (or realise) to complain about it. They accept what they’ve been given, despite how awful it is, because they have basically no other option.
A concrete example: my previous work had to use bitbucket pipelines for our docker builds. My current work uses GitHub actions. GH has my container half-built before I can even click through to the page. Bitbucket took a good minute to start. My complaints about bitbucket fell on deaf ears in the business, and no amount of leaving feedback for Atlassian to “please make builds faster” ever made the slightest amount of difference. Every time MS teams comes up that’s met with complaints about performance (among other things) so people definitely notice that. VS Code gets celebrated for “actually having decent performance”, so the bar is so low that even moderate performance apps receive high praise.
Users definitely notice performance, whether the PM/business cares is a different matter, but we should stop deceiving ourselves by saying it’s alright because “users won’t notice or care”.
Performance is a feature. Or does anyone here enjoy using an old tomtom gps where every tap on the screen takes 2 seconds? If so, please donate your beefy laptops to charity.
"We should forget about small efficiencies, say about 97% of the time: premature optimization is the root of all evil.
Yet we should not pass up our opportunities in that critical 3%."
Using Electron and bloated frameworks that "abstract" things away is the biggest problem for most modern software, not the fact the developers aren't optimizing 10 cycles from the CPU away. It's a fundamental issue of how the software is made and not what the code is. If you need to run a whole web browser for your application, you already lost, there is no optimization that you can do there.I think the sibling comment here in this thread shows what some developers think. Electron is not reasonable. "Developers" using Electron should be punished by CBT and chinese torture.
Okay maybe that's too much. But I suggest a new quote:
"We should forget about Electron, say about 100% of the time: electron is the root of all evil.
Yet we should not pass up our opportunities in rewriting everything in Rust."I sometimes wish performance was an issue in the projects I work with, but if it is, it's on a higher level / architectural level - things like a point-and-click API gateway performing many separate queries to the SAP server in a loop with no short-circuiting mechanism. That's billions of lines of code being executed (I'm guessing) but the performance bottleneck is in how some consultant clicked some things together.
Other than school assignments, I've never had a situation where I ran into performance issues. I've had plenty of situations where I had to deal with poorly written ("unclean") code and spent extra brain cycles trying to make sense of it though.
make it work
make it work correctly
make it work fast
Pretty was never in the picture. But if anyone wants to add it, it should come last.
Make it work
Make it good
Make it fast
That is not at all what "make it work, make it pretty, make it fast" is about. That saying is about prioritization. Making it fast doesn't mean anything if it doesn't work.
However, if you are doing performance-sensitive work then this is a very bad strategy. You need to design a performant architecture up front otherwise you'll likely have orders of magnitude worse performance, even after optimizing your code.
Ex: if your "make it work" design has shared mutable state, you're going to have a bad time when you want to scale that horizontally and unlock 100x better throughput/performance.
So, most code that I write tends to have dead simple data types, but combines it with classes (and subclasses) that represent methods (strategies) on how to retrieve, transform, present and store the data. The 'make it work' phase may do this in a simple script, but the actual data model tends to stay the same.
Far too often I see applications that assume low latency and unbreakable Internet connection. They seem to do almost no caching at all. For example thumbnails.
Also many of applications will be almost unusable (or trigger OOM) when you try to work with a big file. Sometimes a big file has merely tens of MB, sometimes problems start with a 3MB file. Those are the issues that occur without thinking about performance from the start - memory is free, you can copy things around, everything will fit in RAM.
One more thing. When your application consists of a client and a server it may turn out that you will put yourself in a corner when not thinking about performance early on. Everything will work without any troubles at first and then it turns out there are some latency issues with more data and you can't easily upgrade the client for example. Or you had an architecture that allows to spin up more servers and handle the load closer to the client, but it can cut your margins.
Does it though? Where's the evidence for it? The vast majority of people I've worked with over the last couple decades who like to bring up "clean code", tend towards the wrong abstractions and over abstracting.
I almost always prefer working with someone who writes the kind of code Casey was than someone who follows the clean code examples I've spent my career dealing with. I've seen and worked with many examples of Data Oriented Design that were far from unmaintainable or unreadable.
> Prefer polymorphism to “if/else” and “switch”
Algebraic data types and pattern matching (a more general version of switch), make many types of data transformation far easier to understand and maintain (versus e.g. the visitor pattern which uses adhoc polymorphism).
> Code should not know about the internals of objects it’s working with
This is interpreted by many as "don't expose data types". Actually some data types are safe to expose. We have a JSON library at work where the actual JSON data type is kept abstract and pattern matching cannot be used. This is despite the fact that JSON is a published (and stable) spec and therefore already exposed!
> Functions should be small
"Small" is a strange metric to optimise for, which is why I don't like Perl. Functions should be readable and easy to reason about. Let's optimise for "simple" instead.
> Functions should do one thing
This is not always practical or realistic advice. Most functions in OOP languages are procedures that will likely perform side effects in addition to returning a result (e.g. object methods). Should we also not do logging? :)
> “DRY” - Don’t Repeat Yourself
The cost of any abstraction must be weighed up against the repetition on a case-by-case basis. For example, many languages do not abstract the for-loop and effectively encourage users to write it out over and over again, because they have decided that the cost of abstracting it (internal iteration using higher-order functions) is too high.
Open is great for extensibility: libraries can be precompiled, plugins are possible. Changes don't propagate - it is ""forward compatible"" which is great for maintaning the code.
Closed on the other hand is great for matching. Finite is predictable & faster. Finite is self-contained and self-describing because it exposes the data types without shame.
The purpose of the visitor pattern is now clear: it closes the open polymorphism for a finite set. Great, now we only need one kind and we still get matching. Or is it the worst of two worlds? Slow & incomprehensible and all changes propagate everywhere.
So which one is better? Neither. But the reality is that all old imperative languages with polymorphism chose the open kind - the kind that adds more features because it was needed for shared libraries. Leaving you to build any pattern matching yourself and to burn yourself with the unmaintainable code.
If people get burned, they learn. First they say don't do that and only then they replace the gas stove by induction.
I agree. And your "2 cents" demonstrates a depth of understanding greater than that offered by these rules.
A really good developer writes clean code using the right abstraction (finding those tends to take the most time and experience) and drop down to a different level of abstraction for high performance areas where it makes sense.
The fact that bad developers suck and write bad code no matter if they use clean code or not does not reflect on the methodology
I personally haven't seen value from that coding style. There may be some platonic ideal clean code that is better than other methodologies in theory--it is likely that my sample is biased--but from what I've seen, the clean code style tends to lead most developers towards over abstraction.
For juniors which have no experience, any sane methodology is better than none, since otherwise you get even more of a mess.
That said, Clean code has some great advice, some mediocre advice and some frankly bad advice, but the authors point are largely irrelevant to 99 % of software engineering.
It is easier to find an abstraction if we lay out what the program is doing all in long functions, that just "do what they do" until you figure out what needs to be abstracted.
Not just general advice, but advice that is meant to be applied to TDD specifically. The principals of clean code are meant to help with certain challenges that arise out of TDD. It is often going to seem strange and out of place if you have rejected TDD. Remember, clean code comes from Robert C. Martin, who is a member of the XP/Agile gang.
The author cherry-picking one small piece out of what Uncle Bob promotes and criticizing it in isolation without even mentioning why it is suggested with respect to the bigger picture seemed disingenuous. It does highlight that testable code may not be the most performant code, but was anyone believing otherwise? There are always tradeoffs to be made.
It reminds me of Twitter or other companies which starting to change programming languages for performance reasons.
CPU meter when clicking anything on "modern" webpage proves that's a lie.
Also, sure, even if "clicking on things" is maybe 1-5% vs "looking at things" THAT'S THE CRITICAL PATH.
Once the app rendered a view obviously it is not doing much but user is also not waiting on anyting and is "being productive", parsing whatever is displayed.
The critical path, the wasted time is the time app takes to render stuff and "but 99% of the time is not doing it" is irrelevant.
It's all horribly performing turd that needs a 4000$ MacBook to be tolerable at best.
- Slow DB queries
- Lack of concurrency/parallelism
- Lack of caching/memoization for some expensive thing that could be cached
- Excessive serialization/deserialization (things like ORMs that create massive in memory objects)
- GC tuning/not enough memory
- Programmer doing something dumb, like using an array when they should be using a set (and then doing a huge number of membership checks)
With that being said, I have worked on the odd performance optimization where we had to get quite low level. For example, when working on vehicle routing problems, they’re super computationally heavy, need to be optimized like crazy and the hot spots can indeed involve pretty low level optimizations. But it’s been rare in the work I’ve done.
This article is probably meaningful for people who work on databases, games, OSes, etc., but for most devs/apps these tips will yield zero noticeable performance improvements. Just write code in a way you find clean/maintainable/readable, and when you have perf issues, profile them and ship the appropriate fix.
There is this popular wisdom that security must be designed for from the start, and cannot be just added after the fact. Performance is like that too, except worse, because you actually can add security after the fact - worst-case, you treat the entire system as untrustworthy and wrap a layer of security around it. You can't do that with performance - there is no way to sandbox your app so it goes faster. You can only profile things then rip out the slow parts and replace them with fast ones - how easy that is depends on the architecture and approach you adopt early in the project.
In a number of cases you actually can. Casey demonstrated it in his Refterm lectures, it's caching. You still call the slow thing, but at least you don't call it as often because you have that layer of caching to partially insulate you from its poor performance. Good luck if you have to deal with cache invalidation, though.
I'll concede on saying that performance and security are alike - you can add some of either after the fact, but you're better off thinking about both from the start.
> Good luck if you have to deal with cache invalidation, though.
Ain't it the truth. Adding a cache is easy. Understanding the implications of doing it is harder.
In my experience so far, it's typically the case that the 20 second DB query will annoy people for months or years - it won't get solved until enough people raise enough of a stink that someone finally prioritizes it. A large customer suddenly starting to make vague hints about bad performance is sometimes (but not always) helpful.
Some may say that "annoying, but not enough to make a stink about it" means it's fine to not optimize it. But I found that people can suffer a lot, and it doesn't mean it's harmless. People will adjust their workflows to minimize the frustration. When some of your "victims" are in-house users, the "not important enough" performance issue may be silently but continuously losing company money.
Among the many things I’ve learned as a self taught dev: you can be the person who raises enough of a stink if you care a lot. It’s not a thing you want to invoke frequently, but it’s a thing you very probably have power to invoke where it matters most. If you can make a good business case (or any case for user success that impairs your org), you have very good odds of being able to pursue it in any but the most toxic situations. If you can link $thing-you-want-to-pursue to other probably shinier biz/org goals, you’re 99% of the way there.
His final code is definitely simpler than the alternative, which would probably involve several files in another environment.
Sure, he does reach for a benchmark, but that's merely to demonstrate the end result.
I seldom wish I could post images on here, but I really would love to share a photo of the surprise coffee mug my work sent me, with a screen cap of a completely obliterated Y axis on a performance monitoring graph. Granted I don’t get to spend all my time hunting optimizations, but some teams/orgs/companies do very much value performance very explicitly.
Edit: and I’m definitely not a game dev. Though I’ve been itching to borrow some game dev techniques that are quite applicable for my domain (particularly ECS, entity component systems, which I suspect have far broader applicability than their adoption outside of game dev).
Where our opinion seems to diverge is where I accept that as being fine, for the sake of being clean/legible/understandable.
I think the core problem here is that he assumes that everything is inside a tight loop, because in a game engine that's rendering 60+ times a second (and probably running physics etc at a higher rate than that) that's almost always true.
Also the fact that his example of what "everyone" supposedly calls "clean code" looks like some contrived textbook example from 20 years ago strains his credibility.
Edit: come to think of it, the only person I know of who actually uses the phrase "clean code" as if it's some kind of concrete thing with actual rules is Uncle Bob. Is Casey assuming the entire commercial software industry === Uncle Bob? It's like he talked to one enterprise java dev like 10 years ago and based his opinion of the entire industry on them.
Regardless of scenario I will never willingly do a O(n^2) sort when writing new code. Just in case those 10 items suddenly turn to 10000 one day.
Even languages that are notorious for having tiny libraries, like C and JS, have built-in sorts.
The point is that the developers may think O(n^2) is fine because their toy use cases had n=10...100, but then actual users will try to use the software for n=10k, or n=100k, and then either waste their lives working with suddenly slow software, or look for alternatives.
I walked into a case like this the other day. I wanted to do a little semi-collaborative project planning. I found a nice tool, played with it for a moment, figured it has the functionality I need and it's fast enough. Then decided to do the actual plan. Once the number of entries in the system went from 10-20 to 30-40, I started to feel things get a little laggy. 50-60, more laggy. At this point I was committed, so I suffered the tool for couple of months, as its UI kept breaking when handling 100 entries. If I knew this would happen at the start, I'd look for something else. But instead, I walked into a hidden O(n^2) somewhere, that makes me hate the product with a passion now.
Generally, Casey seems to preach holistic thinking, finding the right mental model and just write the most straightforward code (which is harder than it looks; people get distracted in the gigantic state space of solutions all the time). However this requires 1. a small team of 2. good engineers. Folks argue that this isn't always feasible, which is true, but the point of these presentations is to spread the coding patterns & knowledge to train the next gen of engineers to be more aware of these issues and work toward said smaller team & better engineers direction, knowing that we might never reach it. Most modern patterns (and org structures) don't incentivize these 2 qualities.
That doesn't seem quite right. as 100 * (100^2) <<<<< 10000^2
Basically, performance doesn't compose well under current paradigms, and you can see Casey's methods as starting from the assumption of wanting to preserve performance (the cycles count is just an example, although it might not appeal to some crowds), and working backward toward a paradigm.
There was a good quote that programming should be more like physics than math.
Part of the compiler was O(N^2) in `let` block nesting depth. That is
let x = foo(), y = y, z = 2y
...
end
would be a depth of 3. It didn't seem like that should be a problem, N is never going to be 10, let alone 100, right?Until suddenly, `N` was in the thousands in some critical generated code spit out by some modeling software, so that handling the scoping introduced by `let` suddenly dominated the compilation time...
If you are shipping a binary to your users that will never be able to get updates, your cautiousness would be justified. There are other situations where it will needlessly limits options.
There are situations where I have knowingly written O(n^2) or worse, and put a # xxx dragons marker by it. Quick to write, leave my options open, keep my momentum on the problem I care about.
I will grep for xxx issues at some later time. I may end up throwing out the code before that happens. If I hit big-o issues before then, I can refactor.
I once had a system where the important problems turned out to be a series of IO bottle-necks - nothing to do with computation - but that was obscured because good sense had been burnt at the altar of compute efficiency.
He does have a narrow view, but it does not make his claims invalid.
I liked that his POC terminal made in anger made the Windows Terminal faster. But even in that context it was clear that by making some tradeoffs - which the Windows Terminal team can not make (99.99% of users do not run into the issue, but Windows has to support everything) - it could be even a lot faster.
So we live in a world where we cater for the many 1% use cases, which do not overlap, but slows down everyone.
Many gamedevs do their own tools, because they are fed up how slow iteration is. The same thing is happening at bigger companies, at some point productivity start to matter and off the shelf solutions start to fail.
This is perhaps the third time I've posted this on HN, but what you describe is the circle of life for widely-used software projects. Large tech companies are not immune to it, resulting in frequent component rewrites, deprecations and almost-drop-in replacements that shuffle complexity up or down the stack.
Step 1: Developer is fed up by how slow/bloated current incumbent is, so they write a fast, lean and mean project that solves their problems
Step 2: The project becomes popular on its merits, rakes in stars on Github as people discover how awesome it is
Step 3: Users start discovering limitations for their use cases, issues and pull requests pour in
Step 4: Thousands of PRs later, the project is usable by most people and has "won". It is now the incumbent, but no longer is as fast as it once was, but it also ships functionality catering to many niche needs
Step 5: Go to step 1
But i also think that the product, library or framework owner should really box in its project and reject wild growth of features and prevent generalisation of the usage.
Like, feel free to fork away if you like. The core repository needs to be simple and stay true to its goals, and when it updates everyone downstream can update if they want to do that to themselves. But for what it is and does, maybe the project as it is is good enough.
I start feeling almost physically sick when I see the potential for bloat to creep into the software I write. This makes working with scaled software development with others particularly hard, however.
If that's his complaint, then "clean code" isn't the problem. The problem is capitalism and/or human nature.
Once something performs acceptably well, ie good enough to sell it, performance isn't going to get any better. Flashy stuff and features get you money, going from 400ms to 100ms gets you...nothing.
> The problem is capitalism and/or human nature.
Fundamentally, yes.
According to Amazon [0] that'd be a 3% gain in sales (assuming the inverse holds true as getting slower, anyway).
[0] https://www.gigaspaces.com/blog/amazon-found-every-100ms-of-...
It was literally a weekend POC, and Casey Muratori even went beyond the POC part and fixed some emoji/foreign language bugs that were present in the Terminal.
Also, his intent was not to replace the Terminal. His intent was to demonstrate that it was possible to do the optimization in the way he suggested. Originally a Microsoft PM dismissed his suggestions and claimed it would be a "doctoral thesis project" or something.
All this "yeah it's a narrow view" is just moving the goalposts more and more. Not only he has to do a "doctoral thesis project" in a few days, he also has to completely replace a tool that's already written, bells and all? Where does it stop?
I would tend to disagree on this, specially when claims come from the gamedev world. Games are presented as finished pieces (even when they aren't), and not just a release milestone. Ideally, a game is a one-off effort where you write a piece of code and if you're lucky, you won't have to touch it again. So, doing one-off optimizations instead of focusing on milestones and long-term maintainability of the code is not only a possibility, but actively encouraged. That's why until rather recently (20 or so years), assembly optimization for critical execution paths, if not for most of the product.
Most of the rest of the software doesn't work like that. You often implement something that will be maintained, modified, extended and reiterated on for several years, not by you, but by several other teams with totally different experience and backgrounds. Or decades. Doing some fancy trick to skip a cleaner, extensible, maintainable design because you shaved off a couple of cycles on it is literally burning your employer's money and potentially causing huge issues in terms of maintainability, as many programs don't actually rely on a happy path like games do.
The main reason modern systems are slow isn't (just) because programmers are lazy - Its because most software - unlike games - have compatibility and maintainability requirements, and more often than not, a huge legacy support. And also, in these systems, most of development time is actually spent maintaining and extending existing code, not writing new one.
The author's assertion is fundamentally wrong, because software engineering is quite more than performance - even when it matters. Flashback to the beginning of the 90's, and "every game" used bresenham's algorithm to skip usage of the (slow or non-existent) div instruction. In some cases, a couple of bit wise shifts would also eliminate mul operations. These implementations were in some cases 2-4x faster than the classical counterparts, on a 12-40Mhz machine. Two cpu generations later, the Pentium comes out, and both mul and div take 1 clock cycle. The fancy pants implementation is now 3-5x slower at the same speed. Except now the cpu clock is 4x faster and shoveling around registers may actually impede parallel execution of code. All of this in a 5-year window. I envy the relatively stable instruction set of the last decade, where everything is sort-of predictable and assertions of speed can be made on code with a relatively high degree of confidence, but the reality is, silicon is cheap, and for most applications, performance is gained not by throwing away what makes some huge applications barely maintainable, but by deploying hardware. New, fancy, faster, cheaper and more economical hardware. Choosing a single metric (performance) and an instance in time to bitch about something is actually a disservice to the community at large.
Where are you getting this information? Agner[0] lists DIV as taking 17 cycles at best (8-bit operand already in a register) on the P5, and MUL as taking 11 cycles. Even Tiger Lake takes 6 cycles for DIV.
There are ways [1] to beat that, but I don't think you can get it down to a single cycle.
[0]: https://www.agner.org/optimize/instruction_tables.pdf p.162
[1]: https://lemire.me/blog/2019/02/08/faster-remainders-when-the...
Ideally a game pulls in over a billion dollars per year, every year for over a decade. Think World of Warcraft or Fortnite, not Flappy Bird.
Engineering tradeoffs are real. But hiding behind them every time when it can be pointed out that they don't actually apply - and when demonstrated with concrete evidence - is another thing altogether.
If there's something wrong with that advice, I can't imagine what it is...
It will start getting really annoying when you try to add shape ‘hexagon’ and need to figure out all the places where a shape can potentially be used, just so you can update the switch statements.
Conversely, in a codebase organized by objects it's not clean to add an extra method to the base class and each subclass. You have to write an external function and switch over every known subclass inside it, which is also very ugly and will also break when you add more subclasses.
The two designs are actually the duals of each other. Someone compared it to rows vs columns and it's a great comparison.
In OO, the methods are columns and each new row is a new subclass implementing them.
In FP, the types are the columns and each new row is a function that switches over the possible types.
https://journal.stuffwithstuff.com/2010/10/01/solving-the-ex...
He also discusses it in his book Crafting Interpreters.
The author of the video discovered that C++ compiler is dumb when it comes to optimizing virtual method calls (that instead of bare virtual method calls he had to help the compiler to guess the right conditions where these virtual calls could be replaced with guessed static calls). Essentially, all that his video is saying is: "virtual calls bad if-else good". Which is like what every C++ game-dev thinks after few years on the job. Which is amusing in how short-sighted it is, and sometimes even more amusing to discover the "solutions" created by such C++ game-devs that are aimed at replacing C++ objects, but do it in a way that's even worse than C++ original design (who would've thought that to be possible!?)
If your cases are defined externally, and you need to be forwards compatible, omitting a default is wrong.
The Swift language specifically added `@unknown default` for switching over enums.
But we can just re-up the problem by adding 100 different shapes instead of the one. Now you have switch statements with 104 cases each spread through your codebase.
This is the "expression problem": https://en.m.wikipedia.org/wiki/Expression_problem
It absolutely is, on toy problems like the one described in the article.
It very frequently is not when embedded in much larger domains as part of large projects maintained over years by teams.
One of the nice things about some of the clean code concepts he uses is that (as he shows) you can tactically step back from them in key, performance critical areas and reap these wins.
If you stay too low level you get lots of tangled spaghetti code with major performance problems and no obvious way forward besides "make it better."
Using procedures/functions is not exactly "low level". Using switch is not low level. Lookup tables are something you have to do in high level code all the time.
Sure he could have used much better variable naming (CTable?) and probably documentation, but code-wise there's nothing that screams low level there.
I would consider most of his replacements lower level than typical the clean code practices he critiques (especially the ones like iterators that he mentions but avoids in order to steel man the clean code side a little bit), not the lowest level possible. They take into account how the machine actually works and avoid additional indirection which is why they perform better.
Take Listing 27, getAreaUnion for example:
f32 const CTable[Shape_Count] = {1.0f, 1.0f, 0.5f, Pi32};
f32 GetAreaUnion(shape_union Shape)
{
f32 Result = CTable[Shape.Type]*Shape.Width*Shape.Height;
return Result;
}
Is represented quite straightforwardly: {-# LANGUAGE OverloadedRecordDot #-}
data Shape
= Square { width :: Float, height :: Float }
| Rectangle { width :: Float, height :: Float }
| Triangle { width :: Float, height :: Float }
| Circle { width :: Float, height :: Float }
cTable :: Shape -> Float
cTable shape = case shape of -- The "lookup table" or "array"
Square {} -> 1
Rectangle {} -> 1
Triangle {} -> 0.5
Circle {} -> pi
getAreaUnion :: Shape -> Float
getAreaUnion shape = cTable(shape) * shape.width * shape.height
Although it is typically easier to abstract in a "high level" language, abstraction does not require it. This whole debate is rooted on false assumptions and the need to take a side, imo.
Casey has a point, it is just ignored in a typical hand-wavery fashion. "The toy example doesn't scale" is a poor argument, especially when what we can observe is slow software.The stuff proposed in the post is not rocket science, it is a very straightforward implementation of tagged unions. Instead of fetching a vtable and jumping to a value there, he proposes to branch on the tag. This is essentially dynamic dispatch on a known set of types.
Additionally, he shows that this can result in speedups greater than a factor of 1. Any program that wants low latency or high throughput can profit from this observation.
This way of programming is by no means the one to rule them all. It has different advantages and drawbacks; none of which have anything to do with the percieved intelligence of the programmer or later consumers, for that matter.
An objective disadvantage of this style is, that the program can't interface with code, that hasn't been written yet, as a caller. Another disadvantage is that the size of the tagged union is defined by its largest "subclass".
In the end, what he has shown is that speed is often a compromise made unnecessarily. This doesn't really have to do with clean code anymore, as I can see how a compiler could implement what he is angry about with virtual functions in every situation where his style is applicable.
Casey has had a similar thing about the windows-terminal and somewhere in his videos a different, yet arguably worse, problem comes to mind: a lot of libraries do not care about performance enough. If you write a program and care, you may run into the problem that the library you use is your bottleneck. If this library is hard to replace (imagine needing a rocket-scientist), then you are done for. In that specific case it was DirectWrite and some other Windows-API that were slow. So if, for one reason or another, the windows team was required to use both, they'd have a hard limit on how fast they could go, just due to that. There is no "being smart" or "requiring a genious" involved in the forced/strongly recommended library here.
The claim that "Clean Code" scales better or allows for more maintainable software hasn't been proven by anyone, and everyone with enough experience has worked with several counter examples.
The problem of code maintainability is not solved by this coding philosophy.
I'm totally onboard with prioritizing readability over performance for most code, but the style in the book has a lot more tradeoffs than it discloses, and you often don't really appreciate that until you are trying to debug if/how A transitively calls Z in a 10 million line codebase.
It's really hard to have a constructive conversation about this though since it's so subjective, any example is too trivial, and any real system is too large.
If you can't point to an actual example, I don't think you have a solid case.
However, the bigger your task gets the more value there is to polymorphism and in general the smaller the percent of total time goes to the polymorphism overhead.
And note that his attack is only on polymorphism, not the other aspects of clean code. I strongly suspect the compiler optimizes away much of the clean stuff I do but I have never checked. I also find profiling easier on cleaner code, it makes it very obvious where the time sink must be and thus what warrants expending effort to improve. Profiling almost always shows the vast majority of time going into the unavoidable (say, disk reads) and a small number of other routines. Spend your optimization effort on the spots that need it because 99+% of your code doesn't run often enough for it to matter.
Casey's definition of simple actually reminds me of the cliché by Rich Hickley. It's simple, but it's not necessarily easy.
ITs still possible to get a bottleneck in assembler.
Whatever language is used, executable code still needs to be profiled using tools as described here.
https://en.wikipedia.org/wiki/Profiling_(computer_programmin...
It's uncharitable to take Casey as making absolute blanket statements like that, but still, it would not be unreasonable for him to single out Uncle Bob in particular.
The Amazon rankings for Bob Martin's "Clean Code":
Best Sellers Rank: #5,338 in Books (See Top 100 in Books)
#1 in Software Design & Engineering
#2 in Software Testing
#4 in Software Development (Books)
Believe it or not, callbacks are not like interrupts, there is a loop somewhere that checks the status of something and then runs the callback. All computer software today involve things that run in loops, you just don't see it. Web browsers do it! Of course they do.
Moreover, he didn't contrive his example, he said that he in fact used a textbook example used by the advocates for polymorphism and such.
EDIT: To expand on this, why is modern software slow? The reason is because people, thinking their code isn't a bottle neck or performance doesn't matter but adherence to the right abstractions is, they write slow code and call backs thinking everything just happens immediately, with layers of abstractions, and those little innocent steps add up when every piece of code written today is written with that same neglect. The call backs runs slowly, queues fill, promises hang, and so-on and on we go.
So sure, may be your code doesn't seem like it needs to run at 60fps. But when everything I do on a computer is written like it will be run every 10 seconds, it definitely will be noticeable. That's because I don't just run your program or look at your website, I look at 10 of them or even more. If I average all of those per time, then may be your code should in fact be able to run at 1Hz or so or I will start to notice.
People of course were right not to teach new devs not to over-optimize immediately, but the culture has swung so far in the other direction, especially since you all seem to love complexity so much, you've managed to yes make computers that can calculate pi faster than a super computer from the 70s crawl when it renders and handles an editable textbox. There has to be a move back in the other direction, you guys need to give a shit about performance, at least a little.
while (command != 'quit') {
command = readline()
handle(command)
}Game developers have the luxury of starting from near-scratch every once in a while. That exactly what his lauded handmade series is all about. I'm guessing that things wouldn't be so clear-cut if he was given a 10 year old codebase to iterate on.
"Normal" developers see game developers as gods walking amongst us and place far more value on their opinions than they should. The truth is that game developers and "normal" developers face equally as challenging problems, just different problems. As a trivial example, an experienced web developer could probably run circles around Casey in terms of elegantly accounting for browser quirks (conversely, the web developer would probably be stumped about data oriented design). Either could learn the other's discipline, but each would have decades head-start on the other.
The idolization of gamedevs is extremely frustrating, especially when it comes to appeals to their authority.
Every story needs a Hero; it's inspiring when you're trapped in the CRUD gulag (until you see the TC/WLB).
I find what Casey says in his videos to be true. And I though about that stuff before I even watched his videos, which are excellent.
However, I started not to care. I don't want to start fights inside the company, especially fighting alone against many OOP cultists. There's not my money at stake, so if companies as a whole decide for OOP, clean code, SOLID, design patterns, abstractions on top of abstractions, making the code bases giant pile of junks while degrading performance, I am not going to go against the crowd.
Code that I write for myself is quite different than code I write for my employers.
I just hope that the industry as a whole will wake up from the whole OOP nightmare.
I agree with that 100%. OOP is a giant mess, no matter where your stance on code clarity or performance stands. It's objectively worse in both regards.
This is simply not true, and has in all likelihood worked on such problems given his work at RAD whose software has been used in +20 years at this point.
> The problems of software performance come from decades of poorly/quickly executed evolutionary change resulting in bad systems design.
This may be true of some code bases, but it's demonstrably false for new software that's created today. Lots of new software gets built and it's slow.
I agree that not everything is like a game, but it makes me legitimately sad when it seems like nobody cares about performance (aside from a few domains).
I work in game development, largely with optimisation. I mostly work with GPU optimisation, which is a whole different beast. On the CPU, most of the time issues are either trying to do too much stuff in a hot loop (rendering stuff that could have been culled, putting physics on objects that don't need it,...) or doing something in a slightly inefficient way in a hot loop. Because everything in the game is indeed a loop consisting of a series of hot loops.
People in this comment section call his example contrived, but it's very similar to one of the biggest performance improvements I've seen in practice.
Still, each triangle's position, shape, and other properties can change each frame, as does that of the camera. So you cannot avoid doing some amount of work for each of the visible triangles and their vertices each frame.
Since you need to update the screen at a consistent frequency (typically 30 or 60 times a second) and the list of triangles that actually need to be rendered each frame is in the millions... Well, that's a lot of work which cannot be avoided.
As a former .net developer that was often pushed into "clean code", my big takeaway from the video was that actually, not using "clean code" techniques, such as polymorphism made the code so much more readable and easier to grok that the optimisation that followed was completely natural.
It's the same thing with TDD zealots. Or any other fad driven development paradigm, which our industry is filled with.
With a non-pessimal design, not only you are able to pay less money on servers in the long term, you're also able to delay complex scaling strategies. Scaling also costs money.
Not to mention that building something with this kind of overhead also means that a lot developer time was spent in the first place, which is still expensive in our industry.
>Virtually every distributed service built today is able to take advantage of this
most of built today services can take much more advantage in using better system design practices.
If your program isn't slow then you don't need to bother making it any faster. Even Michael Abrash, who specializes in code optimization, once explicitly wrote in his graphics programming black book that "The objective (not always attained) in creating high-performance software is to make the software able to carry out its appointed tasks so rapidly that it responds instantaneously, as far as the user is concerned. In other words, high-performance code should ideally run so fast that any further improvement in the code would be pointless [..] Notice that the above definition most emphatically does not say anything about making the software as fast as possible".[0]
[0] https://www.jagregory.com/abrash-black-book/#understanding-h...
No, though s/any/many/ would make this true.
There are no hotspots. Computing has become lukewarm by default, and just like the globe, it's slowly heating up.
What problems do they prevent?
If you're able to make a program into a sort of puzzle concealing its state inside a twisty maze of tiny virtual functions so it's impossible to see plainly how anything is done, then you become indispensable as the only one who's internalized how the thing works.
And then you insist it's all for easier comprehension… for those sufficiently intelligent to comprehend it.
No it wont say something like that but if you profile Python itself and see that a lot of time is spent in Python you can get a hint that rewriting it in another language that doesn't have Python's overhead might help.
But you can only judge that for the very limited set of hardware you personally have access to. You might have many users -- or potential users -- with slower CPUs.
(ignoring projects made for fun of course)
Of course it will.
If your backend service is already suboptimal, and running at 10x worse performance, optimizing that will give you, well, a 10x performance boost.
Imagine replacing poor in-memory reimplementation of database queries that most graphql servers do with actual opttimised database queries. And a better code on top.
Boom. You're operating close to the speed of light.
but you are actually talking about optimizing system design, and not reducing virtual functions calls :).
And that the point of this thread: you need to optimize parts that slow you the most.
So no, in most cases optimizing virtual calls won't bring you from 1s to 30ms
As an example, I recently worked on a large system that was optimized to do a big data transformation in an efficient way. It turns out that data is transformed back to the original format later downstream. So much for the optimization...
All that is to say, often it is the case that "simple > fast"
Clean code at the time had a lot going forward it, and it was an improvement over a lot of JavaEE code that was written in absolutely procedural ways with less care for the developer reading the classes, functions, or individual statements compared to punching out near assembly-like code and moving on.
On slower languages with slower runtimes, something that is 10x slower than normal code will have much more overhead than in the examples demonstrated by Casey. It won't be about "30ms vs 25ms" as some people are saying. In the past I remember seeing differences between 400ms and 20ms between JBuilder and .to_json in a critical endpoint in a Rails app, to give one example. Sure, one is "cleaner", but in the end it's a 20x overhead that has no place this case.
Also, the myth that "processors spend time waiting for IO" that is spread across this thread is BS. In reality, that's only true for single-user programs. If the app is part of a distributed system, the CPU time can be used to serve more users. This allows you to significantly delay more complex scalability efforts, which is also precious developer (or DevOps) time.
Not to mention that applying "Clean Code" in the first place also takes precious time, which could be used for features or anything making money, even optimizing the DB. Instead, this time is used to mess up the code in ways that have zero proven efficacy, and some developers instead think are terrible.
If you're working with databases, performance often boils down to minimizing the number of times you have to access the database (over a multi-service stack, not just a single web service).
When you hit Amdahl's law, it's because you (or someone else) has made decisions about the high level design/patterns to use in a system. To remove such bottlenecks, you may have to scrap the entire project and start over.
For the inner loops, it's perfectly fine to leave most of the optimization for later. But for overall design, it needs to be right from the start.
Time and time again, I see devs stuck in some paradigm (often front-end devs or low volumen RESTful microservices) that makes it almost impossible to handle non-trival data volumes or traffic, causing new products to fail.
I do agree that overall design needs to fit performance requirements but for the most part that has nothing to do with "clean code" patterns.
Now, your argument seems to be: in the real world, there's so much waste, that virtual function calls pale in comparison.
This does not debunk his main point, which seems to me at least the following: all things being equal, writing code with virtual functions that do a tiny amount of work and "hiding implementation details" makes performance worse, sometimes by an order of magnitude.
Now, there maybe situations where you _have_ to use virtual functions, because you are writing a library for other people to use, and you can't dictate ahead of time how they will use it.
This again does not invalidate the point. You need to be _aware_ of the performance implications of this, and mitigate it. He said the following in the comment section on the article:
> Try to make it so that you do very rare virtual function calls, behind which you do a _large_ amount of work, rather than the "clean" code way of using lots of little function calls.
but all things are not equal. You can spend a lot of time improving performance of you function calls and get virtually nothing out of it. Because if you optimize something that takes 0.01% of overall execution time, 'order of magnitude' performance gain is still negligible.
Also articles like this usually fail to mention code maintenance cost. For example by reducing usage of virtual calls you can make your code unmaintainable/expandable and suddenly every new change will cost you 2x more in development time.
That's why in the real world most of the time you choose clean code and you use optimized nonclean code only on places where you need it. If you look at any lets say web framework internals, you will find a lot of non-clean code, which makes framework faster. But an interface will be done in clean fashion and most of user of the framework will enjoy clean code without need to care about unclean internals.
> Also articles like this usually fail to mention code maintenance cost. For example by reducing usage of virtual calls you can make your code unmaintainable/expandable and suddenly every new change will cost you 2x more in development time.
This never actually happens in real life. I've never seen a codebase that is written with "clean code" principles in mind that is also maintainable and easy to develop on top of.
Don't forget virtual machines and interpreters.
If shape Area is computed often enough that you care about inlining the calculation, why not compute & store it every time the height / width change. That’d be easy enough in an architecture based on information hiding, and might illustrate a legitimate engineering trade-off between those architectural choices.
Indeed. That's the point. Why let someone with an axe to grind define the problem in a way they can solve with their axe? The author is using optimization as the frame for rejecting a particular way of working, I'm pointing out that the definition is set up to make the solution work. I don't concede the terms of the debate before it even begins.
Maintainability and cleanliness is the best virtue code can have. If you have extremely clearly written code that has performance issues I can swoop in with analysis tools figure out where the pain point is and refactor it out. Sometimes this is a real headache[1] sometimes not - what I can guarantee is that if the code is "dirty" it's going to be a headache and it'll take more time.
I'd personally take issue with this article over the polymorphism claim though - polymorphism is a tool but it isn't the be-all and end-all tool. A lot of your data can live as structs/blobs in memory with tight internal type definition but without any OO principals. Personally I am a huge fan of functional programming (but not pure functional programming) so objects that I use are relatively few and far between and exist to fulfill a very specific purpose.
I've had two occasions in working when I needed to break out an asm block - the compiler was being a thick headed dummy and this code needed to receive incoming signals without exception or delay - but once that critical section was passed? Back to high level programming and statements favoring expressiveness over raw bare metal performance.
If you want an interesting experience talk to your closest non-technical manager type - be that a product team manager or the company owner - and ask them if they'd prefer if you focus on reducing how long your product takes to execute by 20% over the next five years or if they'd prefer you to lower the growth of the developer labor budget by 20% for the next five years by focusing on maintainability over performance. With the exception of extremely niche cases maintainability is always the golden standard.
1. For instance, I've dealt with OOM issues that have required transforming all logic on a query result to be lazily evaluated on a data stream after main execution finishes - like the logic goes up and down the stack and only then begins processing results. In this particular case the problem was rather easy to deal with because we essentially swapped out the actual value passing on each layer for a lazy result set being passed around - because the code was clean. Sometimes you'll definitely need to massively re-engineer things though.
I feel like there's a misunderstanding here. Casey is clearly not against writing non-capitalized clean code at all. His code in the end is "cleaner" than what he criticizes IMO. What he is criticizing here is capitalized (and possibly trademarked) "Clean Code", the book and philosophy spearheaded by Uncle Bob.
Definitely get the single-threaded house in order before attempting to speed up by running in parallel.
You can of course only optimize what you are looking to optimize. I am not surprised (honestly) that some engineers do not realize they are in fact primed for the kind of things they will find based on what they are looking by the mere choice of where they want to look.
Honestly, for most software, a well performant "monolith" is probably enough.
We had one service that had a single-thread bottleneck that none of the developers could configure out - the solution was to spin up 30x more virtual machines to run instances of this app to meet average production demand.
Along with the heuristic that hardware and electricity are cheap, developers are expensive. That's probably why managers in my experience almost never ask developers to optimize slow code - they believe it would be cheaper to use more hardware if this fixes the problem (in some areas like HFT or GamDev more hardware is not the answer so optimization do happen). In rare cases I've seen optimization being done initiative always come from IC (who knew that the code could work fast / use less resources).
So nowadays writing code I assume that it never will be optimized later and try to do less dump stuff from the beginning.
Experienced devs - such as yourself - generally know how to write reasonably fast code that is also clean (or easy to extend) first time. In my opinion we should be more explicit about this in the heuristics.
We got slow software because we were ok to get slow software. If being fast is not in requirements it means it doesn't matter (for whoever is responsible for defining priorities). Places where performance matters it is never sacrificed.
I am not okay with slow software, and being forced to use it is frankly insulting. Teams (to pick a punching bag) hogging my system resources doesn't impact its adoption since I'm forced to use it. Teams could be 10x slower and I'd probably still be forced to use it. But it wouldn't be because speed doesn't matter.
What I've seen is slow queries but a bigger problem is actually too many queries. It's easy to do especially when using an ORM.
It mostly happens on change when you want to add something to an existing query the changer just add their new query and slop it into a loop, boom performance is gone.
There's a tipping point where you have so much hardware there's big savings with optimization. Things like Postgres and the Linux kernel have a lot of optimization put into them and there's an insane amount of hardware out running that code.
That said, slow software sucks.
Over the years they've grown and one bin showed up where jobs would sit around for weeks--for that case n went from rarely having a screenful of lines to a few thousand--oops, now more than 2/3 of the time was spent in those searches. (And I have a sneaking suspicion that a good portion of the remaining time comes from using field names to retrieve values. The profiler doesn't separate that out, though, because it's not my code.)
- shit UX "ideas" that trigger 10 new API calls to show dialogs and popups.
- Logging libs under pressure
- overemphasis on testability
The word is TENET. Tenants rent stuff.
I think a lot of the clean code advice in general related to object oriented programming.
I've noticed that once my Lua programs (games) grow to reasonable size, it becomes kinda hard to maintain. And I tend to use an object oriented programming style (of course it also doesn't help that Lua is not typesafe). After I finish my current game, I want to try to make a game using a procedural approach. I wonder if this would solve some of the issues I see in my current code base.
One of the core ideas of procedural programming is that data and functionality is not mixed in classes as we do in object oriented programming. Instead, you might have a module that contains some functions and some data objects the functions act upon. This approach would make some other aspects of game programming with Lua easier as well (e.g. serialisation), but perhaps it will make the code also easier maintainable as the size of the codebase grows. It's something I want to contemplate upon.
For instance, if you write a function to do some operation on an object, you could have written that as a method instead. But ultimately it is the same code, it is unlikely that the difference matters much for either performance or readability.
However if you need to do some operation on a bunch of objects you could pack each operation on the individual objects in a method, and call those methods from a main function. Or you could just put it all in one function, with as many nested loops and if statements as there needs to be. Now the difference is real, you pay in performance for a lot of function calls, and following the control flow is different.
Personally I tend to prefer the one function, but sometimes part of it makes sense as its own function, in particular when I can avoid duplication that way.
There is no silver bullet, but best of luck, changing one's style can be hard.
Oh, it's not. Look at the crap the Java community used to produce. 37 levels of abstraction is not maintainable.
All that says is you should focus your energy on the increasing the value of .1%. It's not actually an argument to not spend any energy.
It's like saying 'Astronauts only spend .1% of their time in space' or 'tomatoes only spend .1% of their existence being eaten' - that .1% is the whole point.
You can debate how best to maximize that value, more features or more performance. The OP is suggesting folks are just leaving performance on the floor and then making vacuous arguments to excuse it.
you're right. some things are inevitably slow so we should therefore never care about performance in any situation.
avoid belittling the efforts here just because they don't apply to all situations.
Assuming it is right, there is something called multitasking, the CPU, RAM, and most importantly, the cache is not all yours, if there is 1000 pieces of software like yours, that's 100%. You may argue that 1000 pieces of software is unreasonable, and you would mostly be right, but it happens, and mostly for the same reason software isn't optimized: quantity over quality.
Another issue is that you have to make a distinction between throughput and latency. You don't have to keep up with a sustained 100 actions per second, people don't go that fast, but you definitely have to respond within tens of milliseconds, because more than that is noticeable. Latency is much harder to optimize and if you are in the critical path, these cycles may matter.
A lot of devices are battery powered these days, and all these wasted cycle are reducing the battery life of the entire system. Mobile devices are crazy powerful these days, but this power is meant to be used sparingly. And even with line powered devices, I think we waste enough energy as it is...
And finally, what is the point of "clean code"? Hopefully not just because it gives software architects boners. The point is usually to make software that will last: easier maintenance, less bugs, etc... But performance bugs exist too, and one of the most common software evolution is to do more of what the software already does. An image editing software will process more and bigger images, a database will store more entries and more details about each entry, documents will get larger, etc... You may even find that your users are using some feature on a scale you never intended, maybe someone is pasting entire books on your note taking app, and it may turn out working quite well... if you cared about performance. Not caring about performance is technical debt, and it may negate the advantage of using "clean code" in the first place.
I don't. how many programs are running in your OS right now? how much CPU do you need to keep those things plus the things you need running in a performant manner?
how much CPU would you need if things performed better? the answer is "less" every time.
better software performance = less money required for hardware to obtain the responsiveness you require.
it's important, and it's important completely independently of how it is framed here.
just wait till you've seen software get slower for 30 years. to put it another way, watch hardware get faster and faster and faster for 30 years while you observe software continually consume all of the available headroom until it feels slow again. watch that happen for THREE DECADES and wait for someone to tell you that everything is fine and that someone saying "software is unnecessarily slow" is wrong because they aren't framing their argument how you think it should be framed.
This works until your computer is old enough to be slower than what a majority of wealthy people (ie desirable customers) are using, at which point you need to buy a newer, faster computer, even though your current one was already "faster than anyone could reasonably need it to be".
This is all harmless enough—a little disrespectful perhaps, to make other people waste their money, but not so terrible—until you consider the environmental impact of all these new computers, which the average spreadsheet absolutely should not need but does anyway. It's also an equity issue—someone on a fixed income can't necessarily afford a new machine.
What would actually happen if Moore's law ended tomorrow, and we were no longer able to make computers any faster than they are today? It would really suck for scientists and hardcore gamers, but I actually think a majority of computer users would benefit The experience of someone who just writes documents and checks email would be unchanged, except that their current computers would never slow down!
Value of time is disproportionately weighted by user attention, which is at its highest right around when user input is happening.
Also, "clean code" (as in from the "Clean Code" book) is generally not good advice for most programs anyway. Not only does it eat performance, it's not all that great for building maintainable, extensible systems.
This talk(Preventing the Collapse of Civilization) by Jonathan Blow disagrees with you.
every place we work at is FAR more likely to scale how many compute instances we are using, than optimize the application.
Wrong focus. Software time doesn't matter, user time does.
Users don't want to wait, even slow typists want instand results as soon as they hit Enter.
That's fine, that's as it should be and isn't an interesting metric. The computer should wait for me, not the other way around. Needlessly waiting for the computer is a sign of s%&t software and not as uncommon as we'd like, huh?
And you're saying for many developers, performance is not the biggest priority.
That's fine. But it doesn't make him even 0.000001% wrong. And he's not applying anything to any niche situations. You just missed his point. Performance.
> I think the author is taking general advice and applying it to a niche situation.
I think the author is taking a general situation and applying common sense to it.
So why isn't my browser at 0% CPU when IN THE BACKGROUND then?
I hear this a lot, but that 0.1% when I'm actually waiting to do its calculation, it better be fast.
And the rest 99.9% of time, it better not eat my battery and memory...
So you might write your code in a straight way, careful to not lose efficiency and then you are going to call slow code.
So his advices are easier to follow on some scenarios than others.
This is bad for sustainability
Bad mindset.
A GFCI breaker spends 99.99999% of the time waiting with zero leakage current. Yet, when it does detect leakage current, you want the breaker to trip as quickly as possible.
See where I'm going?
Imagine if it would take 5 seconds from flicking a light switch until the lights actually turn on. Because the switch is waiting for user input 99.99% of the time, right? Would you install that light switch in your home?
Is it? What processes are running on your computer right now, how many are waiting for your interaction?
Further the 0.1% of the time I do interact, I want the results promptly.
This applies doubly so if you can rely on templates & structural typing to push your polymorphism to compile time. clang & gcc are surprisingly good at optimisation as long when you don't have to bounce off a vtable and code is clean / avoids "manual optimisation".
Also while I'm not saying I don't believe the author here, I wish they would have used https://quick-bench.com/ or https://godbolt.org/ so that readers could trivially verify the results & methodology.
Clean code (OOP, DRY, etc) is optimized for maintainability and extensibility, not necessarily performance.
In fact, I think it’s pretty well understood that clean code is a tradeoff wrt performance, at least that’s the way I’ve always understood it.
Clean code works well for something like a web app that’ll need to be maintained by scores of different engineers over many years or decades.
At least that’s the theory. In practice, at least some level of abstraction makes it a bit easier to rip and replace parts of the app without a total rewrite.
In other words, the user’s entire perception of your program’s performance falls into that 0.1%.
Performance for performance sake is an interesting and appealing challenge to us engineers. I was writing C code in the 90s and I miss being that close to the hardware, trying to spare every clock cycle that I could while working with machines that had sparse resources.
But today I'm building SaaS products for millions of simultaneous active users. When customers complain about performance it is often not what us engineers think of as "performance." They're NOT saying things like "Your app is eating all memory on my phone" or "the rendering of this table looks choppy." It's usually issues related to server-side replication lag causing data inconsistencies or in some cases network timeouts due to slow responding services.
The point is the age old advice that we were giving aspiring game programmers back in the 90s:
Figure out and understand your priorities.
The famous inverse square root function in the Quake III Arena source code is a great example. If memory serves me, they needed this calculation as part of their particle physics engine. The problem is that calculating inverse square roots precisely is very expensive, especially at the scale they were required to. So they exploited how 32-bit floating point numbers are represented in binary in order to do a fast, good enough approximation. This is a good example of a targeted, purposeful optimization.
Back in the 90s we were obsessed over getting the most out of our hardware, especially when coding games. So we picked up all sorts of performance hacks and optimizations and learned how to code in assembly so that we could get even closer to the bare metal if we needed to. The result was impossible to understand and maintain code and so experienced engineers taught us young'uns:
Write clean code first, then profile to understand what your bottlenecks are, then to make TARGETED optimizations aimed at solving performance issues in order of priority.
That priority always being driven by user experience and/or scalability requirements.
Anything else is premature optimization. You're speculating about where your performance bottlenecks are, and you're throwing out maintainable code for speculative reasons; not actually knowing how much of an impact your optimizations are going to have on user experience.
If you are throwing out maintainable code for the sake of performance, it had better be because you know that it's your bottleneck, and that the performance increase is worthwhile in the first place. "Performance for performance sake" shouldn't exist anywhere outside of hobby projects.
I would argue that responsive user interfaces are really important to user experience. Not many people are complaining because everyone is used to unresponsive apps, but that doesn't mean users wouldn't appreciate a more responsive app.
I would also add that there isn't always a tradeoff between performance and maintainability. If you can adopt some performant coding patterns that don't sacrifice maintainability, then absolutely do that. I think Casey's example of "switch-based polymorphism" is one such pattern(and I think the fact that Rust took a similar route to polymorphism is a vote in favour of this pattern).
Well written video games, the kind of thing Casey works on usually, beat the hell out of basically any other category of software in terms of user responsiveness. At least of all the software I use regularly.
The presented transformation of code away from “clean” code had nothing to do with optimisation. In fact, it made the code more readable IMO. Then it demonstrated that most of those “clean” code commandments are detrimental to performance. So obviously when people saw the word “performance”, they immediately jumped “omg, you’re optimising, stop immediately!”
Another irritating reaction here is the straw man of optimising every last instruction: the course so far has been about demonstrating how much performance there even is on the table with reasonable code, to build up an intuition about what orders of magnitude are even possible. Casey repeated several times that what level of performance is right for your situation will depend on you and your situation. But you should be aware of what’s possible, about the multipliers you get from all those decisions.
And of course people bring up profilers: no profiler will tell you whether a function is optimal or not — only what portion of runtime is spent where. And if all your life you’ve been programming in Python, then your intuition about performance often is on the level of “well I guess in C it could be 5-10 times faster; I’ll focus on something else”, which always comes up in response to complaints about Python. Not even close.
For instance he fails to explain why he didn't bother addressing the tail of his unrolled loop. He does in the course, but here he's just assuming it's irrelevant, and doesn't address again potential criticism like "he doesn't even bother to write correct code, look at that lazy unrolling!".
Thankfully there's a free lecture that explains the broad concept with an example. It's this lecture that convinced me to try out his course, where I hope he'll go into more details: https://www.youtube.com/watch?v=pgoetgxecw8&list=PLEMXAbCVnm...
A lot of my performance and code quality chops came from projects where the team was comfortable with the performance of the system but the business was not. They wanted to stop when it was right for them but not right for the situation. It ended up negatively affecting my opinion of them because ultimately I started to see it as deflecting. It's fine because I don't know what else to do, not because this is the best that can be done.
I think it's also important to separate computer science from software engineering. Algorithms and data structures can be reasoned about mathematically, but what about agile VS waterfall or functional VS oop?
Maybe software engineering relates more to philosophy than math. There are theories and some are more sound than others, but there is no objective truth. Most of us agree that it's important for code to be readable, but we still don't have the answer for what's the best readable code. Even first principle such as DRY are challenged, eg by WET.
I recall vaguely that there was a situation where there were two schools of thought with different approaches. One was about hacking away things, the other about having correct programs. Maybe Berkeley vs MIT? Such different opinions at the basic level of a discipline are more likely in philosophy than math, I think.
No one is working with a huge amount of data in big loops using virtual methods to take every element out of a huge dataset like he is showing. That's a false pre-position he is trying to debunk. Polymorphic classes/structs are used to represent some business logic of applications or structured data with a few hundred such objects that keep some states and a small amount of other data so they are never involved in intensive computations as he shows. In real projects, such "horrible" polymorphic calls never pop up under profiling and usually occupy a fraction of a percent overall.
This is largely inaccurate. Video encoders/decoders are typically written in C, with some use of compiler intrinsics or short inline assembly fragments for particularly "hot" functions.
The times that assembly outperforms a higher level language has reduced as well over time, with compiler and CPU improvements over time.
If everything was built with constraints like "this must serve user input so quickly that they can't perceive a delay", we would probably be a lot better off across the entire board.
We should try to steal more ideas from different domains instead of treating them like entirely isolated universes.
Sure but it’s a question of tradeoffs.
If the wall-clock-optimized version creates a bus count of 1 for that domain or is difficult for a dozen engineers to iterate on, then that could be worse for the business and users.
Should we want better software? Yes. Should we learn from other domains? Absolutely. But we should ultimately optimize for the domain we’re in, and writing much product-focused software is ultimately best optimized for engineering team and product velocity.
Another example: would most software benefit from formal verification? Yeah, I guess, but the software holistically—as a thing used by humans to solve problems—might benefit considerably more from Ruby.
I've basically built my career for the past decade by pointing out "yes, we can load our whole working set into memory" for the vast majority of problems. This is especially true if you have so little data you think you don't have CPU problems either.
All in all, I fail to see how it disagrees with my points.
I didn’t see any part of the comment that implied it did
this is exactly how the typical naïve game loop/entity system works.
And importantly, UnrealScript was designed in the 90's, when memory latencies were far less of problem.
Things way worse than that exist. Replace "virtual method" with "service call."
Yeah. I opened discord earlier, and it took about 10 seconds to open. My CPU is an apple M1, running about 3ghz per core. Assuming its single threaded (it wasn't), discord is taking about 30 billion cycles to open. (Or around 50 network round-trips at a 200ms ping).
Crimes against performance are everywhere.
- It makes horribly inefficient use of my CPU
- It needs an obscene number of network round-trips to load
- One of the network servers that discord needs to open takes seconds to respond to requests
This isn't a new problem. Discord always takes about 10 seconds to open on my computer. (Am I just on too many servers?)
It should open instantly. Everything on modern computers should happen basically instantly. The only reason most software runs slowly is because the developers involved don't care enough to make it run fast.
Except for a few exceptions like AI, scientific computing, 3d modelling and video editing, modern computers are fast enough for everything we want to do with them. Software seems to have higher requirements each year simply because the developers get faster computers each year and spend less effort keeping their software tight and lean.
There is truth to that, but also:
* some of them would care if they knew what was possible with reasonnable effort (that's what Casey is trying to address. So far in the course i'm not really seing much that I could apply to the kind of code I write, sadly - but I'm hoping to learn stuff.)
* it's very likely that making performance-aware or optimized code takes just a tad longer than not doing it, and time-to-ship is valued much higher than time-to-run in most industries (this is the point I think Casey is overlooking, or at least not addressing enough. I don't know if it's by design - maybe he disagrees with the trade-off entirely - or if he's biased towards one of the few industries where time-to-run is crucial.)
This makes sense when you're a shiny new startup. But seriously, 10 seconds for discord to open? There's a point in every product's lifecycle where performance is a feature. Discord isn't a startup anymore. Why can't they fix these performance problems? At least discord is pretty snappy once its loaded. The new reddit interface? Its a hog. But despite a massive outcry, why haven't they fixed it?
My pet theory is that they don't know how. And talking about velocity is just a smoke screen.
I think most professional engineers don't really understand the software stack well enough to be able to improve the performance of the software they write. Its pretty understandable - nobody asks about this stuff in job interviews. And the software stack only gets more complicated each year. If you follow React tutorials online, you can get pretty far adding features to a web app without ever needing to understand how react actually works. Or the web browser, and Vite / webpack / whatever and the operating system it runs on top of.
And thats a pretty good deal! More engineers! So long as we don't mind the new reddit site. And electron apps that take seconds to load.
Of course Casey Muratori knows how to write performant code. He understands the whole stack. He knows how to read the assembly that the C++ compiler produces. Thats something more of us should aspire towards.
I wonder if it would be valuable to make an online course talking about performance engineering. I feel like its one of those things that has fallen by the wayside, and I think thats a massive pity.
But hey, those network calls are fast on my loopback interface, or my company LAN, when I'm playing with the dev version, using test set simulating 2 users and 5 posts for each. Surely it'll be just as fast for the real users, over the Internet, on channels with 1000 users and 5 posts per second.
Then again, at some point we had "Lisp machines", maybe some day there will be a computer architecture where memory / computations patterns are adapted to massive simulation - rather than shoehorning on existing architecture.
And those will fail just as miserably as Lisp machines.
I guess only in the places you call methods?
Wait? Are they everywhere? Hmmm...
The reality is that the entire Java ecosystem revolves around call stacks hundreds of calls deep where most (if not all) of those are virtual calls through an interface.
Even in web server scenarios where the user might be "5 milliseconds away", I've seen these overheads add up to the point where it is noticeable.
ASP.NET Core for example has been optimised recently to go the opposite route of not using complex nested call paths in the core of the system and has seen dramatic speedups.
For crying out loud, I've seen Java web servers requiring 100% CPU time across 16 cores for half an hour to start up! HALF AN HOUR!
There is one thing worse you can do (and I caught a C++ compiler doing it when we were profiling code while building an x86 clone years ago) instead of loading the address and jumping to it push the address then return to it, that not only breaks pipelines but also return stack optimisations
I remember doing that in a code generator ages ago because it was easier than calculating the jump offset :-P
It wasn't really used anywhere, eventually i decided to keep the interpreter and move any complex logic in C which ultimately was the simpler approach (and which has been my take on scripting languages for years now: use scripting languages for the "what" and native code for the "how").
The particular example with shapes could be CAD or BIM app where it is usually more then "few hundred times".
So go on all of you, write everything in Python with 90 levels of indirection, my stock will go up.
You're mistaken, the load size has nothing to do with the end result. The result is normalized to give an estimate of how much faster the simple code is than the polymorphic code irregardless of input size. (Kinda like deaths per 100k instead of giving an absolute number of deaths for statistics about diseases).
So yes, your code is running 20x slower than it should be all the time.
Especially when you make every class an interface, with... get this, one implementation! This is based on real world experience and is not a joke. There are real companies with real people that write real code where every single class is an interface with exactly one implementation. Which, as Casey has shown, results in upwards of a 20x slowdown in the worst case.
Obviously, you probably won't get a 20x speedup by getting rid of the polymorphic garbage. But it's equally asinine to assume that polymorphic functions are only called a few hundred times. I guarantee you your PC is making millions of polymorphic function calls per minute between: the OS, the browser, windows Anti-Malware scanner, steam running in the background, oracle running its checks to remind you to update Java, etc. There are hundreds of processes running all the time on a modern device, these devices are wasting enormous amounts of resources.
And when you run this through a profiler, you will not notice how slow your code is, because everything is slow. Slowness is infused throughout the whole system.
As a result some styles of writing code just don’t work for the audio thread at all, and we’d have to simply avoid or rewrite libraries written this way.
There are just some domains where standard practice for cleanliness is different because of your constraints.
I mean, it’s to the point we’ve got die hards in this industry who insist on putting all functions inlined in headers (not that I agree!)
Congratulations. Taps on the back, champagne all around. Customers call. Same complaint.
Programmer asks "Well, did something change at least?". "Loading bar now flickers more", answers customer.
Of course, part of the point of Handmade Hero is to show that you can totally reimplement everything from first principles. Libraries are not magical black boxes, they're code written by human beings like you or me, and you can understand what they're doing.
For instance, he wrote his own PNG decoder[0] live on stream, with hardly any prior knowledge of the spec, even though I'm confident that under normal circumstances he'd just use stb_image. I'm sure he did this just to show how you'd go about doing that sort of thing.
[0] He only implemented the parts necessary to load a non-progressive 24bit color image, but that still involved writing his own DEFLATE implementation.
Jon is not making the by-the-numbers annual entry in the Call of FIFA series here.
But also, as a nit pick, Jon isn't just programming all the time, he's running a business and he is very involved in the indie game community and a founding member of indie fund. And when he is programming it isn't always for his own games. Here is a link to the credits page for him at MobyGames:
https://www.mobygames.com/person/188969/jonathan-blow/credit...
Where in addition to nebulous "Special Thanks" credits, you will note programming and QA credits on several non-Thekla games.
Moreover he's pushing a particular paradigm and view point that is in many ways the opposite of clean code. He pushes for still doing very low level design with minimal abstractions. But even at the time Braid could have been written in python SDL wrappers and probably had similar performance, and the witness could have used unity. If clean code is about maintenance and time to market, the Blow paradigm hasn't proven that its needed or fixes the holes in clean code. This is not to say clean code is perfect just that Blow hasn't cracked the nut either and I don't know why people act like he is the final word, or honestly even a respected voice, in game software engineering. On the other hand, if Blow wanted to talk about managing indie studios or game design my ears would prick up instantly.
I disagree.
> You say he has more credits but only one of those is programming since Braid.
He did start a company after that you know. A successful one that makes money and employs people to make art. I don't imagine that running a business takes no time from his life.
> But even at the time Braid could have been written in python SDL wrappers and probably had similar performance
Braid did a lot more than you give it credit for. Here's a GDC talk about the rewind system in which he explains some of the hurdles he had to deal with: https://www.youtube.com/watch?v=8dinUbg2h70&t=5s Pay particular attention to the discussion of the background particles and how to get that to work within the RAM constraints.
> On the other hand, on a software engineering level those paint by numbers Modern Warfare and FIFA games are both more technically impressive and are designed for fast iteration
Hardly anything about these games change from release to release. They're not exploring new gameplay problem spaces, they're not doing anything super interesting or surprising on a technical level either and I don't get why you think they are. Of course if you keep using essentially the same engine and know exactly what you're trying to make, making another like a goddamned factory is going to happen quicker than if you're trying to make something unique and meaningful.
> If clean code is about maintenance and time to market, the Blow paradigm hasn't proven that its needed or fixes the holes in clean code.
I have heard "maintenance and scaling" as an excuse for poorly performing software for a long time now, yet what I'm not seeing is software that has features added on quick schedules and without bugs. So at best I'd say that it isn't accomplishing what it is supposed to and, at the same time, it is wasting our time and resources by producing slow software to boot.
tell me you've never made a game without telling me you've never made a game
When I was coming up I got gifted a bunch of modules at several jobs because the original writer couldn't be arsed to keep up with the many incremental changes I'd been making. They had a mentality that code was meant to be memorized instead of explored, and I was just beginning to understand code exploration from a writer's perspective. So they were both on the wrong side of history and the wrong side of me. Fuck it, if you want it so much, kid, it's yours now. Good luck.
When we get unstuck on a problem it's usually due to finding a new perspective. Sometimes those come as epiphanies, but while miracles do happen, planning on them leads to disappointment. Sometimes you just have to do the work. Finding new perspectives 'the hard way' involves looking at the problem from different angles, and if explaining it to someone else doesn't work, then often enough just organizing a block of code will help you stumble on that new perspective. And if that also fails, at least the code is in better shape now.
Not long after I figured out how to articulate that, my writer friend figured out the same thing about creative writing, so I took it as a sign I was on the right track.
I do know that the first time I was doing that, it was for performance reasons. I was on a project that was so slow you could see the pixels painting. My first month on that project I was doing optimizations by saying "1 Mississippi" out loud. The second month I used a timer app. I was three months in before I even needed to print(end - start).
TDD provides those guarantees. If someone changes the behaviour of the function you will soon know about it.
That's significant because Robert 'Clean' Martin sells clean code as a solution to some of the problems that TDD creates. If you reject TDD, clean code has no relevance to your codebase. As Casey does not seem to practice TDD, it is not clear why he though clean code would apply to his work?
TDD is about documenting behaviour. Which is why it was later given the name Behaviour Driven Development (BDD), to dispel the myths that it is about testing. It is true that you need to document behaviour before writing code, else how would you know what to write? Even outside of TDD you need to document the behaviour some way before you can know what needs to be written.
A function's behaviour should have no reason to change after its behaviour is documented. You can change the implementation beneath to your hearts content, but the behaviour should be static. If someone attempts to change the behaviour, you are going to know about it. If you are not alerted to this, your infrastructure needs improvement.
That's only true with spherical cows. That something happens is a requirement. When it happens is often only as specific as 'before' or 'after' but tests often dictate that they happen 'between', which is not an actual requirement, it's an accident of implementation. It was 'easy' to put it here.
Nowhere is it written that behavior in a system is strictly additive.
Systems are full of XY problems. When you recognize that, and start addressing that problem, you sprout a lot of tests for the Y solution and block delete tests for the X solution. That behavior doesn't exist in the system anymore because it's answering the wrong question. Functional parity tests can be copied, or written in parallel. But the old tests disappear with the old code (when the feature toggle goes away).
Leaving the code for X around is at best a footgun for new devs, and at worse a sign of hoarding behavior of an intensity that requires therapy.
You're espousing a process whereby you've nailed one foot to the deck, preferring form over function. Whether you believe what you're saying or not I can't say, but it's restrictive and harmful.
For a unit of the same identity to suddenly start doing something different is plain nonsensical, never mind the technical challenges that come with breaking behaviour that should scare anyone away from trying. Logically, a unit is additive until the unit are no longer used, at which point it can be eliminated.
> But the old tests disappear with the old code (when the feature toggle goes away).
Absolutely, but static analysis can easily determine that the tests being removed correspond with units being removed. If (TDD) tests are removed and the unit code isn't, something has gone wrong and your infrastructure should make this known.
Refactors compose. In three months you can completely rearchitect a module without breaking it at any point in the process. That’s the promise of refactoring.
Functions don’t have an identity. There is no such thing. I don’t know who taught you that but they have broken you in the process. Renaming things is a refactoring. We don’t check the entire commit history to make sure that function name has never existed. Only that it hasn’t existed recently. There’s no identity.
One of the reasons to refactor is that the function has been lying about its responsibilities. So you extract steps out of it, create a new call path that fixes the discrepancy, migrate the call sites, delete the incorrect function, and then, if the function name was really good, you might wait a while and rename the new function to the old name. Each step makes sense if you’ve followed the entire process. If you haven’t been following along at all then you have absolutely no idea how things got here until you read the git history thoroughly, which some people can’t do, and others won’t do if they expect the code to be static.
With respect, you should reread the comment chain. You're clearly just repeating what has already been said.
Like visiting a friend who did their own house remodel. Their spouse saw all the steps, all you saw was before and after, and so the fact that the bathroom door is missing is confusing. The bathroom still exists, but now it's the master bath.
IBM’s Visual Age tried to behave that way and it didn’t end well. Eclipse dropped that conceit when it forked.
The only comment that talks about function identity is this one[1], and it is written by you. "His head" refers to your own?
I used the word identity earlier in that thread, but I definitely wasn't referring to functions. The word "function" isn't even found in the comment.
Remember that everyone has their blind spots!
I follow Casey on twitter, and a couple years ago there was a weird thread where he had hung his browser for 4-5 seconds by running some JS to assign CSS rules to ~50K div elements. And Casey was a million percent confident that the hang was due to JS being slow, and had nothing to do with CSS or DOM rendering.
Likewise, I've worked in code bases where performance had been dreadful, yet there were no obvious bottlenecks. Little by little, replacing iterators with loops, objects/closures with enum-backed structs/tables, early exits and so on accumulating to the point where speed ups ranged from 2X to 50X without changing algorithms (outside of fixing basic mistakes like not pre allocating vectors).
Always fun to see these videos. I highly recommend his `Performance Aware Programming` course linked in the description. It's concise and to the point, which is a nice break from his more casual videos/streams which tend to be long-winded/ranty.
There might be some important colums like say a “status” or a “date”, which are fundamental to a lot of queries.
Or you have colums X and Y being used frequently or importantly in where clauses together, then that’s a candidate for composite indexes.
Stuff like that.
There's been a lot more educational material on composite indexes and in particular partial indexes and so I'm not sure if someone with 3 years' experience today can accurately judge a conversation talking about ten years ago.
If there are a few fields referenced in a single WHERE add a single index that includes all of them.
If you have index that has a, b, c then it is as if you also had indexes a, b and a.
If condition in WHERE is = put this field at the beginning of an index. If it's < (or similar) put it at the end. You'll get best results if you have none or only one < in your query.
Just taking the little bit of time to think about what the computer needs to do and making a reasonable effort to not do unnecessary stuff goes a long way. That 2x-50x factor is in fact very familiar. That’s something loading in a second rather than in a minute, or something feeling snappy instead of slightly laggy.
And it matters much more than people say it does. The “premature optimisation...” quote has been grossly misused to a degree that it’s almost comical. It’s not a good excuse for being careless.
It takes 19 seconds for the main menu to load when you push play on our current game in Unity. It's killing me.
Meanwhile in my lua side project, its less than a second.
Can you recommend any SQL book with main focus on performance improvements like this?
> Select * queries
Sometimes you only want two columns, but you ask for 5. Say you query a million rows, where you ask for but throw away 60% of all the data you get back.
> Naive indexes
As in, just slapping an index on a table that doesn't have one makes such a big difference that sometimes it's all you need.
> N+1 queries
This is more of a problem in ORMs, but any time you call the database N times instead of 1 time. A classic example is writing a for loop that asks for one row at a time, instead of asking for all rows once.
You get 90% of the improvement just from that.
Indexes are the whole reason anyone even uses databases. And yet some backend guys think they are optional.
We can all write code that glues a very fixed set of things end to end and squeeze every last CPU cycle of performance out of it, but as we all know, software requirements change, and things like polymorphism allow for much better composition of functionality.
The abstraction is super common and allows you to connect streams to each other without worrying about the underlying mechanism, which 99% of the time I don't really think you want to worry about unless you're sure it's a performance overhead. And that's great, because I surely don't want to write specializations by hand for all the different combinations of streams I need to use if I don't have to.
I use generic readers and writers every day, but they don't have any runtime dispatch. The question was "Why would you do anything except vtables and runtime dispatch?" One answer is that my code that uses generic readers and writers and gets monomorphized gets to also be generic over async-ness.
Also I think he isn't specifically against "clean code" but how the first tool used by various "clean code advocates" seems to be polymorphism via inheritance. I have seen this enough in lots of Java codebases and "Enterprise C++ codebases". "We need X." "Oh first I will create an Abstract Base class for X, then create X, so we can reuse X nicely elsewhere when we need it". It is still on the developers to understand that they may not even need it but for them it is "clean code".
virtual void setPixel(int x, int y, int color) = 0;
and then implemented flood fill etc in terms of that.You can still easily get something working just by implementing setPixel without virtual dispatch, the linker has no problem inlining that call at compile time.
If some arbitrary API needs it to be virtual it's easy to implement the virtual call in just that specific case, instead of burdening your entire system with a virtual call that'll always be static in practice.
Your entire system will not be burdened with virtual calls in places where you use concrete implementations (so long as they're final/sealed, anyway). The overhead is only there if you try to use the abstraction generically, but why would you do that in the case where the virtual call "will always be static in practice"?
It's only a problem when it's done per datum.
Heck, even codebases used in trading can use polymorphism for very high level interfaces (or async tasks) but hot code paths don't use it.
There's lots of ways to do this poorly and well. There's no process for it. That's a feature. I feel like a lot of the flak clean code gets boils down to, "I followed it dogmatically and look what it made me do!" It didn't make you do anything; it's trying to teach you aesthetics, not a process. Internalize the aesthetics and you won't need a rigid process.
Obviously when you do this you probably need more code than you'd normally write. That can be viewed as a maintenance burden in some situations, esp. when you don't have product market fit. Again, this shows that treating clean code like some process that always produces better code in every situation is extremely naive.
The tests he put together here are hardly something I'd call a straw-man argument, they seem like reasonable simplification of real-cases.
The focus on performance here ignores the fact that most programs are large systems of many things that interact with each other. That is where good design and abstractions and “clean code” can really help.
Like all things it is about finding a balance and applying the right techniques to the right parts of a larger system.
To me the interesting point is the reminder that there is an innate tension between going fast and being "clean" (i.e. maintanable/understandable). And once you are aware of it you can make your decisions in an informed way. Too often this tension is forgotten/ignored/dogmatically put to the back ("performance doesn't matter over cleanliness" and the likes).
Mind you, I'm also of the camp that performance is very secondary to cleanliness in modern enterprises, but I appreciate a reminder of just how much we are sacrificing on this altar.
It is not an excellent argument for why 'clean code' is not a better fit for the other 99% of their codebase.
For example, if the code is to be distributed and extended as a third-party library, then the class hierarchy is probably a better fit to allow extensibility.
But if the purpose of the API is to compute the area of given shapes (as in the example), then it makes sense to make it efficient and there is no use to provide extensibility to the outside world.
The advocates of "clean code" that Casey mentions will go for extensibility no matter the use.
I find it very interesting to have numbers to weight what you are leaving when you go the extensibility route instead of the non-pessimistic route.
Unfortunately this is just some catchy stereotyping that probably doesn't match reality.
The process was "think about the model, design a class hierarchy that fits with the model, add the operations". Even if I don't think about extensibility, that's how I thought my code.
Why? Because when I think about performance, I used to think about algorithms and I/O optimizations (for example batching).
Now I'll look into this data-oriented programming-thing and see how that applies to my platform (embedded Java) :D
That's a fairly rare breed for sure, but the real problem was that nobody called him out. Well I did, but I'm no longer working there. I couldn't.
Paraphrasing Russ Ackoff, doing the right thing and doing a thing right is the difference between wisdom or effectiveness and efficiency. What Casey is doing here may be efficient, but calculating a billion rectangles doesn't present a realistic or general use case.
"Clean Code" or any paradigm of the sort aims to make qualitative, not quantitative improvements to code. What you gain isn't performance but clarity when you build large systems, reduction in errors, reduce complexity, and so on. Nobody denies that you can make a program faster by manually squeezing performance out if it. But that isn't the only thing that matters, even if it's something you can easily benchmark.
Looking at a tiny code example tells you very little about the consequences of programming like this. If we program with disregard for the system in favour of performance of one of its parts, what does that mean three years down the line, in a codebase with millions of lines of code?`That's just one question.
You should check if your code is in the hot path before optimizing, because the more you couple things together the harder it is to change it around. For instance, in Casey's example, if you wanted to add a polygon shape but you've optimized calculating area into multiplying height x width by a coefficient, that requires a significant refactor. If you are sure you don't need polygons, that's a perfectly fine optimization. But if you do, you need to start deoptimizing.
Clean code means easy to read, maintain, not full of arbitrary things 'because performance'.
There was a moment in grade school where I was sat down and it was explained to me that you don't have to take a test in order. You can skip around if you want to, and I ran so far with that notion that at 25 I probably should have written a book on how to take tests, while I could still remember most of it.
One of the few other "lightning bolt out of the blue" experiences I can recall was realizing that some code constructs are easier for both the human and the compiler to understand. You can by sympathetic to both instead of compromising. They both have a fundamental problem around how many concepts they can juggle at the exact same time. For any given commit message or PR you can adjust your reasoning for whichever stick the reviewer has up their butt about justifying code changes.
you have Japan infrastructure, and you have Turkey infrastructure
6.1 quake in Japan = nothing destroyed
6.1 quake in Turkey = everything collapses
The engineers in Turkey probably didn't value performance and efficiency
It's the same for developers, you choose your camp wisely, otherwise people will complain at you if they can no longer bear your choice
You act like innocent, but your code choice translate to a cost (higher server bill for your company, higher energy bill for your customers/users, time wasted for everyone, depleting rare materials at a faster rate, growing tech junk)
Selfishness is high in the software industry
We are lucky it's not the same for the HW industry, but it's getting hard for them to hide your incompetence, as more things now run on a battery, and the battery tech is kinda struggling
Good thing is they get to sell more HW since the CPU is "becoming slower" lol
So we now got smartwatches that one need to recharge every damn day
Uh, no. It was corruption. They were standards, that worked, but people didn't do it, plain and simple.
A "blame systems over individuals" version would be that the industry is externalizing bad performance onto users, damaging environment, causing frustration, wasting lives, and occasionally even actually killing people (shitty ER/hospital software comes to mind) - because there's no good feedback mechanism to force software companies to internalize those costs.
I agree, i apologies for the mistake, i can no longer edit the post unfortunately
The reality is that how flexible your interfaces and abstractions are and their design has to be a part of your original design considerations when building something. It's a bad move to just hand wave away performance concerns because you religiously adhere to some design patterns. It's also a bad move to drop down to using intrinsics for everything from the get-go and thinking you know better than the compiler when it's a codepath that isn't even computationally expensive or a bottleneck a priori.
These "clean code" principles should not, and generally are not, ever used at performance critical code, in particular computer graphics. I've never seen anyone seriously try to write computer graphics while "keeping functions small" and "not mixing levels of abstraction". We can go further: you won't be going anywhere in computer graphics by trying to "write pure functions" or "avoiding side effects".
These "clean code principles" are, however, rather useful for large, corporate systems, with loads of business process rules maintained by different teams of people, working for multiple third parties with poor job retaining. You don't need to think about vector performance for processing credit card payments. You don't need to think about input latency for batch processing data warehouse jobs, but you need this types of applications to work reliably. Way more reliably than a videogame or a streaming service.
Right tools for the right jobs, people need to stop trying to hammer everything into the same tools. This is not only a bad practice in software, it's a bad practice in life, the search for an ever elusive silver bullet, a panacea, a miracle. Just drop it and get real.
Not sure about this; in my experience (in a different domain, audio processing) you totally can get away with both of these a lot of the time.
Function inclining works well, so you can write small pure functions in a lot of cases (especially if you accept a function that reads from one buffer and writes to another as pure).
As for avoiding side effects, this is normally more about keeping your state updates small and localised (allowing more parts to be pure), which is often not a problem performance-wise.
IME it's much easier to improve the performance of a piece of code which is easy to reason about and change with some level of confidence that your optimisation will not break things.
I know there's some unavoidable global state in computer graphics, but presumably there is lots of code that doesn't directly touch that.
But I'd point that languages and compilers are built so we can have both.
The problem is some definitions of good really aren't, and it isn't just because they're slow, it's because they make some things "good" at the detriment of others. Uncle Bob's Clean Code is really good at making some portions of the code "good" and "simple", but when put together they are interconnected in ways that are more difficult to understand.
I agree that this is mostly true, but maybe not for beginners in the field
When I was reading "Raytracing in One Weekend" (known as _the_ introductory literature on the topic), I was very surprised to see that the author designed the code so objects extend a `Hittable` class and the critical ray-intersection function `hit` is dynamically dispatched through the `virtual` keyword and thus suffers a huge performance penalty
This is the hottest code path in the program, and a ray-tracer is certainly performance critical, but the author is instructing students/readers to use this "clean code" principle and it drastically slows down the program.
So I agree most computer graphics programmers aren't writing "clean code", but I think a lot of new programmers are being taught them because of introductory literature
In my experience, writing the code with readability, ease of maintenance, and performance all in mind gets you 90% of each of the benefits you’d have gotten focusing on only one of the above. For instance, maybe instead of pretending that an O(n^2) algorithm is any “cleaner” than an O(n log n) algorithm because it was easier for you to write, maybe just use the better algorithm. Or, instead of pretending Python is more readable or easier to develop in than Rust (assuming developers are skilled in both), just write it in Rust. Or, instead of pretending that you had to write raw assembly to eke out the last drop of performance in your function, maybe target the giant mess elsewhere in your application where 80% of the time is spent.
A lot of the “clean” vs “fast” argument is, as I’ve said above, pretending. People on both sides pretend you cannot have both, ever, when in actuality you can have almost all of what is desired in 95% of cases.
Code that is unreadable, tightly coupled, untestable or just messy is much, much harder to work in than code that is readable, loosely coupled, well-tested and clean. This has been proven often and is really a no-brainer. Performance-optimizing is finding the bottleneck, then rewriting that without changing the functional behaviour. For this you need resp. readability (to find bottlenecks you must be able to understand flow and code), ability to rewrite (tightly coupled code cannot be rewritten in isolation) and insurance the behaviour doesn't change (test coverage).
Ergo: a clean archictecture is a requirement to make code more performant in the first place. Even if that architecture is bad for performance in itself, it enables future improvements.
Clean Code is actually a book by Uncle Bob. And Clean Architecture is the name of another book by him.
What Casey is criticizing isn't "good code". He's criticizing Uncle Bob's philosophy.
If you have a large sum of I/O and can see the latency tracked and which parts of the code are problematic, optimize those parts for execution speed.
If you have frequent code changes with an evolving product, and I/O that doesn't raise concerns, then optimize for code cleanliness.
Never reach for a solution before you understand the problem. Once you understand the problem, you won't have to search for a solution; the solution will be right in front of you.
Don't put too much stock in articles or arguments that stress solutions to imaginary problems. They aren't meant to help you. Appreciate any decent take-aways you can, make the most of them, but when it comes to your own implementations, start by understanding your own problems, and not any rules, blog titles, or dogmas you've previously come across.
I actually just had this debate with myself, specifically about shape classes including circles and bezier curves. However, the operation was instead intersections. There was zero performance difference in that case after profiling so I kept the OOP so that the code wasn't full of case statements.
Let's take your example of a program with a lot of I/O. A straightforward way to optimize that is to find a way to reduce the number of I/O operations you do.
And once you do that, the bottleneck shifts. You're spending less time in I/O, both in an absolute sense and relative sense. So you might run into a new non-I/O new bottleneck that was just drowned out in the noise before. So you optimize that ...
And sometimes this goes on for many iteration cycles and you end up with a 100-1000x performance improvement.
Compare this to discussions about FP, new languages like Rust, and so forth. This really demonstrates the primary vogue mindset is increasing complexity and hierarchy to the detriment of all else, and is why the supposed new paradigms of "modern software development" are not really that new but just evolutions of the current paradigms. You really touch what is a culture's sacred cows by that which attracts criticism without any real sincere rebuttal.
It doesn't matter for corporate environments perhaps. It does hell of matter for consumer facing web stuff, both the front end and the backend! 99% is a lot of stuff, so yes it does matter.
And the top reply (right now) is about 99% of the time is spent waiting for user input, and no, that isn't even true. A lot of that "input" is waiting on the network, and the number of requests for any application makes per unit time definitely scales with increasing code complexity.
But anyway, may be that makes a tenuous argument that most code does not care about performance, but again, if it's 99% of code, then yes it matters because it's my entire computer, and that's how we have machines that are faster than they've ever been yet they struggle to edit text compared to say emacs on pentium 4.
I don't think this is a sincere rebuttal at all.
In fact, it perversely proves the point. If 99.9% of apps aren't caring about performance - that's probably why my $2000 computer is slow to open a text file.
I want all of my applications to be well performant. Using the modern web is an atrocious experience most of the time. It's strange when I stumble upon a mostly-unchanged "web 2.0" era site. It loads instantly, like, shockingly fast. It doesn't have all the SPA widgets and "interactivity" but it works, loads extremely fast, and doesn't turn my computer's fans on.
My computer is several orders of magnitude faster than my computer from 2001. Yet the applications I want to use feel slower.
People say that programmer productivity is more important than efficiency, I don't think programmers are more productive. We've just lowered the bar.
And I think neither FP nor Rust discourage the internal-representation-dependent 10x optimization. They only discourage doing this across module boundaries, but the style of programming encouraged by FP and Rust encourages putting your datatype variants together in one module (unlike OOP).
So with those languages you're much more likely to naturally arrive at a fast solution than with traditional OOP.
That is not our job! Our job is to solve business problems within the constraints we are given. No one cares how well it runs on the hardware we're given. They care if it solves the business problem. Look at Bitcoin, it burns hardware time as a proof of work. That solves a business problem.
Some programmers work in industries where performance is key but I'd bet not most.
CPU cycles are much cheaper than developer wages.
And we use them because they’re autonomous, remember exact details and are very fast and reliable.
There’s of course some level of good enough. We don’t write ad-hoc scripts in assembly.
But to say dev time is more expensive than computer time only makes sense if programs are actually fast. Fast, reliable feedback loops matter. Consistency matters. And simplicity matters in many dimensions.
Web application servers that are orders of magnitude (N times) slower than they should be (not even _could_ be) cost us N times more hardware resources, N times more architectural complexity that require specialized workers and tools and so on.
Speed and throughout matter for productivity. Not just ours but our user’s as well. Good performance is important for good UX. Wasting fewer cycles opens up opportunities to do meaningful things.
Among the most popular languages is Python. It is popular in spite of its bad performance, high memory use, and lack of CPU multithreading.
And it is heavily ran on servers.
Why? Because running Python apps is still much cheaper than hiring humans to wait for calls or manage e-mails.
Humans are valuable. They should not be working on easily automatable problems.
The bottleneck is automating AT ALL, rather than automating with a low machine cost. Only at huge scale (i.e. Big Tech with billions of daily events) does it warrant to optimize the code.
Of course, assuming you have a sane computational complexity. If you don't, it doesn't matter which paradigm you use.
But the gist of the video doesn’t disagree at all.
In fact the resulting code was very clear, easy to write and understand. He got a 15x improvement by removing indirection and OO cruft. I don’t think he’s saying “don’t use language X” here, but rather “don’t make it harder for yourself and the computer”.
Hardcoding types and methods (so they compile to simple/fast machine code as the video proposes) takes away flexibility (but you can do that with Cython by the way, but it is not nearly that popular).
The table driven, static dispatch approach Muratori isn’t necessarily “hard coded” in the general case. It’s just data driven, or table driven as he puts it. It’s not less flexible as you can easier add (append!) new cases.
In fact it reminds me of how I do it in higher level languages like JS and Clojure.
Of course we’re paying a substantial performance tax when using higher level, dynamic languages. We’re paying that tax in order to get very tangible benefits.
But we can still program in a way that is data oriented, simple and easy to work with by the compiler and runtime (JIT).
Sure, python web servers are a thing. Heck, we use them a fair bit. But it's a deliberate decision from the start where it makes sense to use it for that purpose.
Exactly that!
There is a time and a place for performant code but it's not the only metric we need to take into consideration. Sometimes you're better off using the slower tool or algorithm as it improves things on a dimension other than performance.
Of course not. But then, nobody is really complaining about apps they think are fast enough. The problem is, when something is noticeably slow, you complain about it, file reports etc, and are met with stiff resistance.
What's important is not that you make the machine go as fast as it can possibly go at all times. What's important is to know how fast it can go, so that you're aware of just exactly how much you're leaving on the table. The actual amount in most cases, would, I expect, surprise most people...
Writing in efficient code has an environmental impacts and a human impact.
I still lean heavily towards clean code, because clean is easier to Grep and bugs have a much higher impact to my career, but efficiency imho should be addressed in the language or compiler, not in the code.
Please just stop with this. It's plainly false.
At $dayjob I recommended some simple database query tuning that 1 developer applied in their spare time. This improved performance from 9 seconds per page to 500 milliseconds per page.
That customer wanted to use auto-scale to expand capacity (nearly 20-fold!) to meet the original requirements, which would have cost about $250K annually.
The dev fixed the issue in like... a week.
What developer costs $250K per week!? None. None do. Not even the top tier at a FAANG.
Not to mention the time saved for the thousands of users that use this web application. Their wasted time costs money too.
Those are the constraints you work with. Those are the business problems you are solving. So yes this was justified. It was justified on the numbers.
My point is that your job is only to tune performance iff there is a solid business case for it.
I spent a week reducing a pages load time because the busines saw the load times as a problem for customer acquisition. The cost of my time was justified. Meanwhile we had a task that for 6 years took over an hour to run. I optimized it to take less than a second. That change has offered zero business value and was a waste of time. It runs periodically on a server that is most often idle. There was no justification for that work, other than it taught me to align my efforts with the business.
CPU cycles are much cheaper than developer hours. But yes, if you have enough of them yes they will cost more than a developer.
I think its very very hard to put a cost on performance.
A few seconds here or there is very draining on people. How do you measure if people are avoiding doing things, or putting off work because their tools are janky. How do you measure how much time people spend complaining about how slow their computer is?
If they are a normal business admin they are very replaceable. So you don't have to make life easy for them unless they can justify the cost of doing so.
How busy are they?
If you have a 1 EFT position filled by an administrator that is only 60% busy then there is no cost in slowing them down until they're at 100% capacity. Then you might think about making life easier for them to avoid having to hire the next administrator. If people avoid doing their job because they don't like their tools that is a disciplinary matter. Time to find an administrator that can do the job required.
That is a mentality which is too horrifiying to find proper words to describe.
> If people avoid doing their job because they don't like their tools that is a disciplinary matter. Time to find an administrator that can do the job required.
It's human nature to avoid difficult things. Discipline only goes so far; it happens subconsciously.
And anyway, small inefficiencies add up to mountains, but it can be hard to see what's going on when everything is just pebbles everywhere.
The actual production platform is just a dozen or so virtual machines, they're not even that big.
Not everybody has their own data centre where they pay cost-price for all-Linux servers that run only free software.
In a typical business that hosts their servers in the public cloud, it's not unusual for a single database index to allow savings of tens of thousands a year.
E.g.: SQL Enterprise on an Azure E8bds_v5 VM is $2,678 per month, but the next step up to E16bds_v5 is $5,355 per month!
If you can optimise a database so it can move from 16 vCPUs to 8 vCPUs then that alone saves $32K annually. This is ignoring ancillary costs such as the second DR instance, etc...
CPU cycles are cheap, SQL Server licenses are extortionately expensive. While the costs are tied to the server they run on if you can offload to a different CPU not tied to that license model you can still take advantage of the low cost of CPU cycles.
But people aren't. Any code that "makes people wait" is wasting people's time. The only way to make code take less time is to optimise it, because capacity != latency. You can't get a 20 THz processor. You can buy more capacity, but you can't buy more speed!
At FAANG scale, it's common to hire "top" developers at hugely expensive annual total comp to tune stdlib code like "string" for just 1-2% efficiency gains because at their scale that might be 1,000 to 10,000 fewer servers.
At a small scale, budgets are tight.
At "enterprise" scale, staffing (user) costs are high, and license costs are high.
I can't think of a typical business scenario where compute and/or associated per-core-licensing costs can be blindly disregarded with a flippant statement like "developers are expensive and infrastructure is cheap".
Say you're a small business with an in house server rack. Not a software company but say a manufacturing business. Not enterprise scale, smaller. You have 1 development resource. It turns out the ETL server is overloaded and it's causing the reports that run on the same server to run slow. You could get the developer to spend a few weeks porting the legacy system to faster modern option to speed up the ETL and maybe improving some of the reports. But it would be far cheaper to buy in another server for $3k and have the developer spend less than half a day moving the ETL onto the fresh server.
On which... it'll run at maybe 20% faster, because that's the scale of single-threaded processor speed improvements these days. Not to mention that now there's a network hop involved, which will eat into any CPU gains.
Very few apps scale well with increasing core counts, and then hit a wall around 64 cores for almost everything.
Okay, okay, fine. The ETL is natively parallel code and somehow, magically, it can read inherently sequential file formats like multi-gigabyte CSV or JSON files in parallel. This tiny org already has 10 gigabit switches, SFPs, and everything.
Did you upgrade the database server too? No? Now the shiny new ETL server is twiddling its thumbs while the database server is getting overloaded.
Suddenly this option is "not so cheap". You have to buy a new database server and... uh-oh... it's Microsoft SQL Server, Oracle, or SAP HANA, and the licensing is going to eat half your tiny little company's profits for the year.
Did you forget the OS license, backup agent license, anti-malware license, and so on? I bet you did. All of those are extra, and either per-machine or per-core.
Someone has to set this all up. Small non-IT shops typically outsource this to an IT service company. They'll explain all the extras that turn a $3K purchase into a $30K purchase (including on-site assistance to install everything).
Or that 1 guy could have just taken a 10 minute look at the ETL logs, discovered that "SELECT * FROM HugeTable" is unnecessary, and fixed the problem.
Imagine it's all Linux no back up required and the database server has lots of capacity.
The lone developer can do it in half a day. No extra cost.
But yes there exist situations where that is no possible. This is just one example of where it is
Yes, sometimes this might make sense --- if the dev fees are going to be exorbitant, or if you just can't afford to pull a dev off a project to work on it. Other times, it makes more sense to pay the dev...
You keep saying, CPU cycles are cheaper than developer hours, but this is nonsensical without quantities attached to each. How many CPU cycles, on what kind of machines? How many cycles do those machines have to spare? What's the performance per watt? How many developer hours at what kind of salary? There's way too much missing info to be making such a statement.
Yeah that is fair. There is a lot of missing information. I'm only trying to across that now days developer time is often expensive and hardware is often cheap and powerful.
I don't like the idea that our goal is always to make the most efficient code possible. It's not, it's to deliver business solution as efficiently as we can with the resources available. Just like you wouldn't want to pay for a mechanic to spend a week making your car more fuel efficient if you took it in to get new engine mounts. His job is not to make your car work for you, not work the best it possibly could.
That said most of the time you do want to be writing efficient code.
If he could get me a 1000x gain in fuel efficiency, like you often see when software when performance overhauls are done, I would sure as hell give him his week. But this is less about maintenance and more about how the car/software is built the first time around.
In that vein, I do expect that if gasoline prices drop to $0.20c/gallon in a decade (hah), that fuel economy on new cars does not drop to 3mpg to match. That's essentially what seems to have happened in software --- the hardware got really fast, so software got really slow.
> It's not, it's to deliver business solution as efficiently as we can with the resources available.
This is true; I guess what I take issue with is externalizing hidden costs to the customer. We keep paying for faster and faster hardware, and have to because that old hardware which is still working perfectly fine can't run the new software, which is much slower. And often, even on new hardware, the software is just this side of "tolerable". If you're writing your own in-house tool and nobody cares, do whatever suits.
> That said most of the time you do want to be writing efficient code.
Yes. That's all I want. Reasonably efficient. Not balls-to-the-wall speed demon witchery like we saw back in the demoscene heyday, just not to be sliding backwards all the time to erase all the gains our hardware got us.
---
EDIT: I get where you're coming from with the business incentives, I really do. But I'm saying I have a lot of issues with the end result --- as is often the case, maximal profit for the business is wreaking havoc elsewhere, in a sort of tragedy of the commons effect. And there are no realistic ways for me as a consumer to alter business incentives. A lot of software, especially the type that people get paid to write is closed source and closed protocol.
An excellent example is Discord --- it works well enough, when it works, but it's kind of a big heavy behemoth. It doesn't run well on older computers (I frequently see it burning a whole core just sitting in a voice channel). Right now, it's using nearly 1GB(!) of RAM, and frequently climbs to 2 or more if I leave it run long enough. This is a program whose core functionality was essentially available to me in 1999, and it struggles if I try to run it on a 4GHz machine from 2012. The search function sucks imo, and various other complaints. Screensharing with audio is broken (because electron), and probably always will be.
And I can't do a damn thing about it. I can't use a different platform, because the platform I use is determined by the people I want to talk to that are already using it. I can't improve or fork the program, because it's closed source. I can't (realistically) use a different client, or even write my own, because it's a closed protocol. So I'm just stuck with this pig of a program, and no amount of rage or frustration that I feel will alter the company's business incentives.
But it's just one program, right? Okay, I've got the cycles to spare, and the RAM, well, I overprovisioned this machine, so (in the case of this one, relatively modern machine), it's not the end of the world... right?
Now add Spotify. It's the same damn problem, so now the problem is 2x. Add a web browser (I mean, one who's job is actually to browse the web). I manage to draw the line there, mostly, but a lot of people are stuck with a lot more (VS Code, etc). It all adds up to a nightmare. And yet, all the time, I hear how "performance doesn't matter" (not exactly your words, but a prevalent developer sentiment).
you can only excuse-kick the blame-can down the time-road so far.
CPU cycles are cheap.
Dev time is expensive.
User time is sacred.
Why? What if the user is a machine operator? He is standing around waiting for the machine to finish it's current cycle. As long as he can get his data entry done in the time he is waiting it costs the business nothing.
But we don't realise it because we have no idea how harmful it really is to waste a few seconds a day for millions of users.
For example: The finance industry heavily uses Excel the inefficiencies this brings to performance are huge. However, the ability to get a domain expert to maintain the 'code' is worth it.
There are always trade offs. Even if you are in a performance sensitive industry your job is still not to write the most efficient code possible. You are there to make the best use of the resources you have to produce the best product you can. Sometimes that will mean doing things that harm performance but help in other ways.
> That is not our job! Our job is to solve business problems within the constraints we are given. No one cares how well it runs on the hardware we're given. They care if it solves the business problem. Look at Bitcoin, it burns hardware time as a proof of work. That solves a business problem.
Cost is a business problem. It is a constraint. The problem is sometimes it doesn't feel that way because a lot of folks around you use bad practices and it's easy to horizontally scale things so performance is often brushed off.
Now that the era of 0% interest has ended, companies are actually starting to take a look at what they're running and... surprise! Taking even 10% longer to write quality, performant code can yield 2-3x (sometimes I've seen 1000x) improvements which means much lower cost.
> Some programmers work in industries where performance is key but I'd bet not most.
Because most don't know any better. Unfortunately this is a symptom of bootcamps, etc. that have created a lot of "programmers" that only know how to code vs. "engineers" that know how to build systems that are maintainable, scale, and solve complex problems.
In almost every company I've been at they inadvertently retroactively look back and realize how wasteful they were. In the ML world at least, it's a bit different I suppose because performance matters from the start.
I 100% agree. What I'm trying to get across is that you have to identify the cost to justify the performance tuning. If the cost of the tuning is greater than the cost incurred by not doing you might not want to do it.
What are your actual costs and how do they line up? Are you writing ML and big data then processing costs are huge. You can probably win by spending money on developer time to reduce processing costs. On the opposite end of the scale are you writing CRUD for a small business, then the development costs are likely to outweigh any costs from inefficiencies in the applications code.
To me the article read as if our job was to make every bit of code as fast as possible. I think we should only spend time on code that meets the greater goals of the system it operates in.
If you identify that run time is a cost then you have to identify what part of the code is the bottle neck and fix that. Then you have look back to see if any more improvements will be worth the time invested.
My interpretation was that he thinks the baseline for performance is too low. There are kind of two parts to performance: non-pessimization and optimization. He’s mostly harping on the former point, that you don’t have to go tuning anything. Just by writing code without following the clean code rules you can get a huge speed up for free, without any tuning at all. It’ll be free of obvious overheads, even if it doesn’t do anything special to fully utilize the hardware.
When every function in your program is virtual, fixing bottlenecks won't give you a significant speedup. You don't need to optimize programs to make them 1000x faster, we just need to stop shooting ourselves in the foot.
> Look at Bitcoin, it burns hardware time as a proof of work
Sounds to me like the job needs other constraints in form of regulations then.
But using "clean code" and SOLID does not make the code faster to write, easier to maintain. It makes it harder to reason about and slower.
The author seems to be neglecting the fact that the whole point of “clean code” is to improve the likelihood of achieving the first goal (code that runs well, i.e. correctly) across months/years of changing requirements and new maintainers. No one (that I’ve ever spoken to or worked with, at least) is under any illusions that you can almost always trade off maintainability for performance.
Admittedly, I think a lot of prescriptions that get made in service of “clean code” are silly or counterproductive, and people who obsess over it as an end unto itself can sometimes be tedious and annoying, but this article is written in such incredible bad faith that its impossible to take seriously.
Yes that is the whole point of "clean code". Thing, is, it failed.
Simplicity is better achieved with other methods. Forget Uncle Bob and SOLID, read John Ousterhout (A Philosophy of Software Design) instead.
That is a large statement you make there. It begs for backing up.
---
As for SOLID, there's one good thing: Barbara Liskov. Her principle have mathematical underpinnings in type theory, shapes Haskell's type classes and likely Rust traits too. The rest however ranges from situational to just crap.
Single responsibility is at best a heuristic for the real goal: keeping a nice and small API/implementation ratio. And it fails way too often, causing you to make tiny classes and one liner functions, whose implementations are so tiny they don't even pay for their interface. Pretty bad overall.
The Open/Close principle is just crap. Don't use inheritance if you can help it, and don't bother with keeping your code open for this or closed for that. Just keep it simple, so that when requirements changes you can rewrite the parts you need to rewrite.
Interface Segregation is a situational heuristic. Just keep your interfaces small, and you'll know when it makes sense to split an API in two or not.
Finally Dependency Inversion is cancer. I mean that literally: it causes your code to grow unsightly appendages, makes everything it touch a tad bigger and more complex, and in most cases it doesn't even facilitates testing. Because surprise, the overwhelming majority of the time, code dependencies are fixed. So let them be. Don't complicate your program with interfaces that only have a single implementation. Let your code depend on the implementations directly. It will be simpler, easier to navigate, easier to modify, and just as easy to test.
---
As I said, "Clean Code" failed. Miserably.
That's not "failing" that's at most "it doesn't work for me", but without explaining why it doesn't work for you, it really is little more than a wild claim.
I've had great success with "Clean code" practices. Though I implement it more according to Alistair Cockburns' "Hexagonal Architecture", it's overall very similar and strives for the same goals using the same methods. So in that sense, N=1 it hasn't failed. It has helped at least one person.
Can't do much better in short HN comments. I've tried at times to be more rigorous than that, but so far I've only scratched the surface: https://loup-vaillant.fr/articles/good-code
That being said, A Philosophy of Software Design looks like a good read. Thanks for the recommendation.
They look like reasonable heuristics indeed, but they're perfectible.
Factoring out repeated patterns, that's good. Keep that one. But I would advise to wait for the pattern to emerge in the first place, so that when you factor it you know exactly what abstraction you need. https://caseymuratori.com/blog_0015 (Semantic Compression)
The size of functions is a fairly poor heuristic. The main size you want to minimise is that of the entire code base. At the function (and class/module) level, it's better to minimise API/implementation ratios. That's how you know the function (or class, or module) is useful: small easy to learn APIs that hide significant implementations give you leverage, and help you minimise what you need to keep in mind whenever you're writing a new piece of code.
Relying on a language's abstraction features… yeah, I guess, though I generally avoid class based polymorphism. It's a bit heavy for my taste, especially when we have the ability to just pass closures around instead. Often that poor man's object is all you need.
> I don't think their critique is really focused on the nuanced differences between the various schools of "how to make your code nicer to read" (at least, I didn't read it that way).
It wasn't indeed, but you'll note his code ended up being quite a bit smaller than the original. All those abstractions are nice when they make your code shorter, or better organised but in this case they just didn't. We could blame the toy nature of the example, but still: all those one liners were terrible for the API/implementation ratio, it's no surprise they could be fused together so concisely.
The more the years pile up the more I agree with the sentiment in this post, generally going for something that works and is as optimal of code as I would get if I was to come back to make it more performance oriented in the future, I end up with something generally as simple as I can get. Languages and libs are generally abstract enough in most cases and any extra design patterning and abstracting is generally going to bite you in the ass more than it's going to save you from business folks coming in with unknown features.
I suppose, write code that is conscious of its memory and CPU footprint and avoid trying to guess what features may or may not reuse what parts from your existing routines, and try even harder to avoid writing abstractions that are based on your guesses of the future.
I agree, but surely you wouldn't really say that his end-state in this particular article is more maintainable than the starting point.
On the other hand, an abstract base class broadcasts the message that we should fold our new cases into the existing design -- that's what base classes are made for. But even when new cases don't fit neatly in the existing design, we often feel as programmers that we need to pay respect and deference to existing design, especially if it was made with extensibility in mind. And so we add more complexity (maybe we need an extra field or an extra method), make more kludges, and soon enough the original OO design is a mountain of complexity and has so much "gravity" that it's nearly impossible to escape it anymore -- nobody can imagine throwing it out and starting fresh, so it just keeps gaining complexity. And for all its complexity, it's also slower.
Having said that, even Java now encourages programmers to use algebraic data types when "programming in the small", and OOP/encapsulation at module boundaries: https://www.infoq.com/articles/data-oriented-programming-jav... though not for performance reasons. My point being is that the "best practice" recommendations for mainstream language does change.
Last time I checked it could not inline megamorphic call sites, evn if implementations were trivial (returning constants). At the same time I saw C++ compilers able to replace an analogue switch that dispatched to constants with a simple array lookup, with no branching at all.
I think OO centric rules are harmful in a word where languages support functional programming etc. Polymorphism isn’t always the best answer. A big nested if statement that reads like the business spec can be easier to follow and reason about.
That aside if easy to understand code makes your app a bit slower, you profile to work out why and fix up the bits that matter making the tradeoff where it is needed.
Writing code for what you think will be performant everywhere, and not caring about readability in the process is a fools errand, at least in most SaaS/Web/Business apps.
Cramping code may be fine for high performance code where the mandays are justified for the two line code fix. But in most Software you likely just want to include the new cache methods, change the used object, add a button or fix an update. And the person fixing it probabaly hasn't seen the specific code ever before. There is sadly no need for high performance software, when you can't sell it in time.
As Uncle Bob says (paraphrasing): Dependency injection (java) is just programming XML where you turned compiler errors into runtime errors
In the end, it is all tradeoffs. If you have a rough mental model of how code is going to perform, you can make better decisions. Of course, part of this is determining whether it matters for the specific piece of code under consideration. Often it does not.
The problem with laying down a bunch of arbitrary rules is that they never apply to all scenarios. As the person coming up with the rules you can easily re-evaluate where and when they don't work, but the novice receiving those rules wont necessarily have an intuition for the reasoning behind them yet, and so wont so considerately apply them. For everyone else, they need to understand that there is no silver bullet, no 10 commandments that will give them the best result, life is messy, and they need to think, develop their own intuitions by interrogating their own code in each new context - but it all starts with caring about your code, not being satisfied with a pile of spaghetti, or a pile of OOP just because OOP, or a pile of strictly pure functions just because FP. Every single rule or programming pattern is wrong given enough contexts, it's all subjective.
Discussing patterns and rules is useful, but only if they are only used as a mental anchor to think about them, not some kind of axioms of programming correctness.
I actually believe "the hardware that we are given" is the entire root of the problem.
Most programmers work and test using whatever hardware is current at the time, but this is makes them blind to possible performance issues.
Take whatever you're working on, and run it on the hardware of 5-10 years ago. If you still have a good experience, you're doing it right. If not, you should probably stop upgrading developer machines for a while.
Whatever your minimum hardware requirements are should determine your development machines. This way, you will naturally ensure your low-end customers have a good experience while your high-end customers will have an even better experience.
My game studio has been doing this for years. It saves money for expensive hardware, it prevents performance issues before they arise and it saves developer time for not having to overthink optimization.
re: the speedup from moving from subclassing to enums - Compiler isn't pulling its weight if it can't devirtualize in such a simple program.
re: the speedup from replacing the enum switch with a lookup table and common subexpression - Compiler isn't pulling its weight if it can't notice common subexpressions.
So both the premise and the results seem unconvincing to me.
Of course, he is the one with numbers and I just have an untested hypothesis, so don't believe me.
What compiler are you using that devirtualizes every class hierarchy? I suspect that Casey is using C++ so he may (unfortunately) have multiple translation units in his program.
Yes, the video uses C++.
Obviously if one compiles a library then the compiler has no way of knowing that other subclasses of `shape_base` do not exist. My point is that when compiling a binary as they are doing for their video, the compiler knows that there are no other subclasses that it needs to cater to.
It might require LTO explicitly, of course. At the very least godbolt doesn't devirtualize without LTO [1], but godbolt itself breaks if I enable LTO [2] and I CBA to test locally right now.
The fundamental problem here is that C++ doesn't have a good abstraction to represent visibility of public types, since any other translation unit - even across the DLL boundary! - can re-declare the type and then derive from it. The only way to constrain visibility is to use anonymous namespaces, and that only works if the type can be confined to a single unit (that C++ compilers seem to ignore the optimization opportunities here in practice perhaps indicates just how rare this actually is).
Refactorings are moves you can make. Choosing when to make them is up to you. In fact, Fowler provides guidance along with each refactoring suggesting when it might be applicable (i.e., not always)
CSE doesn’t work across function boundaries unless those functions are inlined, which won’t happen with virtual functions due to the above.
For one: Most compilers for most languages are bad by that metric, and interpreters don't even get to play. So this is not helpful for the vast majority of people. Waiting around for them becoming good is not a viable option.
Second, say the compiler would perform good in this scenario. Cool, lets go up a notch, or two, and it would start performing bad again, because there are limits to what it can do in a reasonable amount of time.
And if that limit were big that maybe wouldn't matter, but the limit is low, and so it does. Real programs are so much more complex than this example, that even if the compiler got 10 times better, it would still fail to optimize large parts of your real program.
There's so much emphasis on writing "clean" code (rightly so) that it's nice to hear an opposing viewpoint. I think it's a good reminder to not be dogmatic and that there are many ways to solve a problem, each with their own pros/cons. It's our job to find the best way.
"Clean Code" should be called pessimistic coding - a big part of it is to enable OK flexibility in any imaginable direction. But real-life code will not change in all possible direction - in fact, you can predict quite well roughly what can and cannot change. Writing for performance means, among other things, making things easier both for human and the CPU by reducing flexibility in the unlikely directions.
In the toy example from the video: Casey's proposed alternatives baked in the assumption that the program is working with shapes which, for a given computation, all can fit a specific family of equations. Clean code will make it just as easy to add a square as to add a parametric spline surface. Casey's code will make the former trivial, the latter hard without redoing the entire shape-related code. It's a good tradeoff if you're making a program that mostly works with non-parametric simple polygons, because nobody will need parametric splines in it. On the off chance they will, they can pay for the extra effort - and in the meantime, your software is 20x faster than the equivalent "clean code" version.
--
[0] - This thinking alone is a problem. It's not the 1% that needs some optimization work. The entire user-interacting surface and everything downstream of it need it, which means effectively the entire program. You're free to set a cut-off point beyond which you don't care about "less important" features - but 1% seems quite too early.
Where is this in the book? I can't find a single mention of "hotspot" in the book, and even "optimizations" only shows up 3 times.
Ultimately though, data driven design can fit under OOP (object orient programming) as well, since it's pretty much lightweight, memory conforming structs being consumed by service classes instead of polymorphing everything into a massive cascade of inherited classes.
The article makes a good argument against traditional 1980-90's era object oriented programming concepts where everything is a bloated, monolith class with endless inheritances, but that pattern isn't extremely common in most systems I've used recently. Which, to me, makes this feel a lot like a straw man argument, you're arguing against an incredibly out-dated "clean code" paradigm that isn't popular or common with experienced OOP developers.
One only really has to look at Unity's data driven pipelines, Unreal's rendering services, and various other game engine examples that show clean code OOP can and does live alongside performant data-driven services in not only C++ but also C#.
Hell, I'm even doing it in Typescript using consolidated, optimized services to cache expensive web requests across huge relational data. The only classes that exist are for data models and request contexts, the rest is services processing streams of data in and out of db/caches.
If there is one take-away that this article validated for me though, it's that data-driven design trumps most other patterns when performance is key.
Re-usability, OOP concepts or pure functional style, design patterns, TDD or XP methodologies are the only things that matter... And if you use them you will write "clean code". Even worse, the more concepts and abstractions you apply to your code the better programmer you are!
If you look at the history of programming and classic texts like "the art of programming", "sicp", "the elements of programming"... The concept of "beautiful code" appears a lot. This is an idea that has always existed in our culture. The main difference with the "clean code" cult is that "beautiful code" also used to mean fast and efficient code, efficient algorithms, low memory footprint... On top of the "clean code" concepts of easy to test and re-usable code, modularity... etc
A code base that can't be understood and maintained by the whole team, will degrade quickly.
What Casey is doing is showing how bad a hammer is at removing screws. I mean duh, you're removing screws with a hammer.
My comment was a jest (as suggested by the smiley).
What I take from Casey's post is that a simple non-pessimistic representation allows for efficient code. That is, using a table instead of a class hierarchy gives massive performance boost. Compared to a "clever" loop unrolling doesn't give that much of a boost.
So we need simpler representation. IMHO the table implementation is not less readable nor less flexible than the class hierarchy. But it is less common in the code I am used to, in other words, it's not a widely used pattern).
That's the insight, I think. "Clean Code" tells you to use maximally pessimistic representation for everything, because everything could be extended in some way in every direction. Meanwhile, in the real world, you likely have a good idea what directions of evolution are possible, and which of them are even useful.
Casey's example shows you that, if you design your code to make use of those assumptions, you'll get absurd performance benefits for little to none loss in readability (and perhaps even a gain!).
Some may ask, "what if you're wrong with your assumptions?". Well, you pay a price then. Worst case, you may need to rip out a module, rethink the theory behind it, and rewrite it from scratch - likely forgoing some of the performance benefits, too. Usually, the price will be much smaller. Either way, it's still better than being maximally pessimistic from the start, and writing software that never had a chance of ever becoming good or fast.
Techniques like manual code unrolling/inlining, writing branchless code, compressing data to fit into a pointer, etc.
As an extreme example, people choose not to write low level code for a good reason, even if it might be fast.
High level / interpreted code will be more understandable, but will automatically have an overhead in most cases.
You can have fast and clean code, just not the Uncle Bob style of "clean code". Uncle Bob hijacked the meaning of cleanliness. It doesn't mean that code written like that is actually clean, in fact it's usually the opposite: Uncle Bob's clean code is NOT clean.
> the more concepts and abstractions you apply to your code the better programmer you are!
These contradict each other. XP very explicitly opposes introducing (unnecessary) abstractions: YAGNI, DTSTTCPW, etc. And TDD is a good tool for enforcing that, as you only get to write code that you have a failing test case for.
TDD encourages the use of mocks and unit testing to increase code coverage. And unit testing is specially dangerous. You write a test, then program, so the test is helping you (the programmer). Selling the idea that the higher the test code coverage is the better and safer your code is. Not true at all. If your code doesn't have integration tests for example, you will never know how it actually runs. If you mock everything, you are not really "testing" anything but your internal logic. Unit testing and code coverage just checks that a code path has been run. But there are other tools like fuzzy testing or mutation testing... Do you randomize the memory at every test run? Do you make sure that the CPU cache is cold or hot depending on the test? Good testing is hard.
Most unit tests are written to ease the development. After finishing the development, they are safe to delete. Because they don't add any real value as I understand it. I understand that a test is a business contract of something that MUST work in a certain way. Unless the contract changes, the test must never be removed or changed. TDD and exhaustive unit testing make the maintenance process harder because you don't know if a test is useful or not.
If you follow TDD, most unit tests are re-written all the time. Because they were not written to test a business or critical contract, they were originally written to help some programmer write some internal logic.
That's the purpose of unit tests. They do not exclude the need to perform other kinds of test. Integration tests, contract tests, stress tests - all those will focus on different facets of a system.
> Most unit tests are written to ease the development. After finishing the development, they are safe to delete.
This is especially bad advice, unless no one will never touch that codebase ever again.
I saw old unit tests highlight bugs that would have been introduced by new code many times over the years.
> TDD and exhaustive unit testing make the maintenance process harder because you don't know if a test is useful or not.
Then, as a developer, remove unit tests that became useless.
Code coverage is a measurement. If you turn it into a goal, it will become useless. If you have "useless" unit tests, it tells me that some unit tests were written as padding to move code coverage up.
No, it encourages reasonable decoupling, i.e. good design.
If you see yourself introducing mocks (I think you mean stubs, mocks are something more specific) everywhere, you are feeling the pressure, but avoiding the good design.
https://blog.metaobject.com/2014/05/why-i-don-mock.html
> Most unit tests are written to ease the development.
Yes, unit tests help significantly in development.
> After finishing the development, they are safe to delete.
Noooooooooooooooooooooooooooooooooooooooooooooooo!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!
They are your guardrails against regressions.
> If you follow TDD, most unit tests are re-written all the time.
Nope.
And of course that's important because it enables you to be courageous and refactor mercilessly. Which again is important because it enables you to Do the Simplest Thing That Could Possibly Work, and stick to YAGNI, because you know you can change your mind later.
TDD was later given the name BDD (Behaviour Driven Development) to emphasize that you are not testing, but actually documenting behaviour. What kind of behaviour are you documenting that isn't useful and why did you find it necessary to document in the first place?
> If you follow TDD, most unit tests are re-written all the time.
What for? If changing requirements see that your behaviour has changed to the extent that that your unit does something completely different, it's something brand new and should be treated as such. Barring exceptional circumstances, public interfaces should be considered stable for their entire lifetime and, at most, deprecated if they no longer serve a purpose.
The implementation beneath the interface may change over time, but TDD is explicit that you should not test implementation – it is not about testing – only that you should document the expected behaviour of any implementation that may carry out your desired behaviour.
I don't think someone that believes this has a good understanding of unit testing. You can easily get 100% coverage without testing anything at all!
Coverage is a great metric if it's predicated on high quality tests. Even then 100% coverage doesn't equal "safe". It means that a lot of effort has been put into understanding and testing internal behavior.
You still need higher order tests, arguably even more.
Well written unit tests serve to provide regression testing to prevent bugs reoccurring and to keep existing functionality (e.g. support for reading existing data) working.
TDD is used as a way to help write tests for the API surface and usage of your classes, functions, etc.. You should generally avoid testing internal state as that can change.
For example, if you are writing a set class, the logical place to start is with an empty set -- that's because it is easy to define the empty logic, defining accessor functions/properties like isEmpty, size, and contains. The next logical step is adding elements (two tests: add a single element, add multiple elements). Etc.
Later on, you can change the internal logic of the set from e.g. an array to a hash map. You will keep your existing tests as they document and test your API contract and external semantics. Likewise, if your hash set uses another class like an array, or a custom structure like a red-black tree, you shouldn't mock that class.
I had an Android programmer, who was eager to write clean code following GOF patterns, OP and the rest of the fancy things senior developers usually do. Ended up Android team with 3 devs required 3x time to develop same feature compared to single iOS engineer.
You said you're targeting Android, that implies you were using Java, which has its reputation for both a rigidly inflexible language-design team and its ecosystem having more design-patterns than a set of fabrics swatches - that's not a coincidence.
But for iOS, they'd be using Swift, right? Swift's designers clearly decided they didn't want to be like Java: take the best bits of C# and other well-designed languages and don't be afraid to iterate on the design, even if it means introducing breaking-changes - but the result is a highly-expressive language that, as you've demonstrated, allows just 1 Swift person to do the equivalent of 3 Java people. Swift is an actual pleasure to use, but using Java today makes me weary.
(To be clear: Java was a fantastic language when it was introduced, but it simply hasn't kept-up with the times to its own detriment, it feels like its falling behind more-and-more at time goes on - but that's going to be the fate of every programming language eventually, imo).
And perhaps I was lucky, but typically readability and maintainability are orthogonal to performance and efficiency. Sometimes readable will have optimal performance, sometimes not. Then it becomes a matter of tradeoffs.
The first entry even stated that they were guidelines and if you have a valid reason to deviate: discuss it and you'll get an exception.
The discussion thing was mainly for new coders. We had a library part of the code that was used by many programs. Optimizing parts of it for their use case, could make it unusable for the others who used it.
Communication is key. If you discuss, before implementing it, why you're making certain design decisions then everything goes a lot smoother. If there are objections, keep in mind that the worst case isn't throwing away your design and starting over. The worst case is implementing it and screwing over your fellow coders.
That was the case during the 90s and the first decade of the 2000. Just wait for Moore's Law to kick in and in 18 months your code will get faster by an order of magnitude for free.
Single-threaded performance has long-since effectively plateaued: it's 2023 now and a desktop computer built 10 years ago (2013) can run Windows 11 just fine (ignoring the TPM thing) - but compare that to using a computer from 2003 in 2013 (where it'd run, but poorly), or a computer from 1993 in 2003 (which simply wouldn't work at all).
This is not to say that there won't be any significant performance gains to come, such as with rethinks in hardware (e.g. adding actual RAM into a desktop CPU package, non-volatile memory, etc) but I struggle to see how typical x86 MOV,CMP,JMP instructions could be executed sequentially any faster than they are right now.
https://stoneridgetechnology.com/company/blog/the-exponentia...
The rest of the article is concerned with how software today is still written for those "serial" processors in-mind and fails to take advantage of parallel computing hardware - but this is hardly a new nor controversial statement.
My argument was simple: 20 years ago nobody cared about performance because of the strong correlation between Moore Law and MFLOPS and performance in general.
Nowadays with multicore, individual CPU cores are becoming slower, not faster. Then we can't just wait 18 months for individual cores to get faster.
Today we need to write better software.
No, that's not what I said: I'm saying that single-threaded performance is no-longer directly (let alone linearly) correlated with transistor count, and hasn't been for decades - but it's single-threaded performance that matters for most end-user applications on peoples' computers/smartphones/etc.
And Moore's Law makes no claims about performance increases either, only transistor density. It's literally the first paragraph of the Wikipedia article:
> Moore's law is the observation that the number of transistors in a dense integrated circuit (IC) doubles about every two years. Moore's law is an observation and projection of a historical trend.
The link to performance (any kind of performance: parallel or serial) was made incorrectly by another Intel executive, David House, but that assumption simply isn't true.
To summarize: *yes*: bleeding-edge IC transistor density generally doubles every 18-24 months, but this does not translate into any kind of performance-doubling in end-user applications. In fact, on the contrary, perceived "performance" (however you define it) outside of parallelizable programs, has demonstrably stagnated.
I vaguely smell the language of choice has something to do with that, and C++ would benefit more from data-oriented design than from literate OOP or functional programming patterns took from very very different programming ethoses (the programming language is an interface to something else)
Which is all well and good, until you need to hire all the network engineers, systems administrators, devops people, security staff, datacenter operations managers, database sharding engineers, etc to manage the 10x more hardware and network surface area you have to throw at your slow codebase.
I find that viewpoint concerning because the reality is this isn't really a dichotomy. Code can be both performant and clean (note clean is not the same as elegant).
One thing I think is confusing people is the dogmaticism about what's "idiomatic" is especially bad in OOP-heavy languages. This is especially bad in Java where a fetishization of design patterns have led to codebases which are both ugly and unperformant.
The reality is software design needs to consider performance from the get-go. Sure there is such a thing as "premature optimization" but if you've determined performance is a goal then you should follow best practices for high performance from the get-go. That includes not trying to perform math on iterables of objects (since that prevents vectorizatio since the data isn't contiguous), avoiding accumulations, not creating and destroying tons of objects, etc. This can all be done in a clean way! And low level code can be cleanly encapsulated so that other interfaces remain idiomatic and simple.
A lot of people fret that this approach leads to "leaky abstractions" because implementation details inform the interface design. That just means you need to iterate on the interface design so it makes sense.
Now, if we consider only a conservative 2x speed-up, I might not care if my app starts up in 2s or 4s, but I do care if my device's battery last for 20h v 10h.
Can't say it better than Knuth:
We should forget about small efficiencies, say about 97% of the time: premature optimization is the root of all evil. Yet we should not pass up our opportunities in that critical 3%.
Clean code has never been about performance, it's about the other people that will read your code – future you included. Performance of people is more valuable for any product than code performance[1]. Only when the performance becomes a bottleneck, or you want to optimize energy efficiency then sure, don't pass that 3% opportunity.[1] I would argue that it still holds for products that need high(er) performance code like video games, embedded systems, particle physics, ... These products just happen to hit bottlenecks way faster and some have hard cutoffs (eg. 60 fps for a game). Still, not everything needs to be optimized to the extreme: the algorithm to sort an in-game inventory does not need to handle 4B+ items.
You shouldn't take that quote to mean you can entirely ignore performance 97% of the time.
> The improvement in speed from Example 2 to Example 2a is only about 12%, and many people would pronounce that insignificant. The conventional wisdom shared by many of today's software engineers calls for ignoring efficiency in the small; but I believe this is simply an overreaction to the abuses they see being practiced by penny-wise-and-pound-foolish programmers, who can't debug or maintain their "optimized" programs. In established engineering disciplines a 12% improvement, easily obtained, is never considered marginal; and I believe the same viewpoint should prevail in software engineering. Of course I wouldn't bother making such optimizations on a one-shot job, but when it's a question of preparing quality programs, I don't want to restrict myself to tools that deny me such efficiencies.
That paper was published in 1974 and yet it captures the mindset of many a programmer in 2023 perfectly. The part I like about this paragraph is the "easily obtained" sentence; I saw a comment from someone mentioning that they made a JavaScript program 10 times faster by replacing the common functional programming combinators (map, filter, reduce, etc.) with for loops. I think most of us would say such a change is easily obtained, does not make the code impossible to debug or maintain, and gives such a massive improvement that it should be a no brainer to reach for it.
That's why you need to understand bottlenecks in your system before going around and start optimizing things, as it could be pointless or even counter productive.
Like, when you're choosing a big data structure that is going to have lots of by-value lookups, you don't implement it as an array with O(n) lookup first and only move to a hashtable after benchmarking it. That would be absurd. You just use a hashtable with O(1) lookup right at the start. Because in the overwhelming majority of cases that's the right thing to do and it doesn't need any justification. And the reason you can do that is because you're an engineer and you know a damned thing about the domain you're working in.
Structural engineers don't build a skyscraper out of paper mache first, and then rebuild it in concrete when it collapses in a stiff breeze. They just build it out of concrete the first time around.
What so many people thoughtlessly call "premature optimization" isn't "premature" at all, it's just "knowing about the problem and knowing how the computer works". When you deliberately ignore this, what you're actually doing is "premature pessimization". Coding as if you don't know the difference between cache and RAM in 2023 is like coding as if you don't know the difference between RAM and disk in 1973. It's negligence. You know better!
And look at what the Clean Code people want you to do instead. All these rules are geared towards future extensibility. Is that not a form of premature optimization? But optimizing for extensibility, not speed. And you end up with codebases littered with abstract interfaces that only have (and will only ever have) one implementation. But you pay for that premature abstraction in both cognitive load and CPU load. You ain't gonna need it!
I put that into quotes, because to me personally these aren't strict rules - but rather guidelines. And they aren't meant to be pushed to the absolute extreme - but rather be seen as methods/tools used to achieve the actual goal: easily readable, maintainable and modifiable code.
And in my workplace, "performance" isn't measured in cpu-cycles, but rather in man-hours needed to create business value. Adding more compute power comes cheaper than needing more man-hours.
For the most part, it still seems to be a good idea to train new developers to know and understand clean code. It will help them produce more stable, less buggy and more readable code - and that means the code they write will also be easier to optimize for performance, if necessary. But with my work, that sort of optimization seems only ever necessary for very small pieces of code - most definitely not the entire code base.
At least my understanding of the case for clean code is that developer time is a significantly more expensive resource than compute, therefore write code in a way which optimises for developers understanding and changing it, even at the expense of making it slower to run (within sensible limits etc etc).
Depends on the number of invocations of the program.
The “developer time is valuable” mantra is thrown left and right, disregarding how much of that valuable resource will be wasted down the line due to bad implementations.
If we optimize for developer time, let’s optimize across the software’s entire lifecycle, not just that first push of a MVP to production.
We just need to do away with shoddy first implementations, while avoiding premature optimization (or pointless perfectionism).
Which is all good in theory. In practice, management needs to also acquiesce and stop with the impossible deadlines.
I mean, I’ve worked on more than one projects that were sold to clients before they were implemented, and I think I’m not unique in that :)
On the other hand, when I studied dynamic dispatch and stuff like that, I don't think enough people told me "when you do that, you are making this tradeoff".
I feel like it's worth sharing this kind of knowledge in order to make better informed decisions (possibly based on numbers). There's no need to become an extremist in either direction.
The claim of 20x program performance difference is overblown. Compilers can often remove virtual function calls, JITs can also do it at runtime. Virtual function calls in a tight loop are slow but most of your program isn't in a tight loop and few programs have compute as a bottleneck.
Measure your program, find the tight loops in your program and optimise that small part of your program.
(This is not to endorse writing code that is harder than necessary to understand, or failing to document the parts that are necessarily hard.)
This also presupposes that making a fast program is a lot more work. However poor performance is usually due to negligence rather than a lack of optimization effort. All you need to write reasonably fast code by default (without micro-optimizing) is:
1. a good knowledge of available algorithms 2. a good understanding of the problem
1. Is a one-time investment on the programmers part that benefits all future programs they write. There is no marginal cost to being familiar with what's available in <algorithm>. 2. Has a marginal cost, but it's probably a time saver anyway. Measure once, cut twice.
I prefer a much simpler rule: if it's easy for the CPU to execute, it's likely easy for you to read too. That means: no deep nesting, minimise branchiness (indirect calls are the worst), keep the code small and simple.
Content wise: his examples show such increases because they're extremely tight and CPU-bound loops. Not exactly surprising.
While there will be gains in in larger/more complex software by throwing away some maintainability practices (I don't like the term "clean code"), they will be dwarfed by the time actually spent on the operations themselves.
Just toss a 0.01ms I/O operation in those loops; it will throw the numbers off by a large margin, then one would just rather pick sanity over the speed gains without blinking.
That said, if a code path is hot and won't change anytime soon, by all means optimize away.
Edit: the upload seems to have been deleted.
As a small, amusing, s(n)ide note, Casey's rants about the slowness of the Windows terminal[2] that ended up in Microsoft releasing an improved version[3], were based him wanting to implement a TUI game as an exercise in SCG and the terminal being too slow.
[1] https://starcodegalaxy.com/ [2] https://news.ycombinator.com/item?id=31284419 [3] https://news.ycombinator.com/item?id=31372606
I mean, yes, if you do something completely fucking idiotic like put an IO operation inside a tight calculation loop, then all your speed gains will vanish. But I don't see how that refutes anything.
Do you remember the last time you had to do a tight calculation loop, and not only that, but one that significantly impacted the total runtime? Personally I do, it was roughly 15 years ago writing a raytracer.
I can imagine that happening in game/3D dev, DSP, emulators, ML, and maybe some other types of software, but even in those cases one already has dedicated libraries and hardware to extract the performance from where it can be extracted.
I mean, Python is slow as an old dog, yet it gets most of the ML fun.
Well yes, that's.. actually the point. Indirectly, anyway anyway.
> I mean, Python is slow as an old dog, yet it gets most of the ML fun.
Python driving some UI and logic, with a ton of optimized Fortran and C driving everything hot.
I'm sorry to say that this argument is not even wrong.
As programmers, it is not useful to us to think about that 99.9% of time. The 0.1% of the time is literally our entire job.
"Most of the universe is not Earth, so why do we spend so much time thinking about things on Earth?"
Everything follows from this. It's not that game devs are so much cleverer than other devs, they are just faced with first hand feedback of "does the game code hit the frametime budget" constantly and the whole dev org is committed to that.
1. Prefer polymorphism to “if/else” and “switch” - if anything, that makes code less readable, as it hides the dispatch targets. Switch/if is much more direct and explicit. And traditional OOP polymorphism like in C++ or Java makes the code extensible in one particular dimension (types) at the expense of making it non-extensible in another dimension (operations), so there is no net win or loss in that area as well. It is just a different tool, but not better/worse.
2. Code should not know about the internals of objects it’s working with – again, that depends. Hiding the internals behind an interface is good if the complexity of the interface is way lower than the complexity of the internals. In that case the abstraction reduces the cognitive load, because you don't have to learn the internal implementation. However, the total complexity of the system modelled like that is larger, and if you introduce too many indirection levels in too many places, or if the complexity of the interfaces/abstractions is not much smaller than the complexity they hide, then the project soon becomes an overengineered mess like FizzBuzz Enterprise.
3. Functions should be small – that's quite subjective, and also depends on the complexity of the functions. A flat (not nested) function can be large without causing issues. Also going another extreme is not good either – thousands of one-liners can be also extremely hard to read.
4. Functions should do one thing – "one thing" is not well defined; and functions have fractal nature - they appear do more things the more closely you inspect them. This rule can be used to justify splitting any function.
5. “DRY” - Don’t Repeat Yourself – this one is pretty good, as long as one doesn't do DRY by just matching accidentally similar code (e.g. in tests).
Any time you have these common rules of thumb in any part of your life you need to evaluate whether or not they are appropriate. But just because they’re not infallibly universal, doesn’t mean they’re wrong, it just means life throws complex situations at us sometimes, and that we need to be pragmatic and flexible
In the clean code version, your compiler will remind you to implement calculateArea, calculateNumberOfVertices, calculateWhatever, and so on and so forth.
With his version you have to add a new `case` to every switch statement and hope you didn't miss one with a default case, because the compiler won't catch it.
You don't need virtual functions and polymorphism to remove his switch statements. Just compose over inheritance.
The compiler not catching it is a limitation of the language he uses, not the limitation of the general concept of switch / pattern matching. Scala, Haskell, Rust do catch those.
> If you follow the principle "switch statements over [X]", try to add a new shape down the line and see how quickly you run into problems.
And who said you'd ever need to add a new shape? Maybe you will need to add a new operation? Try to add a new operation `calculateWhatever` and see how many places of the code you need to chnage instead of just adding one new function with a switch.
Often you really don't know in which direction the code will evolve. Most often you can't guess the coming change, so don't make the code more complex now in order to make it simpler in the future (which may never come).
You're thinking of defaultless switch statements. Rust and Typescript catch those issues as long as you don't add a default. I'd already thought of it when I typed it
>And who said you'd ever need to add a new shape?
ô_o
>Maybe you will need to add a new operation? Try to add a new operation `calculateWhatever` and see how many places of the code you need to chnage instead of just adding one new function with a switch.
You do realize that your "one function with a switch" does have to cover every case exactly like the clean code version, it's not magic you're just fitting it all into one switch/case, and you'll probably end up extracting those case: blocks into separate functions anyway.
And on top of that you're not safe from a colleague adding a useless "default" case at the end ouf ot paranoia and making your compiler not catch a future problem when a new shape gets added.
Productivity is achieved in spite of clean, not thanks to it.
The difference is, adding "default" cases to complete switch statements happens all the time, you've probably witnessed it, I know I have, whereas colleagues adding a dozen classes to add two numbers doesn't.
You're also never safe from meteorites crashing your building, let's talk about real antipatterns
But the same problem exists with interfaces and virtual dispatch! If you provide a default implementation at the interface / abstract class level, then the compiler won't tell you forgot to implement that method, because it would see the default one exists.
And one should be able to fix things that get out of hand when they start getting out of hand. Architecture should be based on current facts, not on our fantasies about the future.
For some reason we're always making this weird assumption that future engineers working on a problem are going to be less capable than us. So we decouple things in advance for them. Those future idiots won't know how to architect for scale, but we, today, with our limited domain expertise, know better.
The reality is that in most cases we are those future engineers. And, unsurprisingly, "future we" tend to know more, not less.
Worse, I think, is that it hides the condition, which can be arbitrarily far away in space and time.
Certainly dynamic dispatch can be very useful, and the abstraction can be clearer than the alternatives. A rule of thumb is to consider whether you'd do it if you were writing in C, using explicit tables of function pointers. If that would be clearer than conditional statements, do it.
(In case it isn't obvious, I'm talking about languages in the C++/Java vein here.)
The word "clean" in reference to software is simply an indicator of "goodness" or "a pleasant aesthetic", as opposed to the opposite word "dirty", which we associate with undesireable features or poor health (that being another misnomer, that dirty things are unhealthy, or that clean things are healthy; neither are strictly true). "Clean" is not being used to describe a specific quality; instead it's merely "a feeling".
Rather than call code "clean" or "dirty", we should use a more specific and measurable descriptor that can actually be met, like "quality", "best practice", "code as documentation", "high abstraction", "low complexity", etc. You can tell when something meets that criteria. But what counts as "clean" to one person may not to another, and it doesn't actually mean the end result will be better.
"Clean" has already been abandoned when talking about other things, like STDs. "Clean" vs "Dirty" in that context implies a moral judgement on people who have STDs or don't, when in fact having an STD is often not a choice at all. By using more specific terms like "positive", or simply describing what specific STDs one has, the abstract moral judgement and unhelpful "feeling" is removed, and replaced with objective facts.
Probably a better link is the blog post because the author updated it with the new replacement video a few minutes ago as of this comment (around 09:12 UTC):
https://www.computerenhance.com/p/clean-code-horrible-perfor...
First off, the number of problems where having an analytical measure of shape area is important is pretty small by itself. Second, if you do need to calculate area of arbitrary shapes, then limiting yourself to formulas of the type `width * height * constant` is just not going to cut it. And this is where the entire optimization exercise eventually leads: to build a table of precomputed areas for affinely transformed outlines.
Throw in an arbitrary polygon, and now it has to be O(n). Throw in a bezier outline and now you need to tesselate or integrate numerically.
What this article really shows is actually what I call the curse of computer graphics: if you limit your use cases to a very specific subset, you can get seemingly enormous performance gains from it. But just a single use case, not even that exotic, can wreck the entire effort and demand a much more complex solution which may perform 10x worse.
Example: you want to draw lines? Easy, two triangles! Unless you need corners to look good, with bevels or rounded joins, and every pixel to only be painted once.
Game devs like to pride themselves on their performance chops, but often this is a case of, if not the wrong abstraction, at least too bespoke an abstraction to allow future reuse irrespective of the use case.
This leads to a lot of dickswinging over code that is, after sufficient encounters with the real world, and sufficient iterations, horrible to maintain and use as a foundation.
So caveat emptor. Framing this as a case of clean vs messy code misses the reason people try to abstract in the first place. OO and classes have issues, but performance is not the most important one at all.
You’re right that once you’ve done the 15x performance gain that Casey demonstrates, the code is pretty brittle and prone to maintenance problems if the requirements change a lot. But I think we can have our cake and eat it too by maintaining our code with the simple path available to fall back to if we get a complicated new requirement. Need to add in complex cases that weren’t thought of before? Add them to new switch cases that are slower, and then keep looking for more performant ways of calculating things of need be.
In some sense, it is a good thing. It creates natural backpressure to scope creep.
The "Clean Code" developer can add a new complex shape type to their program in 15 minutes; just subclass here, implement the methods, done, no changes to other code. No performance impact - the program remains as badly performant as it was before.
The Casye-style developer will take 15 minutes and come back to you with:
"Oh, we can make the calculations for this new shape approximate to, idk. +/- 50%, by adding another parameter to the common equation; this will, however, cause 1.1x performance drop across the board. We could make it accurate by special-casing, at the cost of 2-3x perf drop for the whole app. We can also spend a person-week looking into relevant branch of mathematics to see if there isn't a different equation that could handle the new shape accurately with no performance penalty."
"Or, you know, we could just not do it at all. Why exactly do we need to support this new shape? Is supporting it worth the performance drop?"
Whichever option the developer and their team chooses, their software will still remain an order of magnitude more performant than the "Clean Code" style. Which is to say, those devs are aware of the costs - the "Clean Code" style is so ridiculously wasteful, as to not even notice such costs in the first place.
Limiting them to a handful of precomputed ones is the kind of limitation only a software developer can live with and love.
And once you can handle arbitrary polygons, writing and maintaining special case code has to be justified with evidence that your code is spending enough time on area calculations that it's worth optimizing and maintaining extra lines of code for.
First we have to write the code this way. Next year we have to write it in the other way. Then it has to be done in the first way again.
It's so hard to prioritize things when you ask someone why something has to be done the way they say and they are not able to give a real answer. I can do my job when option A means faster program and option B means more memory usage but I can't do my job when option A means faster program and option B is just the way it "should" be done.
Should performance be talked about more? Yes. Does this show valuable performance benifits? Also yes. Is performance where you want to start your focus? In my experience, often no.
I've made things faster by simplifying them down once I've found a solution. I've also made things slower in order to make them more extendable. If you treat clean code like a bible of unbreakable laws you're causing problems, if you treat performance as the be-all-end-all you're also causing problems, just in a different way.
It's given me something to think about, but I wish it was a more fair handed comparison showing the trade offs of each approach.
There are far worse crimes than having code structure that spans several files.
Sure, but what's the advantage of having your code split over several files? Yes, you can jump between them with an IDE, but that's still a disruption, and it makes it harder to see common patterns that can help you simplify the code.
As demonstrated in the video, splitting up the switch into multiple classes only hurts readability.
Did I miss an IDE that inlines all those things for you automatically? Because ones that only give you a "jump to definition", or maybe a one-at-a-time preview in a context popup, are not much better than Notepad++. You still don't get to see all the relevant things at the same time.
Plus how often do you need to see all functions? If there's a problem in RectangleArea function, you go to that file, no need to look in CircleArea. If you're adding 'StarShapeArea' you only care that it matches the callers expectations, not worrying about the other function logic.
Yes, if something is understandable in a single file, that’s fine. But also appreciate that too many pieces in the same file can also be confusing and disorienting for people. You also have to consider all the cases in which someone would be opening that file.
You’re basically posing a situation which by its very description is an exception. If anyone is slavishly following these rules, they’re undoubtedly doing it wrong.
Knowing when to apply the patterns is the hard bit. And, in my experience working on codebases worthy of considering these things in the first place, there’s much more to consider than what you’re posing.
This extension is easy enough with in the non-pessimistic case (table based).
That's the alleged selling point, whether even that's true is arguable.
Since while performance can be clearly measured - whether the given "clear code" principles do in fact help with writing "maintainable and extendable code" is something that can be argued back and forth for a long time.
The benchmark is a tight loop where the vtable lookup is a big chunk of the total computation. I don't think one can extrapolate this 1.5x improvement to real code. If anything, it represents an upper bound on the performance improvement you might expect to see.
I also didn't see anything about how the code was compiled. Various optimizations could affect performance in meaningful ways.
No you can't. But other things get worse at a bigger scale. Not all programs make those virtual calls absolutely everywhere so the overhead scales with the program, but many don't pay attention to memory access pattern, and cause their instruction pointer to jump all over the place and trash their instruction cache. Mike Acton have once shown that merely reordering objects by types, while keeping those virtual calls, can help the instruction cache quite a bit just by making sure the same code was called several times in a raw.
Casey prizes performance over everything else.
The problem with virtual calls in a big project is that there is no good way of knowing what is the target of the call, without some additional tooling like IDE. But in case of a switch/if, it is pretty obvious what the cases are.
Then you get people writing their own horrible hacks.
Both clean code and performance oriented design have their extremes.
Clean Code has Spring with Proxy/Method/Factory monster... and hyper performance has the extreme in the story of Mel (i.e. read-and-weep only code).
You get to push back and ask, is it worth the developer time and predicted 1.5 - 5x perf drop across the board (depending on shape specifics)? In some cases, it might not be. In others, it might. But you get to ask the question. And more importantly, whatever the outcome, you're still left with software that's an order of magnitude faster than the "clean code" one.
Clean code "assumes code will be changed/maintained" in a maximally generic, unconstrained way. In a sense, it's the embodiment of "YAGNI" violation: it tries to make any possible change equally easy. At a huge cost to performance, and often readability. More performant code, written with the approach like Casey demonstrated, also assumes code will be changed/maintained - but constraints the directions of changes that are easy.
In the example from video, as long as you can fit your new shape to the same math as the other ones, the change is trivial and free. A more complex shape may force you to tweak the equation, taking little more time and incurring a performance penalty on all shapes. Even more complex shape may require you to rewrite the module, costing you a lot of time and possibly performance. But you can probably guess how likely the latter is going to be - you're not making an "abstract shape study tool", but rather a poly mesh renderer, or non-parametric CAD, or something else that's specific.
Sure, but what happens, once you want to start supporting more operations on the shapes?
Exactly - unless you're trying very hard, you're unlikely to beat C++ polymorphism with your OOP code in a different language. Which makes Casey's argument that much stronger. C++ with its relatively unsophisticated OOP and minimal overhead on everything, is as fast as you're going to get, so it's good for showing just how slow that still is if you follow the Uncle Bob et al. Clean Code tradition.
No it isn't. If your C++ compiler isn't devirtualising at all (implied by the article) it'll get stomped on by anything doing inline caching [0] which will generate the switch-case code. The JVM does that for example.
That was from either the 1960s or the 1970s and I don't know that anything has changed in the human ability to read a mangled mess of someone's premature optimizations.
"Make it work. Make it work right. Make it work fast." how to apply the observation above...
Writing "poor code" to make perf gains is largely unnecessary. Though there are certainly micro-optimizations that can be made by avoiding specific types of abstractions.
The lower level the code, the more variable naming/function encapsulation (which gets inlined), is needed for the code to read cleanly. Most scientific computing/algorithmic code is written very unreadably, needlessly.
He never mentioned that people should write horrible code for performance but more like pointing out. Personally and professionally, We stay away from virtual functions as long as possible due to unnecessary vTable lookup every time you want to call a method
Their tool was so dog slow I could see it paint the screen on a modern computer.
I rejoiced when we yanked it out of our toolchain. Most of the advice that it gave was unambiguously wrong.
The answer is - if every single piece of software was written while already knowing the true requirements, the scope of its use and re-use, and knowing the future bugs and security flaws that would appear, then it could be written one time and be BLAZING FAST.
Many parts would still be written in a 'Clean code' style for the necessary extensibility and testability, etc. But many others would be small and near optimal.
THEN on top of that, if the author or an equivalent talent came along and rewrote or supervised the optimization of the regular software, similar to how the article does, your system would be HYPER INSANE BLAZING FAST.
If we are proponents of OOP or Clean code, we need to acknowledge that fact. (i.e. My code may not be important but it all contributes to slowing down the computing world). And if we think the Author is preaching gospel here, you should also acknowledge that because the future is so often unknown when we write code we often have no choice but to fill it with Clean code that can be easily changed later, and sometimes even 'Shit code' that we thought would never be used by anyone.
This is obviously wrong for performance reasons, as operations tend to have high latency but multiple of them can run in parallel, so many optimizations are possible if you target bandwidth instead.
There are many languages (and libraries) that are array-based though, and which translate somewhat better to how number crunching can be done fast, while still offering pleasant high-level interfaces.
I think he understands way much more than how to optimize for 60fps as many commentators here point out.
Say we have this:
int obj_api(object *o, char *arg)
{
return o->ops->api(o, arg);
}
that's representative of how C++ virtual functions are commonly implemented. It gets more hairy under multiple inheritance and such.It requires several dependent pointer loads. We must access the object to retrieve its ops pointer (the vtable) and then access the vtable to get the pointer to the function, and finally branch there.
To call that function a little faster we can go to this:
int obj_api(object *o, char *arg)
{
return o->api(o, arg);
}
in other words, forward the api function pointer from the static table to the object instance. Ok, so now each time we construct a new object, we must initialize o->api. And the pointer takes up space in each instance. So there is a cost to it. But it blows away one dependent load. And the "clean" structure of the program has not changed; it has the same design with virtual functions and all.We could do this for some select functions that could benefit from being dispatched a little faster.
I don't think there is a way in C++ to tell the compiler that we'd like a certain virtual function to be implemented faster, at the cost of taking up more space in the object instance and/or more time at object construction time.
> We can still try to come up with rules of thumb that help keep code organized, easy to maintain, and easy to read. Those aren't bad goals! But these rules ain’t it. They need to stop being said unless they are accompanied by a big old asterisk that says, “and your code will get 15 times slower or more when you do them.”
He isn't against organised and maintainable code, he just thinks the current definition isn't worth the trade-off.
virtual u32 CornerCount() = 0;
you should be able to declare a virtual data member virtual u32 CornerCount; // default value zero
how this would be implemented is that it simply goes into the vtable. ptr->CornerCount retrieves the vtable from the object, and CornerCount is found at some offset in that table, just like a virtual function pointer would be.There is no need to pull out a function pointer and jump to it.
In C I would do it like this
// Every shape has a pointer to its own type's static instance of this:
struct shape_ops {
unsigned (*area)(struct shape *);
unsigned corner_count;
}
// get_area looks like this:
unsigned shape_area(struct shape *s)
{
return s->ops->area(s);
}
// the corner count isn't calculated so it's just
unsigned shape_corner_count(struct shape *s)
{
return s->ops->corner_count;
}
Everyone can override corner_count with their value. What you can't do is implement a calculation which determines the corner count dynamically, but that can be a reasonable constraint.The enterprise code must be easy to change because it deals with the external data sources and devices, integration into human processes, and constantly changing end-user needs. Clean code practices allow that, it's not about CPU performance and memory optimizations at all.
There are no good metrics that measure how "clean code" (atleast the given rules) make the code easier or harder to change and maintain.
All the Java style "enterprise type code" from my experience is bloated, full of boilerplate getters and setters and all sorts of abstractions that often make things harder and not easier to understand/maintain, etc.
However CPU performance is easy to measure, and sticking to "clean code" rules as given in the video demonstrably sets you back a decade in hardware progress/makes the code run 10x slower.
> Clean code practices allow that
This is what you believe, not something you can actually measure as far as I know
Also, his code must have been compiled with an old compiler or less than -O3 as the switch/table version of the code performs exactly the same with Clang and g++ when compiled with -O3.
disclaimer: not a fan of OO regardless.
In a kitchen, clean is a pretty objective concept: no dirt or grime, objects put away with similar objects. Not sure what it means in code, but it seems many people have strong, conflicting, subjective opinions about it. Doesn't seem like a good recipe for productivity or alignment.
I feel like it would be wiser to limit the concept of clean to the eradication of obviously "dirty" or "cluttered" things, like inconsistent style, or naming a module in a way that is misleading about its contents or functionality. Just as all different kinds of buildings can be clean, a code of "cleanliness" should not be so comprehensively prescriptive about architecture and organization. Use more appropriate names for those dimensions of code quality, rather than "clean" as the single stand-in for every good thing.
Esentially the whole point of object orientation is to enable polymorphism without having big switch statements at each call site. (That, and encapsulation, and nice method call syntax.) When people dislike object orientation, it's often because they don't get or at least don't like polymorphism.
Most people, most of the time, don't have to think about stuff like cache coherency. It is way more important to think about algorithmic complexity, and correctness. And then, if you find your code is too slow, and after profiling, you can think about inlining stuff or using structs-of-arrays instead of arrays-of-structs and so on.
A fun story from https://snakeisland.com/aplhiperf.pdf on the utility of people using standard inner product / matrix multiplication operators, instead of hard-coding their own loops or whatever:
> In the late 1970’s, I was manager of the APL development department at I.P. Sharp Associates Limited. A number of users of our system were concerned about the performance of the ∨.∧ inner product on large Boolean arrays in graph computations. I realized that a permuted loop order would permit vectorization of the Boolean calculations, even on a non-vector machine. David Allen implemented the algorithm and obtained a thousand-fold speedup factor on the problem. This made all Boolean matrix products immediately practical in APL, and our user (and many others) went away very happy.
> What made things even better was that the work had benefit for all inner products, not just the Boolean ones. The standard +.× now ran 2.5—3 times faster than Fortran. The cost of inner products which required type conversion of the left argument ran considerably faster, because those elements were only fetched once, rather than N times. All array accesses were now stride one, which improved cache hit ratios, and so on. So, rather than merely speeding up one library subroutine, we sped up a whole family of hundreds of such routines (even those that had never been used yet!), with no more effort than would have been required for one.
Or, just look at SQL. I sure appreciate not having to write explicit loops querying the correct indexes every time I want to access some data.
There are people that are wrong on both extremes, obviously. I’ve worked with one too many people that quite clearly have a deficient understanding of software patterns and try to pass it off as being contrarian speed freaks. Just as I’ve worked with architecture astronauts.
I’m particularly skeptical of YouTubers that fall so strongly on this side of the argument because there’s a glut of “educators” out there that haven’t done anything more than self-guided toy projects or work on startups whose codebase doesn’t need to last more than a few years. Not to say that this guy falls into those two buckets. I honestly don’t think I know him at all, and I’m bad with names. So I’m totally prepared for someone to come in and throw his credentials in my face. I can only have so much professional respect for someone that is this…dramatic about something though.
He is very much aggressive to a degree I am not a fan of, but when it comes to calling out bad practices, I find him more right than wrong. But he is terrible at delivering his message in a way that won't ensure anyone who didn't agree with him already will get pissed off.
Within one codebase the two behave the same because you can just go rewrite your functions, but if those functions are locked away in someone else's library the "dirty" code flat-out prevents you from ever using more shapes than the library author implemented. Want to add a rhombus? An arbitrary polygon? An ellipsoid? Something defined by Bezier curves for some reason? Well, you just can't; sorry champ.
It's an interesting tradeoff to consider, though. Perhaps we write code like library authors too often, or optimise for extensibility when it isn't needed.
Running speed tests on the much-cited mini-example with geometric shapes and their area is unfair and unrealistic, and it does not prove any point.
I think I can see where this is coming from: 'overly clean' OO style will split concerns into virtual one-liner functions without context distributed throughout the universe. For a simple problem, I prefer 'switch'. But that's not a good rule either. For anything extensible, like a GUI, 'switch' would be the wrong choice and virtual much better.
Programmers need to develop a feeling of appropriateness, and restructuring may be necesaary at times.
BTW, the manual loop unrolling in the article is broken and not advisable at all. I'd be angry in code review about such 'optimisations'.
Similarly we can then look at an iterated design and realize the optimized code is frequently going to be harder to refactor or understand (a precondition of refactoring). So now the time to when a customer gets their answer is delayed again.
Optimization step comes long after clean code. Clean code is most useful in the first 2 of the typical 3 steps[1]
1. Make it work (iterations of what working even means) 2. Make it right (iterations of what right even means) 3. Make it fast.
For example, take a look at ClickHouse codebase: https://github.com/ClickHouse/ClickHouse/
There is all sort of things: leaky abstractions, specializations for optimistic fast paths, dispatching on algorithms based on data distribution, runtime CPU dispatching, etc. Video: https://www.youtube.com/watch?v=ZOZQCQEtrz8
But: it has clean interfaces, virtual calls, factories... And, most importantly - a lot of code comments. And when you have a horrible piece of complexity, it can be isolated into a single file and will annoy you only when you need to edit that part of the project.
Disclaimer. I'm promoting ClickHouse because it deserves that.
I think this obsession with clean code is a natural reaction to the overwhelming number of gotchas that seem to just come with the job.
It's a bit like when a parent watches their kid get hurt outside the house and in a complete overreaction locks the kid in the house for life.
Thing is, if I'm programming an airplane control system, I very much want that kind of pedantry I think. I really don't want to make a single mistake writing that kind of program. If I'm programming a video game, just let me write the code that I want. Nobody's going to die if it blows up in my face.
I'm not sure what should be the lesson from all this... Perhaps don't pick C++ unless you absolutely need to?
The iPhone comparisons are extremely cringe. Real application do so much more then this contrived example. Something that feels fast isn’t the same thing as something is fast.
Would I advise beginner programmer’s to read this book? Sure, let them think about ways to structure code.
If he just had concluded with, that it is important to optimize for the right thing that would be fine. But he seems more interested in picking a fight with clean code.
And yes performance is a lost art in programming
Or, more likely, a straw man.
"Clean" exists to provide some solutions to certain problems in TDD. Namely how to separate your logic so that units can be reasonably put under test without an exploding test surface and to address environments which are prohibitively recreated. If you don't practice TDD, "clean" isn't terribly relevant. As far as I am aware, it has always been understood that hard-to-test code has always had some potential to be more efficient, both computationally and with respect to sheer programmer output, but with the tradeoff that it is much harder to test.
It is useful to challenge existing ideas, but he didn't even try to broach the problem "clean" purports to solves. Quite bizarre.
In the example given, all polymorphism can be removed, and the shapes can be stored in std::tuple structure.
And then the operations would be faster than C, since no switch statement would be needed.
The only difference between “modern, clean” C++ and the author’s switch is probably a concept that requires some type attributes.
The example is contrived, and the realization of “clean” code through runtime polymorphism is both dangerous and odd. The whole point of not using polymorphism is to catch runtime crashes at compile time, reduce overhead and improve readability. I know many people who wouldn’t use an object here anyway. Free functions would do nicely, and are infinitely compositional.
But I think there is a reason for the existence of "clean code" practices: it makes devs easier to replace. Plus it may create a market to try to optimize intrinsically slow programs!
Imagine working as a barista with a disorganised bar, a mat on the floor that keeps sliding and a corner is sticking up, and one bag of beans where half the side is decaf and the other is normal.
Now compare that to working in a more common sense coffee shop: everything is in its place, the mat isn't decrepit, and you have multiple bean bags.
In which one do you think it's easier to make coffee?
I meant to say that organisation helps make work easier, including adding performance optimisations.
You see the same thing with microservices. Any of the reading material by the big / original proponents of microservices is actually quite good at giving you all the reasons why they probably aren’t for you. But that doesn’t stop the game of telephone that intercepts the message before it gets do most developers.
So I really just see this whole thing as someone saying “RTFM”, rather than it being any sort of derived nuanced take.
The sooner a professional software developer can get themselves off the treadmill of garbage trendy educational content, the better.
Or, you know, college. Though I can only speak of my local tech college, not full blown university. I don't bear them any ill will --- there's a hell of a lot to try to teach in two years --- but a lot of the things that were taught in my degree, were very dogmatic.
https://www.manning.com/books/data-oriented-programming https://www.dataorienteddesign.com/dodmain/
Also, some good resources are listed here: https://www.dataorienteddesign.com/site.php
https://dl.acm.org/doi/10.1145/356635.356640
The author of the post fails to articulate how we strike a healthy balance and instead comes up with contrived examples to prove points that only really apply to contrived examples.
Casey tells you that following "Clean Code" can give you a huge performance hit for no obvious benefit. And even if "Clean Code" were to be more maintainable (it's not; in my experience, it's actually worse for maintainability), you should still be extremely aware of the cost you're likely to pay down the track. It's not a contrived example, it's literally textbook "Clean Code". I'll say it again: "Clean Code" gives you slower, less maintainable code, and you get nothing from it. Maybe you can afford it, maybe in your use case it's not a big deal, but you should be informed.
Knuth tells you to measure before optimizing, which Casey did. Knuth does NOT tell you "don't worry about performance, you'll optimize later". You quoted Knuth but stopped right before the best part:
> A good programmer will not be lulled into complacency by such reasoning, he will be wise to look carefully at the critical code; BUT ONLY AFTER THAT CODE HAS BEEN IDENTIFIED [emphasis mine]. It is often a mistake to make a priori judgments about what parts of a program are really critical, since the universal experience of programmers who have been using measurement tools has been that their intuitive guesses fail.
To recap: "A good programmer will not be lulled into complacency by such reasoning" - in other words, just because 97% of the code may not need optimization does NOT mean you should not be thinking about performance.
Knuth's point is that when identifying hotspots, programmers were relying on intuition rather than measurement. That's what he meant by "premature optimization". Knuth did not mean (especially since it was the 70s) you should write "Clean Code" that you know has worse performance for little benefit.
And Knuth does not write "Clean Code", by the way.
> The author of the post fails to articulate how we strike a healthy balance
There is no healthy balance between a good idea and a bad idea. Just eliminate the bad idea.
"I like my code to be elegant and efficient. The logic should be straightforward to make it hard for bugs to hide, the dependencies minimal to ease maintenance, error handling complete according to an articulated strategy, and performance close to optimal so as not to tempt people to make the code messy with unprincipled optimizations. Clean code does one thing well" [1]
He never says to throw all ideas of performance out the window when writing your initial run of code. It's just not worth it to dig down into the weeds and micro-optimize everything ahead of time, is all.
But people just take this quote as liberty to completely ignore all notion of performance in their code. Maddening, and a total disservice to Knuth.
And he argues that it's not a straw man.
I mean, dude discovered C++ compiler sucks after over 40 years of trying to make it not suck so much, but ignores the fact that his tools are broken and proceeds to make completely unwarranted conclusions from that.
Needless to mention that software needs to be first and foremost correct. "Clean code" is about reducing the chance of a programmer of making certain kinds of mistakes. And even in the situation where the compiler sucks, it's still worth doing / paying the price in terms of speed, if you can get more confidence of your code doing what it's supposed to. Just like structured programming, "clean code" is an attempt to reduce complexity the author of the code has to deal with.
----
The proper conclusion that should've been the result of his experiments should've been: maybe something went wrong with the language and tools I'm using that even after a massive effort over several generation of programmers and mega-corporations backing that effort, the tools and the language still suck. So, the desirable properties of my programs (i.e. simplicity and ability to be extended) still come at a huge cost.
This approach can be way faster, but is only relevant when you have a lot of entities you need to iterate over. If you have 3-100 objects it would of course still be faster but by negligible amount
So yeah, the book then goes on for a painful 450 page ramble, opinions, and admittedly arbitrary rules. But Martin was at least partially aware of this:
> “Clean code is not written by following a set of rules” — quote from the book!
So really, the person in the video failed to apply the Principle of Charity, which is fundamental in critical thinking. They end up not addressing the interesting claim, and openly attacking a Straw Man.
As for the deeper points implied in the video, they seem –ironically– less fresh:
- Software is slow these days - Performance matters - The way you write code impacts performance - Don't blindly follow rules and generic advice
Groundbreaking!
If anything, the video shows the failures of C++ as a language. Why aren't languages designed to promote maintainability without sacrificing performance? :Rust enters the room:
The more interesting claim that the video's author missed:
> “It is not enough for code to work.” ― quote from the book
He's talking about "clean code", not maintainable code. The claim that "clean code" is more maintainable is an unproven assertion. Whenever I interact with a "clean code" codebase, it is worse in every way compared to the corresponding "non-clean" version, including in terms of correctness.
Do you have much background in procedural or data-driven programming?
"The more you use the “clean” code methodology, the less a compiler is able to see what you're doing. Everything is in separate translation units, behind virtual function calls, etc. No matter how smart the compiler is, there’s very little it can do with that kind of code."
I suppose even that is true, but JIT compilation regularly walks right around those problems. Yes, your code is written to say virtual this, or override that...but the JIT don't care. Is it looking at a monomorphic call site? Or even if it's not, is it ok to think of it as monomorphic right now? Great -- inline away.
All that being said...I once got into a readability tiff over the use of a Java enum in a particularly performance sensitive chunk of code. I went with ints so I could be very, very explicit about exactly what I wanted, and the rather large performance gain...and lost. Yay!
Your mileage may vary, and your measurements may vary.
The gaming devs were obsessed with framerates and efficiency while server devs wanted to decouple modularize everything.
There's no solutions only tradeoffs
OTOH the class design lets someone come and add their new shape without needing to change the original code - so it could be part of a library that can be extended and the individual is only concerned about the complexity of the piece they're adding rather than the whole thing.
That lets lots of people work on adding shapes simultaneously without having to work on the same source files.
If you don't need this then what would be the point of doing it? Only fear that you might need it later. That's the whole problem with designing things - you don't always know what future will require.
You might feel that there are infinities of potential shapes out there but not really infinities of operations.
So by having some operations exceptional faster you could not only save time also you save energy.
Binary Assembly ... ... C++ ... ... Python ... ... Product Manager speaking with words: "Can you make it have more pizazz?
If you don't care about execution speed, I don't want to use it.
But even on the back end, companies seem willing to scale cloud costs rather than make the lumpy and “risky” investment in hiring.
I don’t agree with it but I see it everywhere.
In the process of "improving" the performance of their arbitrary benchmark they make the system into an unmaintainable mess. They can persuade themselves it's still fine because this is only a toy example, but notice how e.g. squares grow a distinct height and width early in this work which could get out of sync even though that's not what a "square" is? What's that for? It made it easier to write their messy "more performant" code.
But they're not done, when they "imagine" that somehow the program now needs to add exactly a feature which they can implement easily with their spaghetti, they present it as "giving the benefit of the doubt" to call two virtual functions via multiple indirection but in fact they've made performance substantially worse compared to the single case that clean code would actually insist on here.
There are two options here, one is this person hasn't the faintest idea what they're doing, don't let them anywhere near anything performance sensitive, or - perhaps worse - they know exactly what they're doing and they intentionally made this worse, in which case that advice should be even stronger.
Since we're talking about clean code here, a more useful example would be what happens if I add two more shapes, let's say "Lozenge w/ adjustable curve radius" and "Hollow Box" ? Alas, the tables are now completely useless, so the "performant" code needs to be substantially rewritten, but the original Clean style suits such a change just fine, demonstrating why this style exists.
Most of us work in an environment where surprising - even astonishing - customer requirements are often discovered during development and maintenance. All those "Myths programmers believe about..." lists are going to hit you sooner or later. As a result it's very difficult to design software in a way that can accommodate new information rather than needing a rewrite, and yet since developing software is so expensive that's a necessary goal. Clean coding reduces the chance that when you say "Customer said this is exactly what they want, except they need a Lozenge" the engineers start weeping because they've never imagined the shape might be a lozenge and so they hard coded this "it's just a table" philosophy and now much of the software must be rewritten.
Ultimately, rather than "Write code in this style I like, I promise it will go fast" which is what you see here, and from numerous other practitioners in this space, focus more on data structures and algorithms. You can throw away a lot more than a factor of twenty performance from having code that ends up N^3 when it only needed to be N log N or that ends up cache thrashing when it needn't.
One good thing in this video: They do at least measure. Measure three times, mark twice, cut only once. The engineering effort to actually make the cut is considerable, don't waste that effort by guessing what needs changing, measure.
> Most of us work in an environment where surprising - even astonishing - customer requirements are often discovered during development and maintenance.
Another fun fact. The author of the video works in an environment where rapid iteration is absolutely vital. I'd pay good money to see a TV show where his style of programming ("spaghetti", as you claim) run laps around your "Clean Code". Because it would. For example, he wrote a terminal emulator in a weekend to prove that Microsoft doesn't have a clue about how to write code (I assume they also have many Clean Code people, and that it would take them about 6 months to write a terminal emulator from scratch).
The reason why this video mentions performance is probably because 1) the author has a course on performance and 2) it's something you can objectively measure.
If it were me, I'd not even bring performance into discussion, I'd just say that "Clean Code" significantly hurts readability and editability (and thus maintenance). But then you jump in to say that the non-Clean Code version is "an unmaintainable mess", and then we go around in circles. Which is probably why performance is his main point.
You'd like to pay money to be assured that you're right? I prefer to have an informed opinion based on actually trying stuff out and measuring†, I also find that as a result I don't feel the need to pay for validation.
> For example, he wrote a terminal emulator in a weekend to prove that Microsoft doesn't have a clue about how to write code
So, your thesis is that writing a terminal emulator - software which is pretending to be hardware that existed 40+ years ago, shows that this person is great at handling surprising requirements changes during development ?
An insistence that the only thing we can measure is performance and specifically speed and therefore that's the only thing that matters is nonsense. It's just that this technique does so poorly when judged on maintainability that it's no contest.
Try it, first add the hollow box example shape, in the Clean Code this is very easy and we'll notice immediately it's de-coupled, colleagues working with abstract shapes don't care at all about our Hollow Box shape, it all just works with the abstract APIs.
The less-clean switch approach is a little bit hairier now, but it's very possible although we may notice now our objects are all bigger again, even though perhaps few are hollow boxes, they're all bigger as they all need to track the possible state of a hollow box. So that's actually a significant performance degradation for some applications in our supposedly "high performance" solution...
The table-driven approach needs a rewrite though, the F*W*H simplicity doesn't apply any more, there are a few "minimalist" approaches, all of them awful compromises waiting for the other shoe to drop - so perhaps a big bang rewrite is called for. Ouch.
Now, having learned from our hollow box experience, let's add Regular Star Polygons next. These are pretty interesting shapes - but we're shape classes so no reason we can't handle this, the stars have a defined area and a defined number of vertices ("corners"). But while the Clean Code here is very tractable, the dirtier approaches start to hurt pretty bad now.
Notice that under Clean Code the exact implementation of Regular Star Polygons doesn't affect anybody else, their code all still works regardless. For example maybe we should sub-class popular examples like the 5/2 and 6/2 rather than taking p and q parameters, doing this works fine under Clean Code, since it's nobody else's business.
† EtA: One of the most important innovations in years has been Godbolt.org, Matt Godbolt originally worked on this tool to examine exactly this sort of question, it's one that comes up early in the talk you liked - can we safely use actual C++ iterators? Wouldn't an old-fashioned for loop be faster in some cases? Matt's answer was "Yes", you can use iterators, the iterators produce exactly the same machine code and the tool he used to demonstrate that evolved into Compiler Explorer, the godbolt.org site today.
I know I'm right, I'd pay money to see the embarrassment of the presumptuous Clean Code people who think that they can write maintainable code better than those who write software that matters (like Linux or Postgres, as mentioned before).
> So, your thesis is that writing a terminal emulator - software which is pretending to be hardware that existed 40+ years ago, shows that this person is great at handling surprising requirements changes during development ?
No. My thesis is that this guy can write code better than people who are supposed to be in the top 1% of the developers (well, it's Microsoft, not a web app sweatshop).
He works on games, where you not only have to iterate very quickly but you also may need to completely change direction halfway through the project. They can't have an "unmaintainable mess", otherwise they're not shipping the game, so your premise is wrong from the start. Also, games are much more complicated to program than the average Clean Code Crud app project that Uncle Bob bikesheds on.
> An insistence that the only thing we can measure is performance and specifically speed and therefore that's the only thing that matters is nonsense
Nobody said that, how are you coming up with this stuff?
The point Casey was making is that you're paying a significant performance penalty (speed, in this example) by doing Clean Code, which is true. You're denying yourself very basic performance techniques if you close your eyes and pretend to not see the internals of each shape.
> Try it, first add the hollow box example shape,
What if you don't have to? What if those are all the shapes you have to support? But there's ten billion shapes, you chose to do Clean Code and now you have slow code for zero benefit. Ouch.
On the other hand, if you want to add more shapes then the problem you're solving changed, and therefore you need to change the code (not really the tragedy you're making it out to be). I don't get this obsession, "code should change as little as possible"; it's actually a "careful what you wish for" moment because class hierarchies tend to become very rigid and difficult to change. Good luck making significant changes when your 100k+ loc program relies on a particular class hierarchy being in place.
> the iterators produce exactly the same machine code
That's great but it seems you're trying really hard to interpret Casey's video in bad faith. The way I understood it, he used an old fashioned for loop to avoid detracting from the main discussion, as not everyone knows what are the internals of STL iterators and how they translate to machine code.
You believe you're right, which of course you do, you almost can't help it.
Still, you keep mentioning Linux and it's worth a moment to consider that Linux actually does have the flavour of problem the Clean Code is modelling here, and it does indeed solve it the way Clean Code recommends and which Casey warns you will have egregiously bad performance. Several whole CPU cycles slower than Casey's spaghetti in fact.
Let's look more closely. In Casey's Clean example the Shape subclasses have to carry a table of functions to call to find out e.g. the area of that Shape, and then as we walk our array of Shapes we use these tables to call the appropriate area function. This costs us a dereference, which takes a few cycles.
With any luck, as this was described, your knowledge of how an OS works warmed up and you realised, "oh, that's, that's actually how the OS kernel works". Yup. Obviously Linux isn't written in C++ and so it has to actually hand-write the code to make tables of functions, so actually it's a bit clearer in the source, we can see that sure enough the implementation of CIFS for example and the implementation of XFS, and the implementation of FAT all just provide tables of functions.
So if Casey is right, shouldn't Linux be crushed by some 20x faster OS made by Casey or similarly minded games programmers (maybe Jonathan Blow) without these low performance tables ? Nope, there two good things to know here.
Firstly, this flexibility is immensely useful and it turns out most users can't live without it. A product which is 20x faster but can't do what you need is at best irrelevant, at worst a nuisance, a distraction. The demo shape project really needs to be able to be extended for arbitrary shapes.
Secondly, and this is often much more important in practice yet Casey just completely ignored it, this is a fixed cost overhead. Multiplying two numbers together is almost no work, so the overhead dwarfs the real work done, but in real software we are often doing a lot of work, yes even despite the "Single responsibility" rule and as a result the overhead is negligible in practice.
While we're in here it's important to notice that the overhead occurs because of the actual indirection, which was incurred in the C++ by the use of virtual function calls to several distinct types of Shape, and in our Linux kernel example by the use of several different filesystems via a table. You are not paying this overhead merely for the existence of functions to allow separation of concerns although that's what Casey implies.
> You're denying yourself very basic performance techniques if you close your eyes and pretend to not see the internals of each shape.
You're spending an unaffordable amount of your finite engineering resource on handling other people's problems in all of your code if you insist on peering inside everything as Casey does in this toy example.
The reminds me of the argument in Hare (another programming language from people who figure they're smarter than everyone else) that they shouldn't provide generics because you ought to build a custom data type each time you need something reflecting exactly what you needed each time. When I read that I decided to look briefly at how their compiler used a hashmap (IIRC) and of course it was buggy because it's all hand rolled and so it has a typical mistake you might make in your first attempt with a hashmap - as every hashmap in a Hare program is its own custom first attempt. I believe they subsequently fixed the bug after I reported it, so that's nice - until next time.
> He works on games, where you not only have to iterate very quickly but you also may need to completely change direction halfway through the project
Casey's only notable actual game project completed seems to have been The Witness, Jonathan Blow's second and more ambitious but arguably less successful game. Casey has worked in the games industry for a long time, but like Blow he's spent a lot more time telling other people he could do better than he has spent on actually demonstrating that.
In the time it took Jon and Casey to ship one game, John Carmack's id Software shipped Doom, some Doom sequels and Quake and some Quake sequels. Jon and Casey are not people to take your cues from if you want to have agile software development practices or ship products in a timely fashion.
Getting a bit desperate here, eh? On one hand, having tables of function pointers does not introduce any code constraints, you can switch to switches or anything else at a moment's notice; class hierarchies are much more rigid (some random Torvalds quote, "all your code depends on all the nice object models around it, and you cannot fix it without rewriting your app"). On the other hand, Clean Code is fundamentally tied to OOP and classes. Here, straight from the horse's mouth [1]:
"This expectation of polymorphism is the essence of OO programming. It is the reductionist definition; and it is inextricable from OO. OO without polymorphism is not OO. C and Pascal programmers (and to some extend even Fortran, and Cobol programmers) have always created systems of encapsulated functions and data structures. It does not require an OOPL to create and use such encapsulated structures. Encapsulation, and even simple inheritance, is obvious and natural in such languages. (More natural in C and Pascal than the others.) So the thing that truly differentiates OO programs from non-OO programs is polymorphism. You might complain about this by saying that polymorphism can be achieved by using switch statements or long if/else chains within f. This is true, so I must add one more constraint to OO. The mechanism of polymorphism must not create a source code dependency from the caller to the callee."
In short, C is not OO because it doesn't do polymorphism (as understood in the context of Java-like OO languages rather than a mystic "it kinda looks and does the same as OO, thefore C is OO"). Furthermore:
"FP and OO work nicely together. Both attributes are desirable as part of modern systems. A system that is built on both OO and FP principles will maximize flexibility, maintainability, testability, simplicity, and robustness. Excluding one in favor of the other can only weaken the structure of a system".
Which is to say, if you don't do OO then you're not doing Clean Code. On a side note, Robert Martin obviously thinks he could write a Linux that's more flexible, maintainable, testable, simple and robust, but he's leaving it as an exercise to the reader.
https://blog.cleancoder.com/uncle-bob/2018/04/13/FPvsOO.html
> You're spending an unaffordable amount of your finite engineering resource on handling other people's problems in all of your code if you insist on peering inside everything
It's not quite so dramatic, you don't need access to STL's internals, just the structures you're working with anyway. To simplify: if you have an algorithm that deals with shapes then don't abstract away the concrete types, don't try to impose a taxonomy, don't pretend there's a magic shape interface that generalizes everything, don't try to fit the square box in a round hole. Instead, allow the algorithm to deal with concrete types. This is in fact the most flexible approach - you won't find yourself having to rethink your class hierarchy when one of your classes doesn't neatly fit into the general picture. I've been in the situation where at the end of a project it becomes very obvious that the chosen class hierarchy is actually unsuitable for easily adding more features and improving performance, but by that point the effort to restructure the hierarchy is equivalent to a rewrite. But hey, we had Clean Code. The key point is that Casey's approach allows you to easily optimize for performance if needed; Clean Code does not.
> Casey and Jonathan Blow are too slow to deliver products
Fine, take Mike Acton. Same ideas, except he had to ship games on demand. You won't catch him doing Clean Code.
This means that in 2-3 months you end up with a codebase that is very difficult to work with, team members tripping over each other due to bad deps and abstractions and your iteration time start shooting up.
Doesn't seem like a realistic avenue to choose except maybe when coding to a final spec?
there's nothing novel in this video, really nothing to do with clean code. This is same sort of thing you see with pure python versus numpy
This video from 2015 where Casey is interviewing Mike Acton: https://www.youtube.com/watch?v=qWJpI2adCcs
Also this course where he specifically talks about Numpy and pure Python: https://www.computerenhance.com/p/python-revisited
It has extremely clean & easy to discern code. But it’s also not the most performant.
This is something very good to have in mind, but it must be applied strategically. Avoiding "clean code" everywhere won't always provide huge performances win and will surely hurt maintainability.
tl;dr polymorphism, indirection, excessive function calls and branching create a worst-case for modern hardware.
You shouldn’t do things that make your code utterly slow though.
"Game developer optimizes code for execution as opposed to readability that 'clean-code' people suggest".
There are few considerations:
- most code is not CPU bound so his claims that you are eroding progress because you are not optimizing for CPU efficiency is baseless
- writing readable code is more important than writing super optimal code (few exceptions: gaming is one)
- using enums vs OOP is not changing the readability at least to me
I think we can have fast and readable code without following the 'clean-code' principles and at the end it does not matter how much gain we have CPU cycle-wise.
To each their own, but I don't find Casey's performant version less readable, I don't see the need for so many abstractions.
It does create implicit coupling. If you try to add a new shape you will run into the problem.
In the clean code version, your compiler will remind you to implement calculateArea
With his version you have to add a new `case` to every switch statement and hope you didn't miss one with a default case, because the compiler won't catch this one.
It's a crap way to code
The correct way to deal with this, is refactoring to a more maintainable code once you know the amount of shapes will wildly change, As soon as we get too many shapes as the problem has changed. You can only pretend to know what is the best architecture for a problem when you have dealt with it several times.
Clean code apologists pretend their single time dealing with website backend is proof enough that clean code works and that it works for every problem and that it has to be the default approach and is the most readable for most problems. It is a total insanity for something that can't be measured with any tool.
Edit:
I fully understand that "premature optimization is wrong" but using these "guidelines" is premature optimization of scalability and maintainability. Somehow when the "premature optimization" is about things you people want that's somehow okay? pff
Also, I don't find clean code readable, it looks like complex, un-refactorable garbage to me 9 out of 10 times. No wonder why people are so fucking scared of rewriting a class and act is if it will take months to do so, this ideology makes impossible to actually play around with your code, you can't neither make it more readable or more performant, you are locked in with a sluggish collection of dozens of files even for the simplest of problems.
The take on switch statements is covered in "The Pragmatic programmer" as well, which coincidentally is much less criticized when it comes to books about clean code.
The way to fix the switch is getting rid of the class "shape" and making it an interface, then implementing the interface in each shape, as a non-virtual method. And then you don't let people inherit. They can compose instead.
Performance is unaffected, you get rid of the switch, the compiler catches your mistakes for you, and everyone's happy
No. The techniques are there to let you change your codebase as you learn more about the problem.
a) "these techniques allow for an easy to refactor codebase". I profusely disagree, its easier to refactor a function with some ugly switch than a behavior disseminated over dozens of files.
b) "we are actually careful and not draconian about these techniques only when appropriate", which I don't agree either, as in every experience I had interacting with people that believe in clean code in meatspace, was an obsession with having things done their way, assertions about "code smells" which were literally just not doing whatever they wanted.
maybe there's a c) I'm not seeing. But seeing this thread on its own its already high evidence that these two notions are clearly there in the CC community.
"we are actually performant" becomes "well, actually we are readable" becomes "well, actually we are testable" becomes "well, actually humm... just shut it and write on our code style". Some posts on this thread even talk about people not using these guidelines being evil and having to be ejected. Just imagine how people not suck up on this ideology look at it.
But IMO framing it as clean vs performant is a mistake
Wtf? First time I hear about this one, and it sounds like a dumb dogma.
> It’s a base class for a shape with a few specific shapes derived from it: circle, triangle, rectangle, square. We then have a virtual function that computes the area.
Quite literally the first and simplest example for why you should prefer composition over inheritance[^1] (that and ducks and chickens).
Good strawman. I am unconvinced.
1: https://en.m.wikipedia.org/wiki/Composition_over_inheritance
It's like database denormalization, it may violate normalization principals but it when applied to a well designed database is a valid optimization technique when done with proper understanding of the implications of said optimizations.
More importantly though, we are willing to sacrifice raw performance for developer experience and higher maintainability because developer time is expensive, and most stakeholders would prefer that you can add feature xyz in a reasonable time, over feature xyz running marginally faster. If ease of development and maintenance weren't important, we'd just write everything in assembly and bypass all these abstractions altogether.