I think your comment is smug as well. The author of the article has considerably overloaded the term "testing", reasonably having a lot of people have a knee-jerk reaction of "Don't do that". I get the impression that this is by design, and the article is intended to provoke discussion via flamewar.
I think the best way to combat this is to simply avoid too much discussion on such articles until they are rewritten to be more clear and less liable to cause flamewars. I have some views on the thesis of the article, and practical experience backing those views – but I simply won't express them because I don't want to get embroiled in the numerous minor arguments caused by the confusing terminology, with little to no actionable information. An example of a pattern that you'll see repeating across all the comments in this thread:
Person A: "We take our time and test the application thoroughly under all sorts of load in a preproduction environment. We run end-to-end tests on every build. We value work/life balance, and try to minimize testing-in-production."
Person B: "When did the article say not to do that? It never says you shouldn't test outside prod, it says you should also test in prod and that's a superpower! You're totally misunderstanding the article!"
I would disagree. There is zero overloading of the term "testing" as that is an already extremely broad term of art that would seem to clearly apply to every example of production testing provided in the article.
> with little to no actionable information.
The article absolutely provides aome actionable points and breaks them up under "Technical", "Cultural", and "Managerial"
> An example of a pattern that you'll see repeating across all the comments in this thread:
The article repeatedly covers the exact points mentioned in your example exchange so the people you see having that exchange are those who at best only skimmed the article. I find that the light readers can sometimes dominate comment threads early but those comments eventually do become out outnumbered by the more interesting discussion. People who take the time to read carefully and think respond slower and thus tend to be back loaded.
> I have some views on the thesis of the article, and practical experience backing those views – but I simply won't express them because I don't want to get embroiled in the numerous minor arguments
That is unfortunate. The only way the discussion improves is when people do take the time to state their views, even when they don't have the time to follow up on replies. The only way to combat vapid discussion is to plant the seeds of better conversation.
The thing that bugged me was that these were not treated as a cost/benefit trade off but simply as a fait accompli.
Over time I've come to believe that an appreciation of nuance and cost/benefit tradeoffs are at the heart of effective testing, but culturally the practice is steeped in dogmatism and absolutism. This exhibits all of that - e.g. "control freak managers", "only one represents reality" and "saying not today to the gods of downtime".
Maybe it applies upwards as well? The article is kind of all smug about "I test in prod" too :)
We have loads of redundancy, and dedicated "test in production" machines/datacenters to test actual production loads in actual production-scale sets machines.
Now, tests in production usually involves a pair of hands and 2 other people looking behind (dev + ops + dba), and requires a well defined rollback procedure and post-mortem. We still have absurdly high SLA (99.999%).
No amount of testing in the bank has ever been able to spot the weirdest issues, so we continue while really trying hard to make the prod "pilots" (we try not to call these tests) as routine as possible. But, we still find crazy issues, notably those only the client notices (so wrong specs they couldnt prevalidate by plugging their monster system on our monster system in QA)
We split by exchange traditionally because it made the most sense when we started 30 years ago in Asia, but... We re more and more splitting clients in groups where we can and allocating them cpu power indeed.
I envy the stock market people because while they must handle even higher volume, they can shard per stock itself and have just one jursidiction.
But non-critical services? Sure
You can't acceptance test "Google's search", because at a minimum doing so would require some kind of reverse proxy that is itself part of the system and can't be acceptance tested without...and turtles etc.
Another way of putting this might be that there isn't a good way to acceptance test "airplane manufacturing companies". There isn't a set of acceptance tests you can run to ensure that Boeing is performing to spec before having it build real airplanes.
It's clearly not perfect though, except perhaps if the test aircraft turned itself into a crater on the runway, there is very little one aircraft can do to effect the function of all others.
In software it's hard to get this kind of guarantee, a new delete with a missing where clause in your canary environment is going to take some time to clean up after (assuming you have backups etc).
Especially not in a world where it is so easy to control traffic to your application in a way where 95% get on the stable solution and 5% get on your staging solution.
Hell, this is why you have LTS and non LTS solutions of the dev tools you use yourself.
It's not whether production is orders of magnitude greater that preprod. It's whether production is orders of magnitude more diverse. If you've designed your product so that every user is a snowflake, that's on you, or at least the sales team. That functionality surface area is all undocumented features. Bragging about not being able to test all of the features you've promised users is like the opposite of a humble brag. What would you call that, a Dunning-Kruger brag?
I've worked at a few places that didn't understand this. A couple eventually got it (begrudgingly, and with a hint of resentment. The other never did, and hemorrhaged money and good people.