Lessons Learned from Writing Over 300k Lines of Infrastructure Code
blog.gruntwork.io
blog.gruntwork.io
I'm going to interviews right now where people are trying to hire a "Senior DevOps Engineer". They seem to think hiring this somewhat low-level employee will magically imbue their organization with DevOps principles, and suddenly everything will be cheaper and faster. I tell them it may take them 2 years to build a solid DevOps practice, assuming they do everything right, and they frown. Why can't we just have the DevOps right now? To that I reply, go hire a company that specializes in DevOps. And again they frown.
I don't think it should be this hard. I think we can consolidate all the important knowledge and pre-packaged work, collaboratively, to empower teams to actually get shit done quicker and better. I think we can make it so nobody ever has to go through the painful process of learning by trial and error ever again. And, of course, I think it should all be free. So I've started a wiki, and I hope I can convince a few more people to help contribute content.
GruntWork, I love what you're doing, but I'm also hoping to make your business obsolete. :)
Totally agree. Unfortunately, you are up against the industry that is not interested in getting shit done quicker and better.
"Institutions will try to preserve the problem to which they are the solution"
Noble intent, but time for that is not yet. The state of cloud infrastructure practices today is a primordial soup of ideas, marketing and the myriad of tools trying to survive on the market, like Terraform vs Ansible vs Cloudformation :)
Good infrastructure is hard, there is global networking with CDNs, DNS, CI/CD, Git integration, HA, monitoring, security, you name it. Thats why you will not find the code on Github, only blog posts on Medium, code brings bread and butter to so many people, including me :), why sharing it ?
And most important. To create really reusable by many infra code is even harder. You need to feel what features should be there for everyone, and what is bloat. It'll take many times the effort of writing non-reusable architecture for one company.
> .. And again they frown
Made me smile :)
Just an fyi... that goal is only possible for ideas that have stabilized and broad consensus has converged into widely accepted best practices. Yesterday's "hacks" that are insanely complex get vetted (or modified and replaced) and eventually get folded into the baseline of Normal Things Everybody Knows. But... the world isn't standing still and it evolves in complexity. Therefore, there's always new complexity that requires new hacks that everybody reinvents.
At a high level, the cycle looks like:
chaos --> create order --> new chaos again
We see a chaos of hacks and complexity. We then notice recurring patterns in the disparate hacks and extract some unifying principle that can "simplify" the complexity. If we stop here, it seems we've tamed disorder and everything should be able to be documented and people can stop wasting time repeating everyone else's past mistakes. But all we've really done is raise the baseline of systems sophistication. There's always a new delta of chaos and complexity.E.g. In the 1980s, sysadmins might implement a rudimentary sync of userids with a shell script and copying "/etc/passwd" files around. But we notice that's chaotic and stupid and doesn't scale. So we rationalize the system with a more unified technology such as a LDAP identity server. But utopia is still not achieved because while LDAP might be ok for employees' sign-on accounts, it's not suitable for millions of customer accounts of a B2C web business. The complexity treadmill is neverending!
>GruntWork, I love what you're doing, but I'm also hoping to make your business obsolete.
If companies like GruntWork and AWS are properly doing their job, they will never be obsolete because the changing world keeps throwing new demands at us and this constant delta of new chaos is always there that needs to be "solved".
Where can this be found?
The difficulty arises from the underlying complexity. The underlying layers are necessarily complex, and the law of conservation of complexity applies. Documentation can be improved, and it will improve with time - and managed services will get more powerful. But yeah.
Took me a while before I found that, would have been helpful earlier on.
Unfortunately, I've also seen plenty of consulting firms selling "Devops Solutions" and productizing everything because their customers look at everything as products and solutions that they acquire and presume that their culture is fine as-is. While I don't doubt that these firms make plenty of money I don't think the practice will be viable long-term because the problems are entirely due to culture rather than tools or even people.
Deming’s work was not widely accepted in the US, but his approach was accepted at a place many Americans do recognize for superior quality and consistency - Toyota. Every other devops process of feedback and organizational structure is derived squarely from Deming’s principles. I always thought it’s rather ironic that a society known for being relentlessly individualist adopted corporate practices of conformity and a society known for being conformist took on an organizational philosophy that put more control into individuals. Maybe it was done this way to counter natural social inclinations in wider society? I’m not sure honestly. My gut feeling is that Taylorism is precisely how large militaries historically work and following WW2 this was easier for American workers to adopt.
Not that I am surprised when I see hypocrisy. I just like to call it out.
After many years of real world experience I developed the following rule of thumb. A competent, experienced developer should produce about 20 lines of code per day, on average, to have any hope for a decent code quality. I am talking about high level languages with big standard libraries and reach ecosystems like python. At that rate 300,000 lines of code is ~41 menyears.
EDIT: “Autocorrect” got me.
That's simply measuring the wrong thing altogether. Lines of code per day produced is simply not a valid metric at all. It's just a bad metric period no matter what number you put on it. You cannot measure the quality of a dev by the lines of code he produces, at all.
if (false) {
// my daily 20 lines here
// ...
}And I'm saying that's still a terrible measure and not accurate. LOC is simply not a valid statistic to even look at. It's an example of lying with statistics; you could take a thousand good developers and you'd probably find their LOC per day counts all over the board. The amount of code you write is not a measure of how good a developer you are whether it's high or low. That's not how you measure a good developer.
The OP is trying to say looking for high LOC counts are not a measure of a good developer, and he's right; but he's failed to realize yet that the metric itself is what's bad, not the values derived from it.
Good developers don't have to be highly productive, but they might be; what matters are the results of their efforts. Do they produce good programs, with few bugs, that age well over time, and are easy to change to adapting circumstances. Good developers write good code; some write more than others, but LOC tells you nothing about how good a developer is whether that number is high or low.
Like any writing, it's the quality of the code that matters, not the quantity. You cannot measure quality by looking at quantity.
Writing is not the best analogy -- because of the typical churn-edit cycle -- but to use your example, of course reaming out pages of text in one day will, on average, be of a lower quality than a virtuoso author who carefully chooses his/her words and maybe produces, on average, a paragraph a day. If you were to tell a layperson that a good writer can output a paragraph a day, they'd probably find that counter-intuitive because it's such a minute amount; "Surely anyone can write a few sentences in eight hours? Even I could do that!" They're making the same false-entailment, without seeing the craft that went into those paragraphs/day, or realising that there will be days when said author writes nothing and others where they're in the zone and write an almost-perfectly-formed chapter. That doesn't stop "paragraphs/day" from being a bad metric, nor does it imply that anyone who writes a paragraph a day will end up with anything good.
I can't say I agree with your rule of thumb. How'd you reach that number? In my experience the number fluctuates wildly depending on what tasks you're working on. Your output will usually shoot up while a project is young or you're working on a new well-defined feature, and it'll go down as you focus on fixing bugs or improving performance. The best days are when you write negative lines of code.
Lines of code is usually not a good signal for code quality. If I pull in a huge new dependency in order to quickly solve a problem then it might seem like my project is maintaining high code quality, but that doesn't paints a complete picture. For each included dependency you should probably have at least one person in your team responsible for maintenance, which includes stuff like: keeping it updated, tracking security issues, checking releases for breaking changes or bugs, updating your code to handle changes or notifying others if the scope is too large, etc. Ideally, they should also be sufficiently familiarized with the code-base to fix bugs or other serious issues which affect you. These points are only meant to serve as a rough example of some of the responsibilities one might be expected to shoulder, but it should be understood that requirements vary based on a large number of factors.
On the other hand I put out hundreds of LOC per hour, when writing the basic HTML for an internal crud app.
This metric is as useless as it can be. Even as an average. It has too much to do with your role and tasks.
Beyond that, if you're only writing 20 lines of code a day each and every single day, then you're either wasting a lot of your and your employer's time, or your role isn't just programming and you have other things to attend to.
True. ... as Cloudformation template code.
I’m presently going through something similar to the authors situation as I’m migrating my leads into Microsoft Dynamics and learning to make the platform for my use case (real estate brokerage). It’s not “infrastructure” but it’s the foundation for my business
I can say that even something as simple as correct implementation of software is not something that should be done at speed, or without serious consideration towards portability.
For example, because I’m solo and also working another full time job, my emphasis is on completing the migration ASAP energy though I’m aware of small errors in the database that might bite me later.
Also, while I’m comfortable with MS’s record of maintaining enterprise software, it’s slightly suffocating to feel that I will effectively loose my time investment should they decide to abandon Microsoft Dynamics
So, yes, I got to the same conclusion. Basically, solve every part of the problem:
- in isolation/as focused as possible
- as a library
- with the easiest to use and most composable API you can come up with
After that, if you have to, you can still glue them together into a framework-like thing. But at least, anyone can pick any feature out of it without succumbing to madness. Also, this is a great way to avoid the whole "Big Ball of Mud" problem that forces you to pull a whole ecosystem into your project although all you wanted to do was log something.
If you go with largish modules you need to be smart when you build it, if you build smaller modules you don't have to be that smart. Dumb is good, dumb is easier to explain and read for others (including yourself in 3 years).
So, be dumb keep it simple, whatever that is in your case. If the code is easy maybe a giant repo is good, if the team is small. If the team is big maybe you have other means of validating access to your repo, but if not you need to split it up. But then commits need to be synced when pushed over several repos..
Edit: make sure it's possible to make clean commits and rewrite the whole code in small steps. Things Will Change(tm)
I use Ansible quite a bit. When I first started, based on some examples I found, my play books were pretty complicated, with a lot of conditional steps, doing things (or not) based on results of previous steps, etc.
What I found is that it's a lot easier to have a lot of roles that each do one little thing, and then use them in a play book as needed.
A role might be as simple as installing a package, templating a configuration file, or maybe even just changing just one line in a config file.
When roles are small and do just one thing, they are easy to combine in many different ways, and it's more obvious what's going on.
Like how every time I need to renew a SSL cert, I need to reread the man pages for how to make certs and cert requests. You'd think a replicable procedure could be documented, but there's just enough variables so that it's slightly different every time.
Unfortunately, most enterprises are still not willing to outsource their IT platform to the real pros and thereby make it someone else’s problem to keep that infrastructure up and running. They want to build their own in-house competency. But the problem is, if you’re a true DevOps pro, why would you work for some random company in which IT is considered a cost center, rather than working for either a hosted platform provider or a (for lack of a better phrase) “independent provider of DevOps”?
IT and security managers are able to convince business executives that IT infrastructure needs to be kept in-house, even though today’s typical enterprise has no chance of actually building and maintaining a infrastructure that is 1% as good as what they could rent from a dedicated provider.
"The real pros" you speak of are also under the same cost pressures as everyone else. I have encountered many such services where after the "real pros" have set up a system, they are replaced by low cost staff who can maintain and minimally extend while billing...
And if you finally have it more or less done it in your local environment / company, you switch and suddenly you start again.
But you know, my salary still increases anyway so shrug (still looking for a solution)
Let me fix this for you. Vast majority of developers are actively denied access to the infrastructure by their managers. They pretty much know everything what an infra eng knows (and even more, after all all this catchy cutting-edge apps like Kubernetes/Elasticsearch/etc. was written by them) or even more.
Blame the managers and no-one else.
As it should be. Devs make apps, Infra guys, SRE, make these apps run in production. Different skills, PHP vs sysadmin's work.
Ask your frontend guy about SR-IOV, just for fun :)
I worked in all kinds of companies. The lamest and slowest ones were these outdated thinking ones. The most productive and highest quality products were delivered in those companies where the team was responsible for their apps networking, load balancing, monitoring, autoscaling, DB sharding.
You think these are somehow special skills. But you didn't provide any arguments.
You'll get insights, as you correctly worded, good. Thing is, designing production infra takes much more than that. Years of telecom experience for example, where downtime is not an option, deep linux and network knowledge, HA designs and lessons from past mistakes. That cannot be learned overnight, and you get sixth sense if it looks right or not.
Add to that AWS SA Pro cert, Python and you'll be getting somewhere :)
> downtime is not an option
HAHAHA - sure. Of course.
You wouldn’t hire an operations engineer to architect your software, why would you have a software architect setup your network? Software may be tough to change but I can assure everyone that the millions of miles of cabling out there in legacy data centers out there would be much better if someone with sufficient networking backgrounds had gotten involved earlier. Operations is a cost center, but it can be a force multiplier for your revenue centers of less than 1.0
If you don’t have the knowledge and experience to lay out a network and expect to grow the footprint beyond more than a handful of instances, you should probably at least get a 1 hour session over coffee with a moderately experienced engineer that has actually thought about these things before. Heck, I’d do it practically for just beer money because I’d consider it an act of goodwill for any future infrastructure engineer that has to work on the environment.
That said, every time I use terms that aren’t directly related to code or software architecture, I feel imposter syndrome creeping up on me—as in writing much of this post. (I’d welcome comments or additional learning resources!)
Also, I’m not sure AWS is the best example, their network terms and best practices seem far more oddly named and unique to AWS than what I’ve heard from Google’s recent conference presentations for enterprise adoption of Google Cloud.
Your infrastructure / SRE peers will appreciate you more for writing software that is easier to deploy and maintain when you're clueless about networking than if you understand networks and designed a system that is extremely stateful when it offers no technical advantage to be that way (databases / caches get a pass, everyone else writing software in 2018 has no excuse to keep state on a machine for longer than a business transaction window). High performing software teams deploy often and have the culture to encourage it in a healthy manner - there is no excuse to write software that is deployed weeks or months as huge chunks of changes unless you fall into very niche enterprisey domains and even still you should at least be making production-releasable software daily.
And like it or not, AWS is the enterprise cloud by fiat now, so it'll determine the bar and etymology for other vendors to meet and (hopefully) exceed.
We also have another bunch of actors in the industry: devops and sysadmins who because their old cushy position was kind of unnecessitated by the emergence of these new services (after all ANYONE can create an AWS account) realized that they have to come up with new smokes and mirrors to rationalize why are their role is so important.
I am very very much against this stance and proposition and I will fight this thinking in every possible forum.
The possibility is real and the risk/cost is very low to empower dev teams in today's cloud landscape so there is really no need to prevent this happening.
And this discussion also reminds me of an old Uncle Bob article [0] where Uncle Bob summarizes the (back then situation) as follows:
> I witnessed the rise of a new job function. The DBA! Mere programmers could not be entrusted with the data – so the marketing hype told us. The data is too precious, too fragile, too easily corrupted by those undisciplined louts. We need special people to manage the data. People trained by the database companies. People who would safeguard and promulgate the giant database companies’ marketing message: that the database belongs in the center. The center of the system, the enterprise, the world, the very universe. MUAHAHAHAHAHAHA!
Now replace data in the above excerpt with infrastructure and you will arrive to the same conclusion as I did. Q.E.D.
[0] https://blog.cleancoder.com/uncle-bob/2012/05/15/NODB.html
The devops / Agile philosophy of everyone being empowered to do most things works pretty well when people want to do all these things, are invested in the outcome together, and are at least vaguely competent in their tasking. However, the approach has limitations when it comes to tasks that nobody wants or can do but is still important to the business. It's even worse when something is important but nobody even knows it because of groupthink blindness.
I don't think I'm being hypocritical in recommending specialists for topics I know something about while advocating for empowerment in other functions because if my previous employers / clients knew what they were doing with infrastructure and healthy software development practices, they could have grown much more before needing to hire someone to do it full-time for them. I really don't want to have to re-IP another awful network again and have to tell leadership that you have to incur downtime to do it because their software can't handle database hiccups like when failing over to a hot replica. It is boring, unfulfilling, stressful work to me that - even worse - offers no tangible business value when done well but when done inappropriately is an albatross.
Compute infrastructure across different industries is in an overall state of health where everyone loves junk food but is starting to recognize its harm, some vaccines have been developed for the flu but nobody gets it or the vaccine costs $20k per shot for some people, doctors for Hollywood actors and pro athletes debate publicly over which lifting program is more optimal, and the two fitness trends are competitive decathlons and walking from their car to their desk instead of taking a Bird. In comparison, software is much further along with at least a vague sense of a board of medicine in different states (that is determined through a pageant and feats of strength, not experience in Mississippi), people are taught about the dangers of junk food, there is a debate on GMOs (Uncle Bob is strictly against it, I see some positives although DBs are much more controversial now than the well-researched topic of GMOs), and only the literally crazy people don't believe in use of vaccines.
You and I both agree that tons of people needlessly hire a personal trainer when the information to start exercising is out there and basically free now. What I think you're suggesting is to "just start jogging and it'll work out - everyone can run without a trainer helping you" but I think it's mistaken not because I think trainers are required. Right now most cloud providers don't give you shoes for free because they want to sell you Air Jordans or hiking shoes to recuperate their substantial investments, the roads are totally unpaved except for paths through lemonade stands charged by how fast you run, and I've seen a lot of people hit by cars while running because they kept stopping to pick glass out of their feet all because the common theme is they started running with socks and they were "forced" to keep running. I don't think I'm being unreasonable in saying that by default people start walking with socks on because they think personal trainers are too much when flip flops can work really well until you need to start running. By your view every other former sysadmin is now a personal trainer trying to get people into some shoes when we can do fine without one, and while I can see that I'm personally not the typical sysadmin type because I started off as a developer only caring about running fast and have learned starting off on the wrong foot can cause serious problems that can be very cheaply and easily corrected. Perhaps we are disagreeing over how much those flip flops cost or how difficult bad footwear is to discard?
Then you found a bunch of unicorns, congrats :) In real life good app devs and good infra devs are different people. I mean good not as in nice and easy personalities, the skills. Do all your app/web devs have AWS certificates?
And by production I mean Netflix like infrastructure.
But we are talking about devops in the context of gruntwork. Which is targeting developers/dev teams and then making unsubstantiated claims about those devs' understanding. That's why I say what I say.
People who propose that managing their precious infrastructure cannot be bestowed on those inferior developers are the real hindrance to any business which want to go forward fast. They simply don't realize that the lack of decentralization of project management and lack of empowering teams creates a terrible communication overhead and grotesque, half baked solutions.
I wouldn't allow most engineers I know to touch infra code. It's just a very different skillset.
Or did you really want to give arguments but you forgot it?
You cannot get junior engineers contributing to infra, or most Semi-senior engineers either. You need a ton of upfront thinking about releases, policies, backouts, risk, etc.
And tbh, if you're thinking about those things it makes no sense that you also focus on user-facing features.
I've worked in companies where "DevOps" was thought to mean: "Devs have a lot of access, even root/dba/etc. access, and while they're sleeping or if they get stuck on something, they should punt it over half-finished or half-broken to the infra eng / devops group"
^^ This is so incredibly common, especially in small/startup environments. It's a nightmare I actively aim to fix when I join / consult at companies.
There's instability in the environment, inconsistencies, or just a mess, and you parachute in there to fix. Step 1, kick out all devs from any systems they have too much access (unless they're special, knowledgeable ones and know and understand what you're trying to do), Step 2, stabilize the system. Step 2 involves environment separation, real automation that more than one person can use, setting up monitoring, and much much more.
By the description of it, the root cause of the problem is your setup. Dismantle the infra eng team and make the dev team to monitor their own apps.
By reading of your explanation devs should have more access and responsibilities.
You set up a half baked process and you are blaming the devs for it? Not cool.
It's incredibly common.
Not that I blame them, if you can work just 9-5, why wouldn't you?
Dismantle infra eng and make devs do everything? I support that, for one reason: a year from now I can charge consulting dollars when I come to fix these unstable, poorly documented, environments with minimal automation and HA.
Guess what? We knew how to configure private clouds, loadbalancers, sharding and high availability. We pretty much know everything about Kubernetes/Amazon ECS et. al and related technologies and we didn't have any problems with monitoring or low level networking among others either.
Anyone can learn this shit. There is nothing special about it.
Don't hold your breath for seeing your consulting dollars, LOL. :D
Dislike: 300K lines of infrastructure code, but all he talked about was DevOps issues? Surely he learned something about writing good code, too, and I hope he didn't write 300K lines of Terraform code.
If you look at the company, their product is providing infrastructure code, it was not a byproduct.
If your things are wandering out of compliance by themselves, you have bigger problems, imho.
But what I prefer about CF is the “easy button”. If something weird happens or I can’t figure out something, I can use our business support plan with AWS.
As of 2 weeks ago[0] drift detection exists, but coverage is poor[1].
[0] https://aws.amazon.com/about-aws/whats-new/2018/11/aws-cloud...
[1] https://docs.aws.amazon.com/AWSCloudFormation/latest/UserGui...
https://itnext.io/things-i-wish-i-knew-about-terraform-befor...
Furthermore, transitioning from Cloudformation to Terraform is much, much easier than the reverse - you can't import existing resources into CloudFormation management, period (you can't even write the aws-prefixed tags to try to confuse it unless you're an AWS employee or something).
Nothing wrong with starting with CloudFormation or any other provider-native deployment description language if you're 100% in AWS and will stay that way. For the rest of us, Terraform is basically the only choice that can work besides paying for eye-watering expensive boutique tools that were developed a decade ago and have more limitations than Terraform.
What a luxury!
I think you do need to invest in getting over that initial learning curve such that your tools aren't "magic." You have a conceptual mental model of what is going on. And I agree that takes much more than a token effort. Yet for me, that also stops short of mastery.
I always aim for simple to use tools, if I need to I can dig deeper and just-fix-the-problem. Or replace it at a whim.
So I agree with you.
Infrastructure code is really tough when it's scaling, so keep your modules in shape/order and small/easy.
At a lot of companies, asking "can I have a week to play around with these new tools/frameworks before I get started with them" will be met with "no way". Of course, not saying that's good: often the new tools are strong enough that you would still long term get massive development time savings even with the upfront costs. But a lot of managers and PMs are only focused on "shipping" ASAP, never mind the technical debt or maintenance costs (and to be fair, sometimes you need to ship ASAP... but sometimes you don't)
Doesn't apply if the new tools are forced by management, but often it's the developers driving new tech.
That's the reality I was trying to capture with my original comment. For example if you're moving from on-premise servers or a simpler cloud provider (Rackspace, Digital Ocean, or Joyent) to AWS there is no way to escape learning AWS, and the last thing you want to do is manage everthing through the web console so you will need to learn a new tool to manage the infrastrucutre when your old tools don't work with the new stuff.
Management and sales types hear about how great, fast, and agile the cloud is and don't consider that it also takes some time to do things right, and there's is no way around learning lots of AWS specific bits at the very least.
Seriously though, what I was suggesting is that some companies make promises to their customers, on behalf of their engineering team without understanding (or properly estimating) the impact and technology involved. Then it's a mad rush to get something delivered. I understand this is a bad way to do business but it's a sad reality.
Yes, and then everybody is upset that things aren't delivered on time/riddled with bugs. There's no way escape out of sync management.
It feels like there's a personal vendetta here?
And you're absolutely right, we should. It just seems like a losing battle at times, that's all I was trying to convey.
> .. It just seems like a losing battle at times
Bad management. Managers in tech companies must have tech background. I was burned before, joining big banking corps, where managers are just bureaucrats, quit with regret for a lost time.
Next time, I will vet the management layer before joining.
https://www.oilshell.org/blog/tags.html?tag=shell-the-bad-pa...
As for example implementations review cloudposse in depth or take a look at the https://github.com/travis-ci/terraform-config repo.
This answer is skewed towards infrastructure as code. Often conflated are things such as configuration management & provisioning.
In the production grade checklist, there is nothing mentioned about metrics.
Infact there is no mention in the post of metrics or graphs.
This implies that everything is done via logs, which is just horrific at scale.
Everything should emit metrics: o hits per second? metric o Memory use? metric o upstream service response time? metric o That new lib you wrote? metrics
This is especially important with microservices. Open trace is grand, but thats for after you've found where the problem is. Your metrics should be your single pain of glass that indicates the health and performance of your system.
It's a pretty good example of doctoring headline worthy titles. IIRC the author gave a talk of a similar name at the hashiconf recently.
It also includes any random mess of shell scripts and the like that manage the lifecycle of infrastructure.
[2] https://docs.aws.amazon.com/AWSCloudFormation/latest/UserGui...
[3] https://docs.openstack.org/heat/latest/template_guide/openst...
E.g. a simple plan in Terraform gives you the possibility to say how many of what instances should be up in given zones.
Edit: every piece of tool used here is used for infrastructure, controlling them is infrastructure code: (it's massive!) https://landscape.cncf.io/format=landscape
You could instead use something like terraform, which allows you to write code specifying your requirements and then run that code. This allows other devs to review your code, takes a lot of human error out of the equation, and is much more sustainable.
When he says infrastructure code, I think he means code that keeps the infrastructure running.
Some examples that come to m would be terraform files, dockerfiles/docker-compose files, jenkinsfiles, bash scripts, etc. Basically, code that keeps the servers running.
The largest code base I worked on (a map reduce based billing system), also had for the time early 80's a fairly complex set of JCL that could compile and build all the systems module's on dev and also push it out to the 15-16 or so live systems.
This also handled all the glue that held together the map reduce.
The first time you run your template it creates all of your infrastructure and is usually smart enough to figure out dependencies.
After you make changes to your template and run it again it knows based on the changes in your template whether it can modify the existing resources or whether it needs to delete and recreate your resources.
Interestingly, Dockerfiles brought back a bunch of procedural configuration management. We had migrated to ansible for all our server-level configuration. But as we've adopted docker/containerization in recent years, simplifying our applications (now separate containers, rather than severs on common servers) has such a reduced level complexity such that simple docker files with `apt-get install foo` are much preferred.
His points are fine, but shocker its the same principals that apply to all code, infra isn't that special, thought we established that years ago.
Also starting by boasting about the number of lines of code you have written to achieve something is asking for trouble.
It feels like many in the 'devops' community that are from an ops background are rediscovering software principals, this chap isn't alone in that!
What I do see is 'devops' typically changing code in critical areas without taking sufficient care, but thats more a cultural thing than infra code being special.
I don't see a lot of build pipelines with automated tests and deployment for IaC. Standardized processes like pull requests with code review are also not ubiquitous.
With terraform one typically writes a template, does a plan, then apply and hopes for the best. What I've seen and the OP mentions the same in his talk is that things occasionally go terribly awry in ways that common software development practices would prevent.