If stated briefly, it tends to come across as self serving and disingenuous, and I don't have time to add the necessary political and technical backstory to avoid that. But fuck it, here's the brief, high-level.
Starting in the 1990s, I was in the 'devops' team inside of Network Engineering in an enormous and enormously successful (to this day) retail chain. At the time, we were told by AT&T, our biggest WAN provider, that we were the largest centrally managed network they were aware of. Even larger than any one of their own various management regions. When I left there in the early 2000s, there were tens of millions of IP addressable devices on the network, and my team wrote the software that managed and monitored most of 200,000 Cisco and Nortel routers, switches, access points, and other kinds of devices.
At the time, this retail chain was, on average, opening or relocating a new store every day, and each store had a surprisingly large network. Also, we were pioneers in just in time delivery.
So, the network had to be extremely agile (we had people all over the world, all the time, being paid to mess around with our hardware in the form of upgrades, installations, rollouts, etc) and exceptionally stable.
Store shelves would start to go bare after an hour of network downtime at one of our hundreds of distribution centers, across 10 countries. (The 'bareness' would happen many hours later, of course.)
For most of that time, I was the technical leader of the automation team inside of Network Engineering. NE itself had on the order of 40 or 45 network engineers, plus the 5-7 people on my team.
At that scale, automating 99% is about as good as automating 0%. EVERYTHING needed to be automated. And, as things changed, rapidly, my team's automation had to be ahead of it, all the time.
For most of the years I was there, my team delivered spectacularly. We even hit 99.9999% global network availability one quarter. Note: this is absolutely not a bullshit number. The sense of urgency and focus demanded that all assumptions were tested all the time, and availability metrics were very accurately calculated.
The monitoring system I designed and we built was state of the art at the time, and in some ways, even compared to the offerings today.
True story: an NCR contractor (who we widely used to do the in-store hands on work) decided to start stealing our equipment in Southern California. He knew we had redundant store routers; the frame relay circuit plugged into one, and the dial backup circuit plugged into the other. He knew that when one router was to be taken for service and another wasn't immediately available, the procedure was to plug both network lines into the remaining router; our automation handled that and kept dial testing.
So he walked into a number of stores, did that, and walked out with an expensive router.
When he was caught, the data from our monitoring system was featured in the prosecutor's case against him.
We guaranteed within less than a minute of accuracy that any managed device that fell off of the network would be noted and reported. And so it was, and he did go to jail.
From a technical perspective, I'll just say that the entire system was built around the paradigm of message oriented programming. Neither I nor any of my team members knew about erlang at the time, but the system we made organically had many of the same features and strengths.
And, finally, just before I left, I was pushing forward into using machine learning techniques to do truly predictive network monitoring and management. The initial results were extremely promising.