My guess is they're doing both.
Anecdote time. I worked at a company where one project was over-provisioned on dedicated hardware and another auto-scaled in the cloud. The over-provisioned project was much cheaper, had significantly better response times and was easier to manage. It was load tested to handle over an order of magnitude more traffic than the all-time-peak and even though fully over-provisioned, it was cheaper than the baseline usage (and slower, and harder to manage) cloud solution.
There was a very long period of time where SSDs were commonly available from everyone but cloud vendors. For some workloads (like databases), that resulted in a massive difference.
This is still true for bandwidth. You can look it up yourself. Off the top of my head, in the US, you can find a server with 100mbps dedicated port for < $300. But that AWS bandwidth is over 3K. So that's less than 1/10th AND you get a relatively [relative to what you can get on AWS] powerful server vs just the bandwidth.
Back when the C4 instances were announced, I ran unix bench on them as well as a dedicated i7-4770. You can see the i7 was quite a bit faster, and, if I recall, was less than half the price
https://gist.github.com/karlseguin/5a6a45ace2048545b6c3 vs https://gist.github.com/karlseguin/a659ef87b3a4a5d590e9
I think database workload is still where the average app would see the biggest difference. A properly configured server with a battery backed raid adapter and proper dual network NICs will blow RDS out of the water for raw latency and cost and, most noticeable to me, consistent performance.
Unfortunately, fewer and fewer companies seem to be offering servers with BBU and dual NICs. And those that do are charging more...so that's also helped close the gap. IBM really screwed up. They wanted to compete with AWS so tried to turn Softlayer into an AWS clone, rather than focus on what Softlayer did better and fix the issues (automation and ddos mitigation come to mind).
However, I think your rationale begs a question: how much of "the automation and services that major cloud providers have" (scare quotes only for readability) is needed by runners of metal? For instance, autoscaling can be mooted by overprovisioning, which would still incur only a minor cost increase. Multi-region is similarly cheap. A lot of the remainder seem like productized functionality that is fairly implementable locally if desired.
We did two games, one over-provisioned on metal, the other on auto-scaled cloud based infra.
The cost is significantly higher. We ended up with a hybrid to control for cost. But our needs are long-lived sessions which does not fit the elasticity of the cloud model well.
I wouldn’t be surprised if on NYE services like chat take priority and get extra resources from background jobs that can wait a few hours.