Yes, I have worked with these products at a moderately large scale. I have also developed a wiki-like CMS from scratch myself, with templating, macros, and everything. The issues I saw would occur at any scale above "tiny". Essentially as soon as scale-out is needed for a dynamic site, caching also becomes mandatory.
Caching is not well understood at all. People think they understand it, but probably don't actually know most of the pitfalls, especially in the face of failure in the general case.
With Wikimedia, many of the issues are not super important. E.g.: transaction integrity is not relevant. The occasional 404 or eventual (in)consistency problem is also not a big deal.
Similarly, the user-based rendering is also fairly easy to handle. The vast majority of the content (the rendered wiki text) is identical-ish between users. The headers, footers, CSS, etc... do change. This can be handled on the client-side in a variety of ways, typically via JavaScript. Even with mostly static content, headings can be altered on the way out "at the edge" by a CDN or CDN-like system.
This is my point: If you're going to cache things in something like a CDN, then you're 90% of the way there to static content anyway! Take the leap and go the whole way to get the benefits. Some things need to be done 100%, otherwise it's like being almost pregnant.
Some random examples of static vs dynamic problems/comparisons:
- You mentioned templating: In the static world, you can regenerate all static pages based on the new template asynchronously at whatever slow "batch process" rate you desire. In a dynamic system, a change to a template would typically invalidate the entire cache and blow your CPU budget instantly, dragging down the synchronous rendering path into molasses. This can be managed, but it's complex and difficult. Similarly, code changes can similarly result in either mixed/corrupt cache content or instant CPU spikes. At least with static content you can manage the rollout by content instead of by server. A trick you can pull with static content is pre-generate the updated pages side-by-side and swap instantly. This is impossible with dynamic content generation. You either get mixed content as the caches slowly expire, or instant load spike.
- Variations such as mobile/non-mobile pages: These add to cache pressure, and can result in sudden performance cliffs where going from 90% cache utilisation to 110% can cause dramatic spikes in load. If you use static content, your "utilisation" is known in advance and changes slowly. Everything is served at the same speed, always. In fact, with systems like S3, your speed goes up as your data volume increases because of the way the sharding works.
- Memcache and the like are hilarious to me. These days the "standard" is to have layers upon layers of caching to paper over the fundamental bottleneck of the database tier. SiteCore has so many layers of caching that I lost count. Is it ten? Eleven maybe? Whatever. The point is that keeping that straight in your head as a web developer is so difficult that SiteCore keeps all caching off by default, murdering performance for most sites most of the time. It's just too "difficult" to have it on by default because devs would lose their minds. Just the access control issues of providing devs with access to shared Redis clusters or CDNs for purge operations is a task all by itself. In large enterprise, this is almost never done and then it becomes a tradeoff between cache TTLs and freshness/consistency. I've lost count of the number of times I've heard some dev tell users to "clear their browser cache" as the "fix".
- Cache purging: you can hide a fundamental performance problem under a layer of caching for years, have it grow to monumental proportions, and then blow up your production site for days while everything slowly recovers from 1,000% load. This has caused several large-scale outages that have hit headlines.
Look at it this way: Wikipedia is something like 99% read-only access and 1% write access. With a static hosting model the VMs would only need to be scaled to handle the write-throughput, not the read-throughput. The content could be put on cloud storage like S3 and that's it. The whole site could be hosted of a handful of small VMs just for HA/DR!