There's a downloadable zip file there where you could probably figure out who the offenders were. Ripe did say that it was a mix of both ISPs and Cloud providers.
https://labs.ripe.net/author/giovane_moura/dns-ttl-violation...
Edit: There are also probably some corporate MITM type "content filtering" caches that are screwing things up too, by caching web pages longer than they should.
* Design your system assuming a hostile environment and that propagation time is on the order of hours.
* Draw and document a hard line above which you consider it your user's problem; i.e. you start assuming the world has updated after 2 hours and any stragglers can just get errors.
And it doesn't have to be an antagonistic interaction. I fully suspect that many if not most cases of TTL violations either a) have a good explanation or b) are unintentional and easily fixed. Let's open the dialog and start improving things.
I worked on dns gslb for a long stretch at Facebook^WMeta, and didn’t see an excess of bad actors. The vast majority of users follow our dns changes in an orderly fashion. Most delay sources to clients themselves.
The "up to 48 hours" is commonly used because that's the ttl of many things that matter. For instance, NS records for names in the .com zone have 48 hour ttls.
You can give better estimates to your users if you know the state you're transitioning from. For instance, a brand new .com domain would be 15 min negative ttl plus the .com zone file update frequency.
Generally I just assume a good old fashioned "48 hours", like in the olden days, and I have yet to be disappointed.
Perhaps you're overly cynical?
Either way, a good demonstration of the value of empirical evidence.
I suppose it's reasonable that you could provide a better estimate for new domains and transfers based on past experience and existing TTLs. But it will be an estimate. And the estimates would be individual or sub-group ones, like "estimate for a new .com domain" and "specific estimate for transfer of this domain", etc.