Most sites still don't have valid HTML.
EDIT: my link is old. 98% of the top 100 sites had invalid HTML in 2021, in 2022 we've managed to hit 100%, great job everyone!
For example, Wikipedia doesn't validate [0] because its CSS has "aspect-ratio: 1" which looks ok to me regarding the spec [1].
[0]: https://jigsaw.w3.org/css-validator/validator?profile=css3sv...
[1]: https://w3c.github.io/csswg-drafts/css-values-4/#ratio-value
IMO the validator's definition of "invalid HTML" is just too strict; it should only count parse errors and completely non-sensible things. And the specification is also too strict at times; on my own website I have "Element style not allowed as child of element div in this context." This is because on some pages it adds a few rules that apply only to that page and this is easiest with Jekyll. I suppose I could hack around things to "properly" insert it in the head, but this works for all browsers and has for decades and why shouldn't it, so why bother?
If the specification doesn't match reality, then maybe the specification should change...
Today the complexity lies not in the robustness of the specs, but in the sheer number of of them, and their many interactions. I mean, just distance units... There are over forty of them