This is something that should have been caught by automated tests.
This is something that should have been caught by automated tests.
There’s basically no automated test suite even for CSS 2.1 that all vendors can rely on that matches behavior specified by the standards.
Luckily standards like that are so mature that it’s not needed as badly, but it definitely keeps incumbents safe from competition.
But for others who don't know:
https://web-platform-tests.org/
https://github.com/web-platform-tests/wpt/tree/master/css/CS...
They’re also tests that don’t actually point anything out.
Take the CSS2 tests for example, https://wpt.live/css/CSS2/, I can’t do basically anything with these. They don’t tell me whether or not a particular part of the implementation has failed. They’re all basically arbitrary tests that sort of cover the appropriate properties, sometimes, maybe.
The biggest glaring issue is that the test runners run in the browser. Well, if my user agent doesn’t have JavaScript… I can’t “run” the tests.
Further, there are some parts of the standard that can’t be automated simply because it’s up to user agents to decide how the final rasterization turns out. You can’t just compare rasterization output or even reliably look for element heights for verification based on the differences between font rasterization.
This is also why a part of the suite requires manual verification.
What I really need as an implementor is the ability to point my software at some .html files, get some failures, add more implementation, get less failures, and repeat. With, of course, the caveat being that you can only really do this for what is well-defined in the standards.
But you can’t do that easily with the wpts.
I'm not really sure I follow your complaint, so I wonder if there's a point of misunderstanding. The tests written for DOM APIs certainly require running javascript to execute the test — that's a given — but most of the rendering tests are reftests. These consist of two files, one using the feature under test and one avoiding that feature [1]. The test passes when the two renderings are identical. That does make these tests difficult to run "by hand", but in terms of implementation it's possible to automate in anything that's capable of rendering to an image. Reftests are preferred over comparing rendering to a fixed image because fixed images tend to be invalidated due to unrelated/permitted changes in the rendering (e.g. a different font or different antialiasing choices). These will typically apply to both test and reference and so are handled automatically.
Reftests obviously aren't suitable to bootstrap getting very basic things working, so there are some manual tests; I suppose one could hope to avoid that by inventing a special kind of output inspection automation just for those basic tests, but it's hard to justify.
IN general, however, the "the ability to point my software at some .html files, get some failures, add more implementation, get less failures, and repeat" is exactly what wpt offers. There is some work required to stand up the necessary infrastructure to run the tests automatically, but it's common for new implementations of the platform to make use of web-platform-tests (e.g. Servo and Flow have both made use of wpt).
In terms of coverage it's very difficult to demonstrate that all the requirements of a spec, and all its interactions with other specs, are completely covered. There have been various proposals for adding test metadata to try to measure coverage, but these have largely proved impractical. Nevertheless some CSS specs have links from the spec to the relevant test cases. And if you find cases that aren't covered by exisitng tests it would be great to get a PR to add new tests [2]
[1] https://web-platform-tests.org/writing-tests/reftests.html [2] https://github.com/web-platform-tests/wpt
It is unacceptable to me to use the software I’m testing to test itself if I can’t establish a baseline for the rest of the tests.
Otherwise, all of the tests could be invalid one day based on a regression and I wouldn’t know. But this is explicitly done with reftests.
And I don’t know what you’re talking about with the DOM API tests, there are several layout and compositor tests that require JavaScript and it’s unclear why other than the test author decided to use JavaScript to verify DOM metrics. That’s fine, I just don’t want that in my layout tests.
I’m OK with a subset of the tests just not being automated based on font rasterization and line height calculations. The standards make those expectations clear, but everything else feels like fair game to me.
I can either test for it, or I can’t. If I can’t, it’s not automated. And if it’s automated, I need information about what is failing, and much of the reftests don’t explicitly point out what they’re testing. The assert data is useful when it’s available, but even then sometimes test authors neglect to point out explicitly what they’re testing for.
Plenty of the layout tests in particular are just box model renders that should match with no explanation as to what section or text is being tested against. That’s just not acceptable. It’s not even about interactions between specs, the specs isolated are not well tested.
If you’re testing, let’s say, the width calculation of a box based on children boxes’ widths in normal flow, it needs to say that. And I think I’ve only ever see some of WebKit’s test’s help links reference actual sections.
Otherwise the WPT provide little to no value to me.
My team has been better off writing our own tests because we directly refer to the recommended or living standards texts by section, quote it and link to it, and as a result there’s no ambiguity about what property, calculation, or raster output we’re testing for.
Because a cursory glance at Chromium src tells me they do have automated tests for IndexedDB:
https://source.chromium.org/chromium/chromium/src/+/main:con...
However, if you do stumble on some tests for basic things like the box model, I'd love to know where they are! Tests like those would allow other vendors to know whether or not they're adhering to specification, programmatically.
Earlier standards are very descriptive, rather than telling you how to implement the actual web technologies. That's left as an exercise for the reader. Seriously!
I guess applications living in the browser don't align well with Apple's interest in pushing their app store.
There actually are some partial tests from the W3C regarding basic web technologies, but they're not really automated ones in the way you'd think.
I only mean to say that you don't need a whole compliance testing suite to catch a major bug like this.
Its not really a test suite, but back in the early days of CSS we had https://en.m.wikipedia.org/wiki/Acid2 It'd be easy to write a set of tests based on Acid2.
There's also https://en.m.wikipedia.org/wiki/Acid3 but that doesn't quite test things 'properly'.
In addition the specs have changed somewhat so now CSS 2.1 and the modules that existed at that point in time have matured beyond what the tests were written for, so a browser matching the spec today would fail.
They were very fun for end users to see whether or not a browser was in compliance, though!
Such an effort would probably involve a higher level of ongoing support from the vendor side of the fence (in terms of motivation and iterative discussion etc), but relative to other options might be, as I said, the path of least resistance.
</armchair_idea>
https://github.com/WebKit/WebKit/tree/main/LayoutTests/stora...
Testing browsers is a daunting task. A modern browser is 10-15 million lines of code with literally thousands of exposed APIs [2]
[1] https://www.pcgamer.com/a-google-chrome-update-breaks-the-au...
My favorite one was websites added to the home screen on an iPad would have the clock displayed above the websites if you rotate the iPad after opening a website. It wasn't fixed for years and it may still be there.
As with every race condition (cross process, too, so tools like TSAN don't help), testing for the absence of race conditions is hard.