These are very common for web apps, because at the end of the day you don't care about the actual html & CSS, only how they are rendered.
> This is more of an integration test than a unit test.
That's debatable. An integration test generally tests 2 or more systems. This kind of test has 2 systems, the generator and the renderer, and we care about the output of the renderer, so it kind of looks like an integration test. However in an integration test you also have control over the implementation of both systems; a regression can be in any of the systems. But that's not true in snapshot tests: the renderer is a given. If the test fails, it's very unlikely to be due to a regression in the renderer. So in that sense, you are really only testing a single component (the generator) hence it is more like a unit test.
> because at the end of the day you don't care about the actual html & CSS, only how they are rendered
It really depends what you’re testing. I’ve generally been skeptical of this kind of test, similarly to OP, because “nothing changed” vs “update snapshots” feels intuitively low value to me.
Despite all that, I recently added a slew of snapshots (literally >1m lines, yikes) along with a custom snapshot serializer. For this use case I do care about the HTML (and XML) because those are the project’s primary responsibilities. The custom serialization slightly relaxes the snapshot value from “nothing changed” to accept known-insignificant changes: it collapses 1+N whitespace characters, sorts attributes alphabetically because their order doesn’t matter, and trims their values because no downstream users are that pedantic about leading/trailing attribute whitespace. Everything else will be treated as an API contract violation. There are some project-specific details which might result in additional custom serialization logic and creating new snapshots, where downstream users are expected to treat certain markup values as semantically equivalent to their equivalent representation in another output property.
Allllllll of that is a really long way around the barn to get to: these snapshot tests are more valuable than “exactly equal” comparison specifically because there’s a known and finite set of things that can fluctuate and a lot of caution around accepting anything into that category. And adding them at all, with known flexibilities, provides value because the underlying library has very high expectations for stability. It’s very unlikely they’ll ever be updated for any change which isn’t either additive or strategic.
(And the reason they were added in the first place was to allow for much needed performance improvements and refactors to proceed with high confidence that they’re safe. Since adding them, the project’s performance monitoring charts needed to run for a period of time to crop to a whole new Y axis range, and a refactor is in review which will enable it to run in client environments without need for server deployments. None of this would have been reasonable by my team’s standards or my own without a large body of evidence that it didn’t introduce regressions. And we very well may retire it after this exercise!)
> Traditional tests check individual properties (whitelists them), where characterization testing checks all properties that are not removed (blacklisted).
And I’d encourage anyone who finds snapshot testing appealing but problematic to consider this approach.
I have ca. 190 test cases on which I run my software and compare the md5 sums of the resulting PDF. If they are not the same, I create a PNG for every page and compare visually with imagemagick.
The trick is to remove all random stuff from the PDF (like ID generation or such).
This takes about 3 seconds on the M1 Pro laptop. I think this is very much okay.
Links: https://github.com/speedata/publisher/tree/develop/qa (the tests) https://github.com/speedata/publisher/blob/develop/src/go/sp... (the Go source code for the comparison)
An alternate approach: generate the PDF, then run it through a PDF reader library to scrape the text out and ensure it is there.
This was what I built once to do unit tests on a pdf generator. The use case was I was working for a software vendor at a very important financial client. Our software was used in their research group to produce output (massive reports) which fed into their trading decisions and had been heavily customized by the client in ways that our dev team couldn’t test (the client had a very exacting internal confidentiality regime and these reports and the various code customizations required to generate them were heavily restricted). We were trying to upgrade our software by 2 major versions and make sure everything was still going to run correctly in spite of the massive upgrade.
Each report contained >100 pages of 9 widgets per page and all the data going into these widgets was restricted as well as the calculations and outputs themselves. So we needed to somehow prove that the new system would generate these reports exactly the same as the previous system in spite of the fact that everything had changed under the hood and we couldn’t see either the old or new reports.
What we decided to do was treat the entire system as a big blackbox test. I built a junit harness that would generate the pdf of each of a set of reports using the old system and the new system and then diff the output. Initially we literally just used a normal text diff on the pdf file, but once we had fixed the first few (hundred) differences we refined it to snapshot both pdfs to images and produce the diff images because that made it easier to find and fix the problems.
It was a very painful process but an extremely effective test because the reports were the end product of the entire system and nailing all the differences in the reports proved that all the input data, all the calculated outputs and everything else was working correctly. It ended up with us doing the upgrade successfully.
Some people act like there's an obvious definition, and maybe there is if you're doing pure TDD Java as described in one specific text book... but in my experience most developers can't provide a good explanation of what a "unit" is.
And those that do... often write pretty awful tests! They mock almost everything and build tests that do very little to actually demonstrate that the system works as intended.
So I just call things "tests", and try to spend a lot more time on tests that exercise end-to-end functionality (which some people call "integration tests") than teats that operate against one single little function.
sometimes it's useful to know that something changed.