For this particular use case: We previously used VMs and snapshots for that workload. The problems we encountered were:
- snapshots aren't really intended to be portable (we used ESXi and also KVM on LVM backed volumes), hence we had to write and maintain tooling to have a "snapshot repository", versioning for these and distributing them to the target nodes - that did all work, but was even slower (startup times, which can include shipping and setting up the snapshot, was 15 - 60 minutes)
- a VM will duplicate all of the services and the kernel, so all of that has to be started as well, where as the containers only start the services under test.
- Using VMs makes it much more difficult for our developers and QA to retreive a particular image and replicate a failing test on a given version of the tested software locally, especially when considering a wide range of client OSs we see in our not that large group (MacOS, various Linux distros, even Windows)
A general observation on trade-offs: With containers layering and COW you get fast development, but have to pay the performance bill when you happen to download and apply an image for the first time. Similarly, taking a snapshot of a LV under LVM or on a ESXi VM is fast, but applying the snapshot is slow.
We had therefore at one point considered using Ceph RDB and their COW clone snapshots[0]. It would let us do "cheap" restores of snapshots. Our initial tests showed that the network bandwidth requirements[1] would have needed some serious infrastructure re-engineering in order to keep up with our I/O expectations. And again, the containers slot in nicely with commonly available local resources and allow working offline, to a degree.
[0] https://docs.ceph.com/en/latest/rbd/rbd-snapshot/
[1] the environment in question uses 10 - 40 Gb/s network links, 40 on the gitlab git and container registry side, 10 on the blade side.