> didn't say they did anything with a VLM to check the output vs the original library
Assuming VLM stands for vision–language model: We validated that all pixels are bit-identical across those 30M GIFs. So we didn't need a language model for that, we just compared expected and actual pixel buffers for strict equality. :)