Ah! Think of this more like software testing that goes in CI/CD rather than an ML test or validation set. We're providing this testing for applications built on top of language models.
For example if you're a SWE working on bing chat, you can make a change to how retrieval works and quickly know how it affected accuracy on a range of different test scenarios. This kind of evaluation is done by contractors today, and they are slow and inaccurate.