> the author may be misunderstanding what practitioners mean
I'm a "practitioner", well a researcher. Most of my peers that I talk to believe evals/benchmarks are a strong indicator of real world performance as well as much more abstract notions such as ability to think or intelligence or whatever those mean. But benchmarks and evals are fairly narrow.
> They do not _guarantee_ software performance in the same way that unit tests do for traditional software.
Unit tests do not guarantee code correctness. TDD is a flawed paradigm. Tests are good, but they shouldn't be the driver nor the measure. The fatal flaw in TDD is it is entirely dependent on your ability to foresee all possible failure points as well as there being an absence of black swans. It's success is highly dependent on the person implementing it, not the method itself. The common hubris of "this should never happen"
> I am not aware of any workable alternative.
The great problem in ML that we're currently facing is that all measurements (of ANY kind) are proxies to the thing you wish to measure. It is easy to forget this, to believe that even with a ruler in front of you that you are measuring meters (inches, whatever unit you want to pick). You are instead measuring increments of your measuring tape, which is hopefully well aligned with an actual meter. This is almost always "good enough" because the misalignment is less than uncertainty of the measuring device itself. The problem is, when you get to measuring much more abstract things, it is harder to know how well you are aligned. We forget to even question this notion until long after it becomes a problem.