I agree with that Evals are a scarce commodity rn. A business needs to define what good looks like. This is a laborious, and sometimes politically controversial, process.
Given good Evals, frontier llms are a magical tool that can automate tasks and do them more accurately than human ops teams.
Exactly which llm to use depends on the mixture of speed, cost and quality of the output.
Jev makes a claim to expand some regions of the Pareto frontier. I look forward to testing if this is true.
There are many areas of work we can’t automate rn. We cannot create good Evals either because time horizons are too long, or it’s too difficult to create good Evals.
That doesn’t mean there’s anything wrong with building good ai engineering systems in areas where it works magically.