One thing that evals are super important from the get go are where the harness+model inference is part of the product, e.g. if you are doing voice ai, building out a test harness to test the system is a non trivial first step.
I use pi and built a harness for just an llm that calls a bunch of tools. That got me 50% of the way and it would be fast. Then build it for STT and TTS, this will be slower but it will get you far. There are a bunch of tools out there for building basic harnesses.