Good data and good evals are two legs of the 3-legged stool that a lot of AI teams are missing.
Ok the sarcasm got too thick but my point is if the engineer has to spend the time to comb thousands of examples then you don't have AI you have a man in a box pretending to be a machine that plays chess.
Are humans just other humans hiding in boxes pretending to play chess?
I’ve resorted to building my own annotation apps.