I’ve validated Jev’s confidence score. Accuracy scales linearly with confidence for the 3 use cases I tested. >.9 it matched a human labeler. I immediately discovered a user behavior I didn’t expect for ~$3. I can now mitigate in real-time due to low cost and latency. This may have a major positive financial impact for all our customers.
Not sure why anyone would feel the need to dunk on this thing without showing a real failure example.