Yes but don’t worry there’s another YC startup who is building AI to check on AI that’s checking on AI.
Gold
Like I said above, we don’t use AI agents to grade other models. Instead, we run in-house evaluations tailored to each category of clinical AI, giving hospitals an apples-to-apples comparison between similar vendors.
How are you going to protect AI which optimise against your tests instead of actual data ?