However:
- the LLM gets basic facts wrong about the solution
- the LLM fails you for not solving issues that you weren't supposed to solve (it's a live outage scenario where you're supposed to restore the system ASAP, not go on side quests for bugs that aren't being triggered)
Looks like the guy is too amazed by one-shot evals and believes they should be used for humans too, but then his LLM makes basic mistakes evaluating the results (such as: did you or did you not commit to a git repository).
When emailed about these issues, first he ignored the email, then when emailed about it again he got really combative and started gaslighting about how well his AI driven job interview process works. This shows that he's not a guy who's capable of dealing with the least amount of disagreement. Definitely not someone I'd want deciding about my continued employment.
This is the full email he sent once he finally responded:
> That kind of attitude doesn't help. We have a business to run; we can't be answering email all day long.
>
> Our harness has survived 500+ tests, I doubt it is incorrect. However, there were three bugs; you only found one.
(I wasn't giving him attitude.)
500 garbage AI generated tests are still just garbage my dude.
Avoid at all costs.