Before you replace, can I show you how the LLM integration works on HackerRank?
For additional context, we are currently using problems which have test cases, and it feels as if it doesn't correctly reflect the day-to-day of a data scientist. In fact, many interviewers pass a candidate iff they pass the test cases.
I'd love more open-ended problems.