The private evaluation set is private from the public/OpenAI so companies can't train on those problems and cheat their way to a high score by overfitting.
If the models run on OpenAIs servers then surely they could still see the questions being put into it if they wanted to cheat? That could only be prevented by making the evaluation a one-time deal that can't be repeated, or by having OpenAI distribute their models for evaluators to run themselves, which I doubt they're inclined to do.
Yes that's why it is "semi"-private: From the ARC website "This set is "semi-private" because we can assume that over time, this data will be added to LLM training data and need to be periodically updated."
I presume evaluation on the test set is gated (you have to ask ARC to run it).