ARC-AGI 2 went the same way because it was basically the same kind of dataset except this time with some attempt to further defend it against LLMs with restrictions on the compute budget. And now ARC-AGI 3 is saturated within ... what is it, weeks? since its release. The fact that it's the public set that's beaten doesn't matter, when the score is 99%. Systems that can score ~90% on the public sets of the previous ARC's can comfortably reach 70-80% on the corresponding private test sets, as far as my eyballing of results suggests.
It is time to accept that the whole idea of ARC is for the dustbin. It does not measure what it's supposed to measure -fluid intelligence, reasoning, whatever it is today. Its original premise, that a system could only beat ARC if it possessed human-like core knowledge systems (a-la Elizabeth Spelke's theory) has been comprehensively refuted: none of the systems that have ever performed well on any version of ARC has made any attempt to represent core knowledge systems in any way, shape or form.
Ultimately, if your machine intelligence (let alone AGI) test relies on tricks like only giving a few examples or keeping a secret test set to defend itself against the dominant approach to machine intelligence... then it's not a useful machine intelligence test. Or it just doesn't measure machine intelligence but... something else. Who knows what.