Skepticism warranted since ARC-AGI hasn't been reproduced and there's at least some debate about validation data leaking into the training set: https://github.com/sapientinc/HRM/issues/18
I'm eagerly awaiting reproduction. I hope the results hold up and someone finds a way to bolt language onto it.