I think the crates trumping other complexity metrics isn't entirely obvious to me. For problem 15 in the OP's post, author says it was too expensive to compute at runtime in the browser. From a human perspective, it's not apparent why, as a large part of the solution is very repetitive. It feels as if there should be a more condensed representation for iterating over problems like that one.
If I may gauge your opinion on it, have you looked into MazeBench? It comes from LLM benchmarking circles, but seems to suggest a search space that's too difficult for LLMs, even with tools, to solve. Curious how much overlap the PS/MIS solvers would have with solving something like this.