Having it with clear hw requirements tiers is a nice differentiator. The only issue is that the benchmarks would 100% need to be closed, no other way around it. And then you have the issue of creating and curating good evals for every "stage" of the project. That's a hard task even for "honest" lab-internal evals. And you'd have to publish those evals after each round (for trust purposes), and start over for the next round. Doable, but it would cost a lot (probably more than the prize pools) and you'd have to keep doing this.