But can an AI construct Terminal Bench 5.0 or GDPval 2027 ?
Without this, there is just no path towards RSI.
But can an AI construct Terminal Bench 5.0 or GDPval 2027 ?
Without this, there is just no path towards RSI.
Doesn't everyone get their agents to construct evals it can't pass? There's nothing magical about this.
For example if frontier models are able to one-shot a database query across 20 columns and 10 tables add one additional relationship then test. Keep doing this until the pass-rate drops below acceptable and now you have your new frontier eval.
My thought experiment was along the lines of "Let's say I'm Anthropic and I want to significantly improve my frontier model's performance on, say, theoretical physics research. How do I build a fully autonomous process capable of constructing an eval that's somewhat outside the current capability in some useful direction (decided by the autonomous process itself)?"
Would love to hear folks' ideas. :)