Scaling Auto-Research Loops for Efficient Agent Harnesses
nvlabs.github.io
nvlabs.github.io
Not sure how useful it may be to me, in my daily use. My tasks are way more encapsulated (e.g. edit the CSS of this website for me) than the "long horizon" that Edge-Bench aims to represent (e.g. working on a math theorem)