Maybe I'm a control freak, but asking agents to one-shot random apps is nothing like how I actually use AI in software engineering.
surprised it isn't a bigger thing, eg artificial analysis doesn't report anything like that
still doesn't measure the human-agent interaction part, but that's pure vibes atp
I suppose it’s interesting to see how they make better greenfield apps. But I am much more interested in how they solve hard problems in existing gnarly codebases.