because that's how agents are marketed.
because that's how agents are marketed.
heaps of people on this site expect them to be omnipotent then claim it’s fake when it doesn’t read minds
I don't care what the company claims, I just use the tool the way I want to. I work very closely with the AI. I'll tell it to plan a change, review the plan, then execute. Then I'll test the changes and have it fix whatever I'm not happy with one thing at a time. I'll specify in detail both what to do and loosely describe how to do it or if I'm not sure I'll ask it to plan the change then review the plan and ask for changes if I want them etc. I also review my own PRs before I submit them to colleagues.
This way I maintain full control of everything, it just saves me hours of googling, planning and typing code - which I do miss a bit but I can't really justify writing code myself when I can achieve the same thing just by loosely describing my idea instead. It also saves a lot of time debugging, I think I'm generally a pretty good programmer but the AI makes fewer mistakes than me. It'll often catch some logic error I made during planning and suggest a good alternative.
A lot of developers seem to give up control entirely and then complain that they're no longer in control. Trying for that 10-100x speedup doing weeks of work in a day. I'm happy doing one week of work in a day. There's a limit to how much I can oversee without compromising quality.
It might do Y, but Y does not excuse so high valuations and investments, so they claim X.
It is ok to judge them by !X.
- https://openai.com/index/introducing-the-codex-app/
- https://www.anthropic.com/news/claude-3-7-sonnet
anthropic specifically brags about how good claude code is every annoucement of a new model. I will surrender that none of them claim its "to perfection", but IMO its implied because no one would claim that their model one-shots any issue to dog shit quality.
no where does this document suggest that codex can "one shot everything to perfection with just a prompt". It describes using a prompt plus agent skills (which are essentially many other prompts) to develop a playable game.. nothing about it being perfect or anything more than being in a playable state.
Besides, other people's claims about something doesn't give you license to abandon all critical thinking. Though it's evident they don't claim what you say they are.
To be clear that's not what I'm thinking, even Fable 5 produces some hilariously bad results under some conditions and sonnet 5 produced great results under others.
But you didn't provide the evidence for that. You shared some links and then admitted they didn't claim it.
It kinda seems like "because I think they're a little too positive about their product, I can set my expectations to anything I want and la-la-la it's their fault."
And I don't see the problem with agents building test scaffolding as they go. It might be too defensive at times, like testing a shell script you don't run often, but big deal. It's kinda cool imo, and it's trivial to make it stop.