I recently tried it with "You are in a directory with a web app that does ... and you want to implement feature .... In each step, you can use ls or any other bash command and I will give you the output".
It was pretty hillarious how the LLM found its way around the codebase with ls, cat, find, grep awk and actually even managed to edit the code that way and do a commit.
Giving the LLM a bit better tools, like a version of "cat" that prefixes the lines with numbers, and a "swap" command that can do "swap 179 250 ..." to swap out the lines 179 to 250 with "..." would probably be enough to empower the LLM to be pretty efficient.
The next step might be to let the LLM manage its context window by allowing it to remove the last output with a command like "forget". So when the LLM does "cat somefile" and realizes that the output is not interesting, it can follow up with "forget" so the output will be replaced with "You deemed the output to be not interesting".
Those tools would probably evolve to make coding and managing the context window more and more efficient. Like "nicecat 100 200" to see the lines 100 to 200 with numbers prefixed. "keep 200 300" to forget the last output except lines 200 to 300 etc.