One related experiment, however, already suggests the result. Someone tried to get ChatGPT to solve advent of code challenges: https://github.com/golergka/advent-of-code-2022-with-chat-gp.... These challenges are very clear, and it already seems to struggle to get the answers. One the second day, it took 12 tries to get it right. If it struggles this much with clear requirements, I don't think it'll do well with vague ones.
Case in point, I got it to write a chrome extension to highlight new comments in Hacker News: https://github.com/HartS/gpt-hacker-news-extension
It's not perfect, it doesn't keep track of what the user actually saw for example. It just stores a timestamp for each opened thread and highlights the comments that are newer than that timestamp.
I wouldn't have thought to solve it that way, but I definitely think its successor will be capable of coming up with "good-enough" solutions for a lot of problems humans currently work on.
Actually, humans are often just coming up with "good-enough" solutions in the first place.
Regarding new comments: You can actually ask dang by mail for this exact feature. It already exists but still in beta or so. I have it and adore it. It works like this: When you go to a comment thread you've been before, it shows all new comments with a vertical red line to left side of the comment. Great feature. I already use it for years?! Could be rolled out actually... :D
Heh, I wonder when economic AIs will arrive
We're going to design a (language) application that performs the following tasks, delineated by semi-colons:
Then I listed about a dozen high-level user stories, separated by semi-colons, and ended the list with a period. Then I clarified my instructions like this:
We will design each step together, where you ask me for any specifications needed to complete the step and I respond with those specfications. When you feel we have completed designing a step, ask me if I have any questions before proceeding to the next step.
It started with a high-level recap of my user stories, asked if I was ready to proceed, and then began to describe the first step. Over about two hours of back-and-forth, I ended up with a 20-page document of high-level architecture interspersed with code samples, db schema samples, and answers to questions that popped up during the "discussion." And I got about halfway to a working prototype during those two hours.
BLEW me away how well it worked.