> Recent research and anecdotal experience has shown that LLMs perform quite poorly with short prompts.
I'm aware of that. The actual prompt was more elaborate. I was just mentioning the gist of it here.
Besides, you would think that after 30 minutes of prompting and corrections it would arrive at the correct answer. I'm aware that subsequent output is based on the session history, but I would also expect this to be less of an issue if the human response was negative. It just seems like sloppy engineering otherwise.
> Specific is the specific thing that statistical models are not good at
Some models are good at needle-in-a-haystack problems. If the information exists, they're able to find it. What I don't need is for it to hallucinate wrong answers if the information doesn't exist.
This is a core problem of this tech, but I also expected it to improve over time.
> Tho you should give aider.chat a try
Thanks, I'll do that eventually. If it's slow, it can get faster. I'd rather the tool be slow but give correct answers, than it slowing me down by wasting my time error correcting it.
Thankfully, these approaches can work for programming tasks. There is not much that can be done to verify the output of any other subject.