Here is what HN was talking about, nearly three months ago -the exact same type of ‘auto-prompt-gen’ tool: https://news.ycombinator.com/item?id=35660751
I was reminded of the same thing. What a lot of it boils down to is that LLMs have no innate ability to self-reflect. They can pretend to do it, but no more effectively than an untrained human would.
Which is exactly as much as Generative AI should be trusted.
The problem is that it will simultaneously say that "cow eggs are bigger than chicken eggs", with the same confidence (and in a way that correlates well with human evaluators).
https://www.reddit.com/r/Funnymemes/comments/10ohd2n/chatgpt...
So when you get an evaluation you are playing the russian roulette - you may get a decent result, or you may get cow eggs.
The point is that the tool fails, and it is known to fail, so much that we even have a name for the times when it fails - hallucinations. I have been calling them cow eggs because that's a nice mental image and I didn't want to have to remember for the proper English term. I will continue calling them cow eggs.
If you're prepared to accept that GPT-4 can answer questions just as well as humans can, why do you even need to do prompt engineering?
* What's the best way to get to Radio Shack from here?
is not the same as
* What's the easiest way to get to Radio Shack from memory when riding a bicycle from here?
You may be under the impression that annual U.S. deaths from medical errors being in the hundreds of thousands miscommunicates but that is truly your opinion. You are merely jumping to conclusions at places another person might not.
And going on to rely on the LLM to validate your perspective is a lossy process. It may not lose your perspective but it loses someone else's and you don't even seem to notice or care.
The post you replied to was saying that the deaths were caused by miscommunication, but you interpreted it to mean that stating the number of such deaths is somehow a miscommunication itself!
You can't use the thing you're testing to evaluate its own performance. This applies to rulers, speedometers, and AI. It's the difference between a "subjective" and "objective" metrics. If you want an objective metric, you need to have it based on something external, based on reality, objective. Otherwise, you have metrics and ideas that have to held themselves up.
Source: My day job is test and measurement. These concepts go back centuries. You never trust your measurement system, you verify it against a standard.
Yeah I agree there. Unless you can check against the output, it’s not really telling much.
Could you post the link please
A positive correlation just means better than chance. "Strongly" is vague, and might not be much better than chance.
No, adverbs like “strongly” modify adjectives (or verbs, but that’s not relevant here) not nouns; “strongly” is an intensifier that modifies “positive”, its not a separate adjective that modifies the noun “correlation”.