> individual 'hallucinations' can't be treated as bugs to troubleshoot
You are wrong here - my company can fix individual responses by adding specific targeted data for the RAG prompt. So a JIRA ticket for a wrong response can be fixed in 2 days.
You are wrong here - my company can fix individual responses by adding specific targeted data for the RAG prompt. So a JIRA ticket for a wrong response can be fixed in 2 days.
At scale, your solution looks like bolting an expert system on top of the LLM. Which is something that some researchers and companies are actually working on.
That would be as dangerous as any other function: you still need personnel verified as trustworthy in processing unreliable input.