Today there might be people who can't extract enough juice from LLMs, so it is not entirely useless to say "I was able to extract this info from a LLM, because I am good at it and you seem to struggle", instead of throwing "just ask Claude!".
Today there might be people who can't extract enough juice from LLMs, so it is not entirely useless to say "I was able to extract this info from a LLM, because I am good at it and you seem to struggle", instead of throwing "just ask Claude!".
IMO they were right, I changed my approach to those kind of questions, and since that I try to answer like "A quick search in Google says that the CEO is Mark Zuckerberg (link to the search)". In StackOverflow I tried to go "As it says in the <a href='manual.html#section'>manual</a>, the params for that function are A, B and C, blah, blah...", so a mild RTFM. And now I do the same quoting the LLM paragraph that gave me the key info. It is like you say "this is how you can figure it out on your own the next time", and feels less aggressive than "go figure that on your own".
At least with "I don't know" the asker can move on to someone who might know faster.
Different reward function, but the same behaviour emerges.
The idea is that you generate fake llm transcripts using your classical training data. E.g. look at some training data, generate q/a transcripts. Generate radom questions, RAG against your whole dataset and look for relevant stuff, if there is nothing there, train a "I don't know." reply.
A moderately sized LLM operating some tools to access more information behind the scenes, perform tests and correct its own errors can write transcripts simulating a much larger and smarter llm.