AI should not behave as if it has inherent knowledge, there should be aknowledgement similar to, 'i saw on the web the other day'; or 'i was cURLing on coolrecipies.com and it gave me a new idea for halloween cake'
AI should not behave as if it has inherent knowledge, there should be aknowledgement similar to, 'i saw on the web the other day'; or 'i was cURLing on coolrecipies.com and it gave me a new idea for halloween cake'
"I saw on the web the other day" is not attribution, it's basically an interjection that makes the conversation smoother. Most of time it means absolutely nothing, except for maybe "it's not my direct experience" at best (and AFAIK LLMs today don't really have any agency, so this disclaimer is moot/noise).
Sure, "I've read an article on Acme Daily" happens, but personally I typically do this as a cue for the listener to cut me off with "ah, yes, I've read this too", saving us both time. Other use case is to give signal about authenticity of the information: not a credit, again - just an indirect indicator of trustworthiness or reliability (when I hear "The Onion reports" - it's surely not about the website, it's about the following being satire). YMMV, of course. I sure want an LLM to write this, but only when it matters to me personally (not the website authors, they aren't a party in our conversation). Just like a human would.
Similarly, I don't think anyone ever said "I've found this on coolrecipes.com" unironically if the conversation is about the recipe instead of a source. Obviously, attributing it to someone both parties know - like a relative, neighbor or a celebrity is a different story - roughly the same as with the news example above. But if you hear "found on coolrecipes.com" it it's most likely an ad. And what I want to say is here is that LLMs today are a breath of fresh air - compared to modern enshittified search engines - specifically because they're not ad- and SEO-ruined (yet). Let us please keep it this way for as long as possible.
ChatGPT is not a human. If a news site was to publish some information on their website without attribution, this would be a problem. ChatGPT is more analogous to a news site then some person you have a conversation with.
the idea is to enable the promptor to review the information, and see for themselves, should they have need.
I'm not being flippant here. We all use massive amounts of reference material to even think the most basic thoughts. Expecting you—or an algorithm—to be able to quote or even know the reference material that let to anything you say is a pretty high bar.
Just like in literature - tropes are standalone entities, and they evolve over time. While a trope can be attributed to some book or author, it evolves, gets mixed up, gets deconstructed and may end up with something barely recognizable. Should LLM be forced to always give a nod to Azimov or Čapek when talking about, let's say, Bender from Futurama - if this is not explicitly relevant to the conversation? I highly doubt so - it would make conversations intolerably stuffy. Talking about virtually any topic would end with a footnotes list multiple pages long.
Our (human) "normal" common sense to attribution is to ditch any and all, unless it's relevant for some reason. Because we generally try to stay focused. People or conversation machines attributing recipes to website addresses is a Black Mirror episode material, a corporate wet dream.