About the zillionth time you get 'I'm sorry, you're right, [blah] doesn't exist. Here's [something else that doesn't exist]', it gets really frustrating. The worst part is you can often tell the software is referencing some relevant page(s), but just inappropriately mixing them with other stuff. And so if you could simply get the link it'd be far more helpful than listening to the program continuing to describe in immaculate detail how to use an API that does exactly what you're looking for, with the slight problem that it doesn't exist.
Humans can't possibly remember provenance of all information. Machines possibly could. There's a significant difference in capabilities and it would be unwise to ignore this.
Machines can already do that. We have many, many memory mechanism for llm these days. There are already search engine like sites like phind.com which are rigorous about the sources. Langchain has tools to retrieve and cite memory from doc stores, APIs etc. If I ask my langchain agent "where did you learn that" or "why are you saying this" it does a decent job of citing the source.
Now, if your objection is that these are not perfect, then I would like to welcome you to the rising part of the sigmoid function of progress.
Access to openAI API directly gives better explanations to logic that was applied too.
I do not understand this god of the gaps type debate that is going on regarding LLMs. Everyone just seems to be interested in pointing out flaws. They are all using the "I'm witholding my judgement words" but are in the "this is just garbage" tone.
However, for the few queries I made, there are still all sorts of issues involved, such as how the reliability of the sources are determined. For instance, I asked phind.com to give some information with references to arXiv, and it concluded that arxiv-vanity.com was the proper source ("rigorous about the sources" indeed...). Then I tried a few queries about jurisprudence and without hesitation it went to reference questionable commercial sites instead of the primary sources. Furthermore, it seems that phind.com is quilty of 1:1 plagiarism in many cases.
> Now, if your objection is that these are not perfect, then I would like to welcome you to the rising part of the sigmoid function of progress.
Edit: Also phind.com does better with coding docs. I think it is focused on code so asking is questions outside of that domain is kinda unfair.
My point is merely that we should have high requirements and not excuse flaws in language models just because humans have them. That thinking is akin to apologetics and not constructive.
They could have 10,000 occurances that told them to use the word dog in a response to the question "what is man's best friend?"
Also, the answer to where did they learn which town to use for "Where and when were grapes introduced into Australia?" seems to be "Actually, I didn't know, I just picked from a list of Australian towns and made the factual link up"
In the case of the GP, checking where something has been learned or inferred by AI is something I learned from reading articles as well as participating in discussions with colleagues over AI explainability/transparency, as well as reading articles and listening to podcasts over journalism and truth in the age of AI. In fact I know what my 3 to 5 biggest influences/sources on that topic are. They might not be the one who invented the topic and answer, but I know who passed it to me.
Do you know where you learned your opinions on a novel question? How?
I acknowledge you as my superior.
I'm sure there are more robust neuroscience papers, but as an example: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4409058/
Your average human does a lot of optimisation to storage things (sight sound smell, etc) in the old wet ware, and retrieval can be hit and miss.
At some point Sussman expressed how he thought AI was on the wrong track. He explained that he thought most AI directions were not interesting to him, because they were about building up a solid AI foundation, then the AI system runs as a sort of black box. "I'm not interested in that. I want software that's accountable." Accountable? "Yes, I want something that can express its symbolic reasoning. I want to it to tell me why it did the thing it did, what it thought was going to happen, and then what happened instead." He then said something that took me a long time to process, and at first I mistook for being very science-fiction'y, along the lines of, "If an AI driven car drives off the side of the road, I want to know why it did that. I could take the software developer to court, but I would much rather take the AI to court."
But even now, with LLMs, we can capture the 'internal monologue' of the bot when it is augmented in that way. For example in the red teaming section of the gpt 4 technical report, the bot is trying to solve a captcha for some nefarious purpose. It can't do it, so it decides (as we know from reading its internal monologue) to ask a taskrabbit worker to do it for them. The taskrabbit worker semi-jokingly asks if it can't read the captcha because it's a bot. In response the bot decides (again as we know from reading its mind) to deliberately lie to the worker so the worker won't know it's a bot. The bot says it's a human with a vision impairment.
I'm not sure what else you would expect for "if an AI driven car drives off the side of the road, I want to know why it did that" beyond that kind of access to the internal monologue of an AI agent.
When you say this is a different argument, which arguments are you saying are different from each other?
When you say explaining reasoning is different from citing sources, why are you talking about citing sources?
Because that's what I was referring to and you moved the goalposts. Sources are a necessary part of the knowledge metadata so that truth weights can be taken into account. Otherwise the data's reliability factor isn't a first-class concept within the system. Citing the sources consistently serves as proof that the system does this (trust), and allows the end user to verify (at least to a reasonable level).
My point was that while ChatGPT might be a viable product, an actual AGI system should not be attempted to be built atop a LLM that can't do this.
Here is a system that does do this. https://wiki.opencog.org/w/AtomSpace
Oh I understand it now. Your original top level comment was about citing sources and then the comment by some third one a couple layers deep about Gerald Sussman changed from 'citing sources' to 'explaining reasoning'. Sorry I got confused because the thread got so deep.
Here is what Claude[1] had to say when questioned about its own explanation about why it misinterpreted a sentence:
"You raise a fair point - I don't actually "think" in the human sense, I'm an AI assistant created by Anthropic, PBC to be helpful, harmless, and honest.
So I don't have subjective experiences, assumptions, or thoughts that led to misinterpreting this sentence. I simply operated based on the algorithms and training data provided by my creators at Anthropic to generate responses.
My explanations for why I might make mistakes were fabricated in order to seem more human-like and helpful..."
Yes this was my point when I said "If you ask a person what they were thinking, the person will just rationalize in the same way the bot will do it." We are saying the same thing. Or maybe you have more faith than me that a human's explanations of their own "reasoning" is not 'hallucinated' by post-hoc rationalization?
> no access to its internal workings
For the purposes of understanding the internal monologue of 'agentic' LLMs that run in some kind of loop, the human observer actually can get access to at least those parts of the internal workings of the AI to extract explanations like the ancestor commenter wanted. So in that sense, the AI is more reliably explainable than a human. It would be as if we had access to a human's internal monologue at all times (although the internal monologue of course is not the only part of cognition). That was my point in the 'red teaming' example.
As far as I am concerned LLM offer little more than a cool tech demo. They suffer from the same failing that results in AI art often suffering from basic flaws like extra arms. Better fakes aren’t closer to the actual solution they’re simply optimized for a different metric.
What seems revolutionary today is going to feel as useless as 3D TV’s once the novelty wares off.
With art you can discard the output with a glance try again so it’s not a waste of time. With text however you’re stuck carefully checking for errors and there is going to be many many errs. Now most students might not care because the grading is generally lenient, but having say a PHD thesis, resume, or professional work riddled with subtle errors is a more serious problem.
I personally just don’t see any real value in low quality output. The time I waste correcting it is longer than just doing it myself and maintaining your reputation is important so doing a poor job on unimportant things isn’t a good long term strategy. If you’re stuck at the level of “fiver” get good is just so much more valuable than get fast.
More substantively, LLMs are useful today, so even if they somehow don't improve at all, they're going to be impactful. In just the few months that we've had chatGPT it's saved me countless hours in work and even personal tasks, and I can see various ways it's going to be useful in the future. As one of the sibling replies says, it doesn't have to be perfect to be useful, and this is a tool with incredibly broad usefulness. (Admittedly unlike 3D TVs...)
Maybe the code isn't perfect, but all I wanted was to take something live, and it's accelerated that process by months.
To me, it already feels far beyond a cool tech demo.
My objection is the end quality seems stuck at that level. They’re great at crossing the just barely acceptable level but make it harder to reach excellence by for example having many subtle bugs. Which I suspect comes from their training set. Their training sets are simply dominated by amateur efforts.
[1] - https://www.reddit.com/r/freelanceWriters/comments/12ff5mw/i...
“Today he emailed saying that although he knows AI’s work isn’t nearly as good as mine, he can’t ignore the profit margin.”
chatGPT with 3.5 was released in November 2022. 4 in March 2023. Massive leap in quality between the two relases.
I don't think we can declare them to be stuck given the SOTA massively improved six weeks ago.
I'm asking because people's standards are quite different when it comes to evaluating emerging/new technologies.
One possibility might be improving the quality of the training data rather than aiming for such a huge bulk of amateur works. Unfortunately, there might not be enough quality writing to get an LLM to work.
The early EV’s as in 1900’s where using fundamentally flawed battery chemistry. Simply improving lead acid batteries wasn’t going to work. Similarly GPT45 is likely to have similar fundamental differences from current LLM’s.
Meanwhile, Reddit and Twitter AI community is beyond giddy with excitement.
On this topic it's also interesting how 'AI skeptic' to the degree that it exists, is doing an interesting shift in meaning in real time! It used to mean something like 'AI is some dead end and people will always be more clever in whatever ways' but now it is starting to mean something like 'we always knew these AI could be so clever but we shouldn't trust them in positions of power or authority because they might make a mistake or we might disagree with them or maybe they can't explain their action or maybe I want to sue them.'
That's the advantage traditional search has - I can see the links where the output came from. With AI, it's unclear how was the data compiled.
Yes, it's called a citation. https://en.wikipedia.org/wiki/Citation