Tim Cook is 'not 100 percent' sure Apple can stop AI hallucinations
theverge.com
theverge.com
I'm not claiming it's perfect, and like Tim Cook I'm not 100% sure it's actually possible to get to a 0% error rate, particularly considering it's trained on data written by humans on the internet, and humans write a lot of really dumb shit on the internet all the time, and human curation can't possibly stop all of it. It's easy to tell the bot to not scrape obviously sketchy sites like Infowars or something, and maybe block some specific subreddits, but there's still always going to be very dumb people posting very dumb opinions [2] in the "mainstream" areas as well.
It's not weird to see really stupid stuff on actual news websites like CNN talking about "ghost sightings", and I'm not sure how you correct for that in your training. You could block CNN, but that's a pretty big repository of news that you're losing training on, and moreover how do you actually block a news website impartially?
[1] Not to undermine the cool stuff that's being done in the space, just my rough understanding of how the algo actually works.
[2] My HN history is probably included in this.
That’s not really an LLM providing references, but a separate db as an extra step providing the references.
If I take llama3, for example, and ask it to provide a reference.. it will just make something up. Sometimes these things happen to exist, often times they don’t. And that’s the fundamental problem - they hallucinate. This is well understood.
Rather, such AI systems have two parts: an LLM that writes whatever, and a separate system that searches for links related to what the LLM wrote.
I can think of a counter-point:
Humans can deal with 'error-cases' pretty readily. The effort of humans to get those last few 9's might be a linear effort. For example, one extra checklist might be the thing that was needed, or simply adding a whole extra human to avoid a problem. OTOH, computers to get that last 0.001% correct might need magnitudes more effort to get right. The effort of humans vs computers does not scale at the same rate. Why therefore should we think that something that humans can do with 2x effort would not require 2000x better AI?
There are certainly cases where the inverse would be true. Where human effort to get better reliability would be more than what is needed for a computer (monitoring and control systems are good examples; eg: radar operators, nuclear power plants). Though, in cases where computer effort scales better than human effort - it's very likely those efforts have been automated already. That high level of reliability in aviation is likely thanks in part due to automation of tasks that computers are good at.
Even if it does, that puts AI ahead in what, 22 years with 2x improvements every 2 years? The simple problem with us humans, we haven't notably improved in inteligence for the last 100.000 years and we'll be beat eventually. It's not even a question barring some end of the world event, we already know it's completely possible because we are living proof that 20 W can be at least this smart.
And it's really just that the upfront R&D is expensive, silicon is cheap. Humans are ultra expensive all round constantly, so the longer you use the result, the more that initial seemingly ludicrous cost amortizes to near zero.
First, is AI actually improving 2x every 2 years? Can you link to some studies that show the benchmarks? AFAIK OpenAI with ChatGPT was something of an 8 or 10 year project. It being released seemingly "out of nowhere" really biases the perception that things are happening faster than they actually are.
Second, is human intelligence and AI even in the same domain? How many dimensions of intelligence are there? Out of those dimensions, which ones have AI at a zero so far? Which of those dimensions are even entirely impossible for AI? If the answer is "AI" can be just as smart as humans in every way, and we still don't even understand that much about human intelligence and cognition, let alone that of other animals... I'm skeptical the answer is yes. (Thinking about sciences view of actual intelligence, animals, and for a long time the thought was animals are biological automotons, I think shows that we don't even understand intelligence, let alone how to build AGI).
Next, even if the intelligence raise is single dimension and is actually the same sport & playing field, what is to say that the exponential growth you describe will be consistent? Could it not be the case that the 1000x to 1001x improvement might be just as hard as all of the first 1000x improvement? What is to say the complexity increase is not also exponential, or even a combinatoric growth?
> 20 W can be at least this smart.
20 W? I'm not familiar with it.
For example, humans can prove that 5 times 8 is 40 in a variety of ways. While you might be wrong in arithmetic, you can check your answer. AI can't check its answer, it does not know when it is wrong (it picked an answer it 'thought' was right, ergo it has no ability to consider that as a wrong answer, otherwise it would have chosen a different answer).
Why do we hold AI to a higher standard than humans? It's the same with self driving cars. "Oh it had one accident, must not be safe!" and yet humans have 100s of accidents a day.
The main problem here is expectations -- everyone expects the machine to be perfect, and when it's not, it breaks expectations. In the past, machines were generally a lot more accurate, they just didn't do a lot of stuff. Now they they can do a lot more stuff, their accuracy is coming down to human levels, and it throws everyone off.
I don't think we need to fix hallucinations, we need to fix expectations.
Humans don't have 100% perfect recall of every fact they've ever learned, why do we expect AI to?
Expecting more from AI is being cautious about wanton stereotype machines that will remove plenty of "redundant" humans from the workforce without demonstrating that they are capable of redundancy against error.
If we're digesting and calcifying human knowledge, it's likely we have only one shot to get it right, so we had better do it correctly.
Ask yourself if you would have been fine with calculators if they did math at an equal error rate to humans.
Having tools that shift blame and confidently lie to you is a problem. Here's a recent example that bothered me: I tried using Meta AI to generate an image and it failed. When I asked why it failed, it kept lying and making up excuses for how it was my fault or my browser was doing something wrong. When in reality the probable explanation was that the image prompt included something that was probably flagged as sensitive content, which is why the image generation failed. (It wasn't even anything nefarious, I just tried several variations of catgirls.) If something goes against some content guidelines I want the tool to tell me that, rather than gaslight me with lies.
I'm with you, it would be nice if LLMs could do those things, but also I don't expect my human friends to do that, so I'm not sure why I should expect it of my AI.
This right here is a perfect example of the problem of expectations. Do you have 100% accurate recall of every fact you've ever learned?
But we don't over-react when humans are fooled by optical illusions/magic tricks, or when they conflate two things, or when they confuse correlation and causation, or when they selectively prioritize some information over other information, or...
This describes the current state of AI, doesn't it?
I think that it does not happen at a level of persistence/pervasiveness that would cause the equivalent human to be hospitalized.
It might be more expensive, but perhaps Apple is uniquely positioned to synthesize/compare answers from multiple models to provide more accurate results.
As Segal's law states: A man with a watch knows what time it is. A man with two watches is never sure.
Well yeah, no shit, that's the technology they're using. Imagining making the opposite claim: that their system will never suffer from the inherent limitations of the underlying technology. The article even has a short paragraph mentioning that literally nobody else can make any stronger claims than this, yet it's presented as a headline-worthy statement.
Would it be the same world that some CEOs seem to think that we live in now?
Would it really become the job the killer that is feared, at just two nines?
Not an expert though, but isn’t that behaviour inherent to how it works? Bit of a misnomer and giving people the wrong idea of what is going on here.
The hallucinations/errors compound and can misguide decisions if you rely too much on AI.
for example, this is bullshit because it’s words with no real thought behind it: “if being correct was important, the developers would have taken an entirely different approach”
Your second comment is more flippant than mine, as even AI boosters like Chollet and LeCun have come around to LLMs being tangential to delivering on their dreams, and that's before engaging with formal methods, V&V, and other approaches used in systems that actually value reliability.
But those who sell AI products don’t want that.
It’s normal in small degrees even for mentally healthy individuals.
metaphors grounded in reality are fine. otherwise don’t call it a “generator”
A 'fact' that is not confirmed is an opinion. A 'opinion' sincerely held to be true without data is called faith.
I do think humans are quite bad at recognizing that most of what they think they know is opinion, and/or the 'facts' are often based on personal experience (which is by definition a cherry-picked data set), and thus is also an opinion too.
I think as well that those that practice science extensively are more practiced to really identify what is fact vs unsupported opinion. Without that practice, it's a lot easier IMO to then think a person knows a lot more facts than they really do. Which is to say, humans generally know very little, and it's not terribly comfortable to acknowledge and feel that way.
Which, leads us to the ultimate reality. Most of what a lay person has to say would be opinion, it's not really worth much - and it's even worse when data no longer matters.
Though, is the AI baseline even worse though?
Computers are nothing more than a set of persisted electrical 0/1 signals and a series of logic gates with intricate timing. There is no 'facts' in that world.
For example, "Chat GPT" does not understand questions, it does some pattern matching to find responses that probabilistic-ally would follow. Another example, AI does not understand it is drawing "a hand", it's just a probability algorithm that indicates these pixels are likely to be "good" following these other pixels.
If someone can build a 'reasoning' step into AI - that would likely be a game changer. An AI that generates an answer, and then compares that answer against an actual fact-database; and then modifies that answer in light of the existing facts. Even better, one that can also challenge when an understood fact is likely to be wrong. To do all that, the AI would need to understand abstract concepts. To give a baseline of where we are on that, my cat is able to do that and AI is currently at zero capability to do it.
It could be that those names are not poorly chosen. People want grant money, companies want venture capital money. What better way to do that than to use specifically chosen names (also known as propaganda!)
Next, there is certainly the "main character syndrome", and the hope of AGI and that these things are on that path, and we are the ones to unlock this. Someone with that point of view, would have a lot of incentive to make those comparisons.
Humans will not be outdone on other things in my or my children’s lifetime.
The answer I come to is "fancy autocomplete." If you know there is a significant chance (non zero) of a wrong answer, you have to validate every answer. To me, that describes auto-complete. Useful, but not something you can just use blindly (which greatly limits the utility, ie: human still required).
https://en.m.wikipedia.org/w/index.php?title=Michael_Crichto...
It's Gell-Mann Amnesia by another name.
I'm impressed with ChatGPT et al right until I query something I know a lot about, and it then proceeds to "hallucinate" far worse than any journalist or newspaper.
It makes me question if every other piece of information it's given me is factual.
It is the biggest "elephant in the room" and if/until it's corrected, every single one of these companies is on a trajectory towards worthlessness. It is a much bigger problem than "People make mistakes too."