AI hallucinates software packages and devs download them
theregister.com
theregister.com
People blindly accepting the output of LLMs seems similarly crazy to me and I think it's only a matter of time before we face a real reckoning over this. The lesson here I think is just because a lot of people are advocating something that seems reckless doesn't mean it isn't reckless.
https://www.howtogeek.com/162282/what-activex-controls-are-a...
That's basically what's happening every time a developer runs "npm install" or "pip install". Sure, it technically not be a "binary", but it makes no practical difference given that no code scrutiny is given the majority of the time.
The introduction of LLMs doesn't change much. People are already predisposed to copy paste random commands from blogs/forums. LLM is just one more source on top of that.
I was thinking about this the other day. I wonder what my leadership would say if I told them I spent the day scrutinizing some of our open source dependencies. I assume even a day would be treated as wasted time, especially on the product side.
FWIW, I used to do this back in the early Rails day and was encouraged to do so. I ended up contributing heavily to the Rails ecosystem because of it, and it was all encouraged by my employer at the time, but they were a relatively small startup at the time and viewed things very differently than the FAANG I work for today.
Irony in the "Greek tragedy" sense (but who is εἴρων and who is ἀλαζών?)
So many online arguments about AI safety, where someone says "it's fine: just don't let it out of the box", and here it is, accessible on the internet — and for some, even this isn't enough: they want the weights; and for some, even the weights aren't enough, they want the training data.
So many arguments where someone says "it's fine: just switch it off if it goes wrong", and not only is the switch out of reach to those affected, most people aren't even checking if it's going wrong in the first place.
The labs themselves ring the alarm bells, put disclaimers on the start page of their own products, call for binding regulations, and there's always someone dismissing this as a "5D chess" marketing move.
So many where someone says "it's just speech, what can it do?", etc. etc.
I'm glad the AI labs themselves have so many signatories to the "let's pause for a bit" open letter. Cool though the tech is, it's definitely acting outside the expectations many have for it, and a simple dichotomy of "better/worse" isn't sufficient to describe the ways in which it is outside expectations given it can be both at the same time.
And it is very clear that the AI companies do want you to be scared of the former, because it makes their tech seem more impressive and distracts from its very real flaws.
Stupidity has of course always existed and you can't regulate it away, but modern tech and the business interests behind it seem to me to be making it worse and potentially more destructive.
I'm not talking about just the latest hype here, it has been a trend for more than a decade now: "don't think for yourself, just consult the magic oracle in your pocket that surely knows best and isn't purposely designed to manipulate you into buying more crap and subscribing to the worldview of a small elite"
But like with the AI apocalypse, or any conspiracy theory, the really scary thing is that there doesn't have to be some kind of intelligent plan behind it, something that you could understand or fight against, or maybe join in because you think they have the right goals. Rather, it's just mindless forces of the market and social dynamics.
I really think it's all going to collapse within another 15 years or so, and can only hope that civilization will recover and perhaps learn from the mistakes made.
"They're the same picture".
Imagine someone who wants to rule the world, just asks an AI for such a plan. AI doesn't need to want anything itself, just be given free rein. I'm sure Boris Johnson would have taken this option if it had existed, though I doubt it would have made any difference to the character flaws which were his ultimate downfall.
Even the "paperclip maximiser" scenario is just someone in the business of making paperclips asking an AI for help, and who doesn't look too closely until it's too late.
Myself, I'm… relatively optimistic, in that I think these systems are likely to be fragile to distribution shifts they themselves create, in ways that mean the rest of us can likely stop them from being existential threats if they start down a dangerous path. (This requires that we never figure out how to make them learn from as few examples as humans need, which may be wishful thinking, but for now seems to be an acceptable guess).
For some reason this reminded me of the old joke about the engineer, the physicist, and the mathematician's differing approaches to noticing a fire; I hope humanity doesn't take the mathematician's solution: https://jcdverha.home.xs4all.nl/scijokes/6_2.html
I’m still waiting for a similar reckoning over the average JS project’s tendency to pull in hundreds of largely unvetted transitive dependencies. But, aside from left-pad, it seems not to have really happened…at least that we know of.
Humans are just such stubborn creatures.
chatGPT3.5 hallucinated packages. chatGPT4 has found me about 10 awesome python packages I didn't even know existed.
Personally, I don't really care if anyone drinks or not. I am super hydrated and if people want to stay dehydrated they are just going to get smoked.
There are already real-world consequences of this:
https://arstechnica.com/tech-policy/2023/05/lawyer-cited-6-f... https://www.reuters.com/legal/transactional/us-judge-orders-...
Before this there were already lots of real world consequences of "computer says no" anyway, so its just a continuation and perhaps escalation of that.
This is what leads to people even in the industry to take at face value a human-like response in helping answer a query; and in this case, willingly downloading malware.
If I imagine a very naive translation system that just does a dictionary substitution, and the dictionary has some bad entries, the correct and incorrect results look the same, but it's still a bug.
The system as a whole still messes up because "dictionary substitution" is not at all sufficient:
Hydraulic ram -> aqua ovis -> water sheep
… this probably counts as just "hallucination" while arguably not being a "bug", because the dictionary is correct (or close enough, I don't speak Latin), but the system as a whole can't ever be better because the model it relies on can't represent the right things.
But in an LLM, there are no explicit rules (or only in some obscure background layers), it's all statistics. Statistics stacked on top of more statistics in a fascinating self-stabilizing way, where each guess provides some support to its peers, like how the weak paper slabs support each other in a house of cards. But it's statistical guesswork all the way down. Even the most correct answers are merely statistics playing out favorably in a case that may or may not have been very easy to get right.
OpenAI was founded with that name in 2015; transformer models were introduced with 2017's "Attention Is All You Need"; Even ChatGPT's immediate in-house predecessors in the form of InstructGPT and GPT-3 were basically ignored by the general public.
No, what made these companies valuable is that ChatGPT specifically crossed the threshold into being vaguely interesting, and then everyone else copied it.
Likewise the finite state machines controlling NPCs in video games — and the search tree in Deep Blue and the learning from self-play system in AlphaZero — are called "AI": they're good enough to be interesting, regardless of your position on the question "what is this 'thinking' thing anyway?" which led to Turing coming up with the eponymous test.
and other emoticons?
Actually people did this before software is even a thing.
That despite every example of an LLM failing in production giving results no competent human would actually produce unless they were committing fraud or had severe mental deficiency, that and no one else would find acceptable.
You sure it's "precisely" and not "analogously"?
Ironically, one common human failing is to use binary classification.
> or had severe mental deficiency
It seems I have worked with more human idiots than you. And at least one time where I was the idiot, despite literally scoring "off the charts" on a cognitive abilities test at school.
I'd agree that LLMs are not as smart as humans — they had to read 10% of the internet just to be as competent as an intern — but to me they're more like newspapers and the Gell-Mann Amnesia effect than "severely deficient", even when indeed unacceptable.
This is completely unnecessary.
Maybe it was this paper: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10619792/
The authors’ statement that ‘unlike hallucinations, confabulations are not perceived experiences but instead mistaken reconstructions of information which are influenced by existing knowledge, experiences, expectations, and context’ is pretty compelling.
A real intelligence would be able to either put together the information needed and fill in gaps with research or it would be able to tell you that it does not know.
"Hallucination" is yet another anthropomorphism of the model. A mind hallucinates; a model just produces incorrect results.
LLMs are very cool but I'm really sick of the "intelligence" branding. If a person who didn't know about something just confidently made shit up, I wouldn't call them intelligent.
And I agree with you, I know people who make shit up about something they don't know about, and it goes "actual humans > LLMs > these people".
Well, now it does I'm affraid :-)
I know we have had this conversation with our daughter several times and we have three languages to work with...
Where all the "cloud" is, is an server in a datacentre with a fancy gui.
I guess we've needed a new buzz world to keep the population excited.
It seems to be a good recipe for success in many fields (business, politics, art, apparently somewhat in social sciences if you make up the evidence for it too) so I'm not too sure about that!
Then you have a highly flawed understanding of humanity... to the point you should check out of tech for awhile and go read about psychology instead.
Extremely intelligent people confidently make shit up when they don't know the answer all the time.
I'm not a python user, but in most languages, libraries are referenced in (what most devs would consider) boilerplate code. Purely conjecture, but perhaps without boilerplate code, the LLMs are left guessing the names of popular libraries and just merges together two common naming conventions "huggingface" and "-cli".
ai is a bad non-losless search engine.
For huggingface, the tool is called huggingface-cli, but the package is called huggingface_hub[cli].
IMO, thats bad naming. If you make a tool called X, just publish it in a package called X.
I think what's worse is publishing a package under a different name than its root namespace.
I question if one should ever do that... Packages are free... Just make one package per binary, and then a metapackage for those who want to install the whole suite.
Like getter/setter in java for all your attributes.
I would never consider imports boilerplate code.
But isn't it also a waste to use data centers full of GPUs to process terabytes of text to accomplish the same thing better programming language design could?
If it wouldn't be beneficial for writing code, which it is, we wouldn't use it.
But yes if we could create better languages or systems, it would be a waste. But we tried multiply new programming languages, we have no code platforms etc.
It does look like though that LLM is still better than all of those approaches.
But thats not the point of it anyway?
Its about writing code faster and potentially better. Cloud GPUs can also generate unit tests etc.
I primarily use it for languages i don't use often enough, nonetheless its only a question of time until it doesn't make sense anymore to write code yourself.
You an I seem to have different values. I have never desired quicker things that are worse.
There is already something like this for e.g. `nvidia-tlt` which exists on PyPI, but just as a placeholder telling you to go and add Nvidia's pip repository.
[1] i.e. specific to my niche where i didn't know the answer immediately.
lol
If it's not already easily googleable, the LLM will just spit out some boilerplate apology and then double down on the previous wrong answers, provide something correct about a tangentially related topic, or just reword my prompt as a course of investigation that - if written by a person - would just be a polite "fuck off and go figure it out yourself."
ChatGPT seems to be able to fix, I'm going to say "about half" of it's own mistakes, but that's my gut feeling and not a detailed analysis even within the domain of questions I've been asking (which itself probably isn't representative).
The truth is, while both humans and AI can make errors, and both can be malicious as well, the actions of AI will be counterintuitive and confusing and we won't know how to counter them in the same way that we counter human follies.
This is just one aspect that shows that AI is making society worse on average and it should be destroyed.
The output of AI can often be confusing, yes, and indeed this is part of the issue Yudkowsky has with AI in general: we can't predict it in detail, especially not in domains where it out-performs us (like Chess).
However, this is tangential to any question about "making society worse" — even when it's making society better (because it's more capable than we are at something which we care about), Yudkowsky would caution that this may be a 5D chess move to gain more power.