AI is growing faster than companies can secure it, warn industry leaders
venturebeat.com
venturebeat.com
We might say the same thing about people.
Ultimately, just how problematic is it to label something as a hallucination? Are investors about to be massively duped? If I create a mechanism to reduce hallucinations and I call it therapy, is that really problematic?
No, it would be fair to say they have a "grasp" of predicting the next word in a given sequence of words based on a set of words in their training set. Hallucination then is what people call their inherent tendency to run into a situation where the probability of the next word being predicted "wrong" is high. And once one "wrong" word has been predicted the probability that the next word is also "wrong" rises exponentially.
LLMs do not have any grasp of reality. They just predict text based on trained patterns. Too many people have been fooled into believing that LLMs can understand anything about reality, but a word-based description of reality is not the same as reality.
> We might say the same thing about people.
If you want to reduce a human being down to being a word predictor, then I guess you could say that?
I’m not sure we really suffer from it. In our internal analytics AI tends to slow employees down and make them less productive. This is in general. The exception is experts using it, where LLMs increase their value output by an ok margin. Especially within our own programming team LLMs have proven a real challenge. On one hand they are fancy auto-complete which will speed an experienced developer up by so much it’s hard to ignore. On the other hand it takes an experienced developer to know when they get things wrong. Not the things that literally won’t work, but the things that will work poorly. I haven’t been too involved outside of our team, but I imagine it’s the same in any field.
Which is where the “hallucination” term sort of back-fired. It was a good way to make people buy into the value of LLMs by making the mistakes oddities almost negligible. The issue is that those mistakes can have such massive impacts that the entire trust in the AI industry falters. I mean, we had one of the CEOs ask if we could switch to Linux now that Windows includes AI… obviously we can’t do that without going bankrupt, but it tells you something about the worry in the non-tech enterprise top.
LLMs know a whole lot more about the uncertainty of their predictions than they say.
GPT-4 logits calibration pre RLHF - https://imgur.com/a/3gYel9r
Language Models (Mostly) Know What They Know - https://arxiv.org/abs/2207.05221
The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets - https://arxiv.org/abs/2310.06824
The Internal State of an LLM Knows When It's Lying - https://arxiv.org/abs/2304.13734
LLMs Know More Than What They Say - https://arjunbansal.substack.com/p/llms-know-more-than-what-...
Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback - https://arxiv.org/abs/2305.14975
Teaching Models to Express Their Uncertainty in Words - https://arxiv.org/abs/2205.14334
People want some improved automation, but not too much.
that's right, they want automation to reduce their costs, but not replace their jobs/business. the people employing them also want the same thing.
I'm surprised things worked out so well.
> Zhou shared a striking example of how AI-generated content could lead to real-world consequences. “Some of the initial stock images of various ingredients looked like a hot dog, but it wasn’t quite a hot dog—it looked like, kind of like an alien hot dog,” he said. Such errors, he argued, could erode consumer trust or, in more extreme cases, pose actual harm. “If the recipe potentially was a hallucinated recipe, you don’t want to have someone make something that may actually harm them.”
There’s absolutely no reason Instacart has to show customers AI-hallucinated recipes from stock images. They choose to do it, then beat the drum about AI security as if they actually give a shit. It’s like Boeing self-certification.
I think we should program and train LLM with universal recognized agogic principles instead of being neutral in this regard, to encourage critical thinking and prevent 'reality tunnels' in the mindset of the users and perhaps incorporating this also in future training and curating techniques [2][3][4]. How to raise GenAI and future AGI well.
There are LLM training techniques to alleviate hallucinogenic, contradictory and confusing answers. These might be augmented with insights from chaos engineering, cognitive psychology and persuasion - making them agogic truth machines scoring very low on hallucination benchmarks [1].
I think we should program and train LLM with universal recognized agogic principles instead of being neutral in order encourage critical thinking and prevent 'reality tunnels'. Perhaps incorporating this in future training and curating techniques [2][3][4]
* Data curation Ensure data used to train AI models is balanced and diverse helps in preventing biases that could lead to hallucinations or harmful outputs. So curating data from a wide range of sources, cultures and viewpoints. Implementing quality control during data collection and preprocessing to filter out unreliable, outdated, or biased information.
* Targeted post-training (fine-tuning) After initial training models can be fine-tuned using datasets specifically designed to emphasize helpfulness, harmlessness and alignment with ethical principles. Embed ethical guidelines in datasets, for example include scenarios to handle sensitive topics, avoid hate speech and promote fairness.
* Red-teaming Red-teaming involves stress-testing the model by simulating adversarial attacks or intentionally providing challenging prompts to see how the model responds. This helps identify weaknesses, such as susceptibility to generating harmful content or hallucinations. This can be used to improve the model's robustness and safety.
* Post-training datasets focused on responsible AI principles Incorporating datasets that help the model understand context and nuance of various topics, ensuring it can provide appropriate responses to the situation.
* Refusal-aware instruction tuning While data curation, targeted post-training, and red-teaming help to prevent the introduction and propagation of false or harmful content, R-tuning directly enhances the model's ability to recognize its limitations. Enabling the model to refuse to answer questions beyond its knowledge.
* Iterative user feedback based refinement Continuously collecting and analyzing feedback from users and independent review teams helps identify issues that may not have been apparent during development.
[1] Vectara hallucination leaderboard https://github.com/vectara/hallucination-leaderboard
[2] On epistemic black holes: How self-sealing belief systems develop and evolve". Maarten Boudry and Steije Hofhuis in the journal Theoria August 2024 https://onlinelibrary.wiley.com/doi/epdf/10.1111/theo.12554
[3] Costello, T. H., Pennycook, G., & Rand, D. G. (2024, April 3). Durably reducing conspiracy beliefs through dialogues with AI. https://doi.org/10.31234/osf.io/xcwdn https://osf.io/preprints/psyarxiv/xcwdn
[4] BriX: Reducing polarization through Bridging and eXposure https://research.qut.edu.au/genailab/projects/brix-reducing-...
AI creates a ton of garbage unmitigated, and cracking down on when that garbage is harmful will be a good way to reduce the amount of garbage these companies put out.
It is the users, and laws, that must control this.
Companies should add traceability to their AI products.
Govt.'s should ban non traceable content.
Holding users accountable for their actions is the only way forward to be safe and innovative at the same time.
If a hammer manufacturer sometimes produced a hammer that exploded like a hand grenade on impact, they would be held liable for that.
Of course then they couldn't leak the latest exponential growth stories to the press, which now appear every week.
After all, it's the same industry that came up with pervasive and invasive tracking, automated insurance refusals, automated credit lookups and checks, racial ...ahem... neighbourhood profiling for benefits etc. etc.
The security aspect mentioned here is that ai generated recipes could poison you when they hallucinate. The fix is “governance” which isn’t really described or defined, but no doubt it’s as necessary as it is costly. We could probably just not use cooks that seem to randomly poison the food and not create a new industry of equally suspect chef-policing but hey, where’s the fun in that?
This is the key. Folks are looking at current capabilities rather than the trend line. We need to be ahead of the development AI. There probably ought to be laws regarding how AIs can be used or not used. There probably ought to be required disclosure when AI is used to create a work of art.
Laws are going to do absolutely nothing here except slow down the wrong people. The world is full of people who don't play by rules. It's a level playing field already. Everybody is doing AI at this point: the Iranians, the Chinese, the North Koreans, criminals, sociopath entrepreneurs in Silicon Valley, and everybody else with bad intentions you can think about. Every idiot with half a clue as well.
Laws are for obedient citizens that aren't a threat. It's all the other people we need to worry about. Besides, US law is worthless across its borders. And I don't even live there. I live in a place with actual anti AI laws (Europe) and I think they are a really bad idea. I'd prefer to be in a place that isn't awaiting the inevitable meekly with their hands tied behind their backs.
There is a nice little howto build and LLM from scratch article featuring on the HN front page today. Assume everybody else read that too. And then some. This is all public knowledge at this point. You can get pretty far coding an LLM with the help of an LLM. Even the offline ones are probably not that bad for this. Any idiot can build an LLM at this point. Even me probably. So that means that world + dog has been figuring out how to use LLMs to their advantage for the last decade or so. Probably with a sharp uptick in interest about 2-3 years ago. We've seen nothing yet.
Given all that, the safest course of action is working on the assumption that this already happened or if not that it will be happening very soon. The trick is not to prevent the technology but to be better at it than everybody else. This is an arms race and we can't afford to let the wrong people win it.
Regarding the fears about art and impersonation. Same thing. That too is going to happen. Nothing we can do about it. No amount of virtue signalling is going to help here. So, I would propose the exact opposite. Let people do what they will (again, inevitable that they will). But require them to sign all their work and assume any unsigned work to be fraudulent, malevolent, and worthless. Illegal even (laws can work to our advantage here). AI work should be signed too. So we can trace it back to who, when, and how. AI will be able to fake a lot of things. But not cryptographic signatures.
It's a lot easier to regulate proper use of signatures than abuse of technology. And the beauty of this is that we've had the technology to do this for decades already. But we still send unsigned emails, publish unsigned articles on blogs, and don't bother with encryption in general, etc. This is stupid and it has been for quite some time. It's only the unsigned work that's easy to fake.
Some scandals around LLM generated content could be exactly what we need here. I'm talking reputation damaging scandals that have the likes of the NYT having to explain in public how they messed up so badly and paying millions in damages to the victims. Money is a great way to incentivize companies to level up their game. The the likes of the NYT need to get paranoid about checking authenticity. And everybody else as well.
A few geeks wearing tin foil hats aren't much of a solution. PGP flopped for that reason. But it's not to late for it to make a comeback in some form.
IMHO the Fediverse is an obvious place to start. Why accept any content there that isn't signed? It should be stupidly easy to level up its protocols to add and verify signatures to content and profiles. I've actually considered having a go at it at some point. No time and other priorities. So, I havent. But why isn't this a big topic in the wider community?
There are such laws:
https://ec.europa.eu/commission/presscorner/detail/en/ip_24_...
According to the polls, yeah, it's unpopular.