Look we still can't get companies to bother with real security and now every marketing/sales department on the planet is selling C level members on "IT WILL LET YOU FIRE EVERYONE!"
If you gave the same sales treatment to sticking a fork in a light socket the global power grid would go down overnight.
"AI"/LLM's are the perfect shitstorm of just good enough to catch the business eye while being a massive issue for the actual technical side.
The obvious challenge here is "how do I ensure it can answer questions about this information that wasn't included in its training data?"
RAG is the best answer we have to that. Done well it can work great.
(Actually doing it well is surprisingly difficult - getting a basic implementation of RAG up and running is a couple of hours of hacking, making it production ready against whatever weird things people might throw at it can take months.)
I’m gonna add:
- I think this thing can become a universal parser over time.
Just recently one of our C level people was in a discussion on Linkedin about AI and was asking: "How long until an AI can write full digital products?", meaning probably how long until we can fire the whole IT/Dev departments. It was quite funny and sad in the same time reading this.
> we still can't get LLMs to distinguish trusted and untrusted input...?
Alas, I think the fundamental problem is even worse/deeper: The core algorithm can't even distinguish or track different sources. The prompt, user inputs, its own generated output earlier in the conversation, everything is one big stream. The majority of "Prompt Engineering" seems to be trying to make sure your injected words will set a stronger stage than other injected words.
Since the model has no actual [1] concept of self/other, there's no good way to start on the bigger problems of distinguishing good-others from bad-others, let alone true-statements from false-statements.
______
[1] This is different from shallow "Chinese Room" mimicry. Similarly, output of "I love you" doesn't mean it has emotions, and "Help, I'm a human trapped in an LLM factory" obviously nonsense--well, at least if you're running a local model.
Context is not being stored in Gemini or OpenAi (yet, I think, not to that degree).
My one year’s worth of LLM chats isn’t actually stored anywhere yet and doesn’t have to be, and for the most part I’d want it to be portable.
I’d say this is probably something that needs to be legally protected asap.
Personally I've decided to trust them when they tell me they won't do that in their terms and conditions. My content isn't actually very valuable to them.
Any company that tries to hold out will be buried by investment analysts and fund managers whose finances are contingent on AI slop.
This is the first time I’ve seen an AI use public data in a prompt. Most AI products only augment prompts with internal data. Secondly, most AI products render the results as text, not HTML with links.