A new way to build with Large Language Models
blog.fixie.ai
blog.fixie.ai
No wait, that doesn't have a great acronym.
If developers on your platform are building code that accepts prompt input from untrusted users do you have any measures in place to avoid them tricking the system into executing functions in an inappropriate or unintended way?
One thing I'd find useful would be a limit on the number of functions that can be called during a single execution - and rate limits and budget calls on functions generally might be smart as well.
I'll never understand this cultural phenomenon. Anybody can open the browser inspector on a random social site and tweak the page to say whatever they want and send screenshots around "implicating" the poor bastards - but none-hypothetically who actually cares?
Sort of akin to running a pub and having a drunk run his mouth. That's not on the establishment that is on the individual.
If you're going to give a LLM-driven program the ability to execute functions that can change state in the world you need to understand prompt injection, so you don't accidentally build something that you really shouldn't have built.
The first public prompt injection vulnerability was a Twitter bot which started spam-mentioning people and threatening the president. https://arstechnica.com/information-technology/2022/09/twitt...
Sure. But also, so? Seems like a storm in a tea cup. "Our twitter got hacked", the end. There are countless such examples that have nothing to do with prompt injections. If anything the phrase "all publicity is good publicity" comes to mind.
I remember in the mesozoic era of the internet (early 00's) there was something like a KFC promo website where they put a guy in a chicken suit and had him respond to chat suggestions directly from users 24/7. Hilarity, as you can imagine, ensued. Shenanigans even.
If some user submitted content (even if it is sieved via a model) ends up "inappropriate" you can just delete it and move on.
I probably wouldn't hook it up to a shell running commands but the same generic advice as with all untrusted input applies..
Developers who don't understand prompt injection will continue to make nasty mistakes like that.
- Rate limits are enforced to provide caps on agent and function usage.
- Execution depth is capped to prevent the LLM from getting into loops.
- Function output is sanitized to prevent corruption of LLM state.
- Functions execute in a completely separate environment from the rest of the service, including the LLM, to reduce the impact from bad functions.
Note that this doesn't entirely prevent against "; DROP TABLES"-type hacks against the implementation of the function, but that problem isn't unique to us. It may however be possible for the LLM to look at function inputs and flag overtly malicious ones.
So, the first thing I'd want to try is hooking up some basic arithmetic to see if its math gets more reliable. And if not, why not?
Seems like someone would have tried this already?
I assume the output of the LLM is monitored by fixie and the result of the function inserted allow the LLM to respond again?
Having internal users ask questions then have a ChatGPTesque system answer with injected data would be nice. Would be very tightly controlled (or it couldn't answer what it doesn't have).
Something between a rigid chatbot that's exists now and ChatGPT that just makes up plausible answers.
Stage after that: GPT, or some other AI constantly adapts / curates the agents for me based on my notes, conversations, etc. (Although a privacy nightmare)
What's the interesting demo here? Did I miss something?
What exactly is running on my infrastructure? Do I need to buy GPUs? What state needs to live on Fixie cloud infrastructure?
One of the conditions for this, is for chatGPT models to start being able to write their own code in order to produce models of themselves that are more accurate and more efficient. Given this ability, and some fitness criteria, genetic algorithms may be used to create new LLMs. This sounds like science fiction, but once the compute requirements come down for these models (by a couple orders of magnitude), I believe this may be possible.
To what extent does your model allow for semantic models to create semantic models that are themselves more efficient in relation to some fitness criteria? Can I tell a model "You (model) I want you to reproduce using interaction with these other models (some collection of other models) and have the child model offspring be more efficient according to this criteria [for example the resultant models will create short stories that are more likely to receive high ratings on a subreddit devoted to short stories]".
You would need to get around the "model pollution" problem in which LLM models pollute the space for which the models generate data because other models are producing web artifacts (Ted Chiang's Xerox of a Xerox problem). I call this the problem of alpha (direct experience). One of the ways I've thought of to fix this is to have models trained on direct user input (such as cell phone video and pictures from a single user) - I have to admit that I got this idea from Neal Stephenson's Snow Crash (see Gargoyle). If your platform can integrate with visual processing this may have a high information density - object detection in daily videos demonstrating how objects are related to each other in the real world of the user and correlating these into a semantic network.
I'd also suggest that Obsidian integration might be useful.
This is exciting, thanks for making the Fixie SDK public.