The resulting system won't have the unbounded flexibility that our existing models have, but if they're provably safe that will make up for it.
The resulting system won't have the unbounded flexibility that our existing models have, but if they're provably safe that will make up for it.
That would essentially require a "non-Turing-complete" prompt language. Because if the prompt language was effectively Turing complete, it'd be impossible to determine whether every possible prompt would produce a "safe" outcome or not. This would severely limit what the LLM could do even compared to GPT3.5.
>Again, we did it with type systems and proof assistants.
Proof assistants require a human to provide the actual proof whether something is safe (correct) or not; they can't do it automatically except for very limited, simple classes of programs.
> Proof assistants require a human to provide the actual proof whether something is safe (correct) or not; they can't do it automatically except for very limited, simple classes of programs.
Finding a proof is in NP (at least if you restrict yourself to proofs that are short enough that a human might have a chance to write it out in their lifetime). So computers can do it.
We will just do to LLMs what we are already doing to people.
What we're talking about here is social engineering of LLMs. That's currently pretty easy. It will get harder but it cannot be made impossible.