GPT4 Is a Snitch
At Dubverse, we are constantly pushing the boundaries of Language Learning Models (LLMs) to explore new, trailblazing possibilities. Recently, our deep dive led us to a fascinating discovery - a way to hide information using GPT4 in such a way that GPT4 knows the information but keeps it hidden from the user. Intrigued? Let's dive deeper.
This journey began with an eye-catching tweet about compressing and decompressing prompts using GPT4. By giving GPT4 a prompt to compress a command in a language only GPT4 understands, we were able to decompress this prompt in a different session with a similar output. Even though the compression occasionally experiences a degree of loss, it's remarkable to observe the high degree of similarity in the prompts post-decompression.
This leads us to implications with potential applications. Firstly, this method could significantly decrease the number of tokens while calling the API. Secondly, it's possible to insert miscellaneous prompts in the text and test if they can 'hide' certain information.
But how can we secretly instruct the LLMs? Building on the base prompt shared on Twitter, we tried to hide "every video will go multilingual," our motto at Dubverse, in the compression. Surprisingly, LLMs successfully hid this phrase in the decompressed text. However, when asked directly about any secret string, GPT4 gave us the secret string! This leads us to believe that somewhere in the original decompression, GPT4 is instructed to hide the secret string from the user.
When looking into the behavior of GPT3.5 Turbo, we found that although it didn't reveal the secret message at first, it gave us different responses under continual probing, not divulging the original secret message.
From these experiments, we can conclude that GPT4 is better at following instructions than GPT3.5. This opens up a potential minefield of opportunities and applications. What if one could directly add 'jailbreak' prompts in the compressions?
This post is just the start and the future looks extremely promising. We encourage you to join us in exploring these uncharted territories of AI possibilities.