I couldn't find a way to contact the researchers.
I couldn't find a way to contact the researchers.
But I also have extreme doubts that proper redaction can be done robustly. The design mockup image suggests that this will all be done as a step subsequent to response generation. Given the abundance of "prompt jailbreaks", a determined adversary is going to get around this.
Redacted = dangerous
https://images.openai.com/blob/047e2a80-8cd3-41b5-acd8-bc822...
I've prompted ChatGPT to make a bit more detailed explanation: https://chat.openai.com/share/42e55091-18c2-421e-9452-930114...
You can probably prompt it to further to generate python code and unmask the file for you, in the interpreter.
Incidentally, this use of GPT4 is somewhat similar to the threat model that they are studying. I'm a bit surprised that they've used plain GPT-4 for the study, rather than GPT-4 augmented with tools and a large dataset of relevant publications.
"No GPT-4 tool usage: Due to our security measures, the GPT-4 models we tested were used without any tools, such as Advanced Data Analysis and Browsing. Enabling the usage of such tools could non-trivially improve the usefulness of our models in this context. We may explore ways to safely incorporate usage of these tools in the future."