Capturing the Flag with GPT-4
micahflee.com
micahflee.com
// returns 1337
return result;
Sometimes the comment stated the correct answer to the puzzle, but the script returned something else since it had a bug.It’ll tell you what a typical output for the command might be, and the more complex the script, the more wrong and full of hallucinations it will be.
There’s a huge difference.
Specifically, you have no way of knowing the difference between accurate outputs and inaccurate outputs, without running the command yourself, making it largely worthless.
Without access to environment, it’s not possible to magically know what the output of a command will be, it doesn’t have an embedded understanding of code, specifically not iterations and mapped data or mathematical functions.
For trivial obvious outputs it’s good, but it’s not executing the code; it’s generating what seems like plausible output; and if the output is trivially derived from the input, it’ll be impressively accurate.
…but, as the complexity of the task increases or the task deviates from “standard problem space” the triviality of generating accurate output decreases and it stops generating impressive outputs.
Tldr; yes, but it doesn’t scale well beyond trivial outputs.
The former is necessarily the case given the Halting Problem; the latter is falsified by the fact we can reason about code despite the Halting Problem.
> the latter is falsified by the fact we can reason about code despite the Halting Problem
i think wokwokwok's point holds true in practice.
Our patience and working-memory is far more limited than what is essential to accurately model all the necessary details of even moderately complex algorithms in our head.
One of the main reasons to limit code-complexity to improve readability/maintainability.
I'm not talking in general terms, or describing 'what the code does' in a summary or bullet point high-level form. No one is arguing that it can't summarize and describe what code does. These models are very good at that.
I'm talking specifically about generating the output of the command, as the OP specifically mentioned.
It does generate the exact output for commands and scripts if you request it, sometimes even if you don't, just as an example; they're just, often, hallucinated rubbish.
Being impressed that GPT can invent from 'thin air' some creative writing (fiction) when you tell it `pretend you're a docker container and now run 'ls'` is I feel, missing the boat, in terms of understanding or being impressed by the capabilities of these LLMs.
Nobody is impressed that it "can invent from 'thin air' some creative writing (fiction)", but that it often does not and in fact produces correct output. You're right we can't rely on it producing the correct output as it currently stands, but that it is capable of doing this at all is impressive.
For a trivial example, what is the output of:
```
while True:
pass
print("goodby world")
```(this is also proof that leaving out the curly braces makes code harder instead of simpler #python-lie-to-me. multiple edits to get this to render correctly on HN )
But it's important to note that just because there's no algorithm that works on ALL programs doesn't mean that the semantic properties of all programs are undecidable. Clearly for the particular programs where the program is bounded and guaranteed to terminate (e.g. no unbounded loops or recursion allowed) we can determine such properties, and I believe theorem provers in fact only allow such programs. And similarly you can restrict yourself to only the programs that you can prove will terminate in N steps (which might be excluding some programs that do terminate but require more than N steps of compute to prove).
It doesn't even have to be complex. Ask it about a program in a language that's not very popular and the odds that it'll completely screw up its answers is high.
I understand that the current LLM architectures are fundamentally incapable of goal-seeking and that they lack any concept of "correctness". However, I also recognize that somehow, the incredibly "wrong" architecture of ChatGPT is able to be useful.
I'd like to think that with access to a sandbox runtime environment, a significantly larger context window, and perhaps additional copies of the LLM "supervising/orchestrating" multiple "lower" copies of the LLM by breaking down large work into smaller tasks, that the current ChatGPT architecture could scale well beyond trivial scripts.
And then I hope we abandon this LLM architecture and develop architectures which can actually internally work towards "going beyond" in terms of quality of output for a given task.
I often have that same problem myself.
I agree that LLMs will often hallucinate. There is obviously no guarantee that the output is correct. But sometimes it is correct anyway, which I notice by actually running the code.
Here is a trivial example which I only mention to bring the conversation back to the reality of GPT-4 actually being able to do things like this:
Me:
You are a Python interpreter. Please give the correct output of the supplied code, with no commentary.
>>> a = ["wokwokwok", "says", "i", "have", "no", "understanding", "of", "code"]
>>> a.append("!")
>>> " ".join([w.upper() for w in a])
GPT-4, on first attempt:WOKWOKWOK SAYS I HAVE NO UNDERSTANDING OF CODE !
> Tldr; yes, but it doesn’t scale well beyond trivial outputs.
It is weird that you seem to agree that it is capable of performing algorithmic simulation, while discounting that with "but it’s not executing the code; it’s generating what seems like plausible output", in a way that seems suspiciously close to defining anything it simulates correctly as "trivial", and anything it would fail at as "executing the code"...
The query it wrote wasn't even valid SQL but it was close enough to make you think it would work.
It's like playing with an aimbot. You may be beating the other players, but where's the fun?
If I, as a person running a CTF, did not want my players to do this, I would set up a few problems which would have incorrect (but not obviously so) "solutions" generated when fed to LLM.
The Shamir's Secret Sharing reminds me of the time I was playing DEF CON CTF Quals, and had the bright idea to try to scan for challenges. I found one - it involved hiding fragments of a split secret in a modified version of ADVENT. I solved it. Even when the board was fully opened, it was nowhere to be seen. You does your hacks and you takes your chances...
Can you explain what this means? I don't understand, except for the split secret part.
Also: How did you do it?
http://point-at-infinity.org/ssss/
DEF CON CTF quals is (or was) "Jeopardy" style with five categories with five problems each. The thing I found was not one of the 25 problems.
Sorry to the author and thanks for the article, I learned about lagrange interpolation.
It would be a pretty big flex.
The MIT Mystery Hunt has occasionally managed to get the NYT crossword on the day of the hunt to contain a clue answer or two.
It looks like a bog standard function for converting to a base X number. All GPT-4 had to do was paste 27 in.
But the way the challenge was designed, it's more about just changing argv[0] rather than the actual executable path.
The binary was not setuid, but was only executable (not readable) by the user used.
Ah, then ptrace/gdb could have been used to dump it out as well :). Looks like a fun CTF, too bad I was too busy for bsides this year..
It's been a long while since doing basic linux administration. I am getting rusty.
Should've had *GPT*-4 proofread this.