G-3PO: A protocol droid for Ghidra, or GPT-3 for reverse-engineering
medium.com
medium.com
That’s a script for the reverse-engineering tool Ghidra that uses GPT-3 to de-compile machine code and to write plain English explanations of what a piece of code does.
The article is quite detailed and describes both its capabilities and its limitations. That G-3PO script is open source, MIT license: https://github.com/tenable/ghidra_tools/tree/main/g3po
There was also another HN story about what at first sight looks like an alternative implementation of the same idea: “GptHidra – Ghidra plugin that asks OpenAI Chat GPT to explain functions”
https://news.ycombinator.com/item?id=34165291
This one is more recent and lacks that good write-up mentioned above. The script is smaller and it seems to have fewer features.
I suggest checking both of them.
https://github.com/tenable/ghidra_tools/blob/main/g3po/g3po....
As a point of comparison, I fed 10.2.2 a copy of gojq 0.12.11 that I had lying around and this is pretty representative of its output
if (DAT_0075c6c8 == (code *)0x0) {
ppuStack_38 = (undefined **)0x45ecd7;
FUN_00462ee0(&DAT_0075da28,local_10,iVar3,iVar4);
*(undefined8 *)(in_FS_OFFSET + -8) = 0x123;
if (DAT_0075da28 != 0x123) {
ppuStack_38 = (undefined **)0x45ecf8;
FUN_00460dc0();
}
}
for further comparison, I fed it actual jq and it did much better about the string literals if ((((((iVar6 == 0) || (DAT_00108018 = DAT_00108018 | 1, local_58 != 0)) &&
((iVar6 = FUN_001045a0(pFVar20,0x72,"raw-output",pcVar13), iVar6 == 0 ||
(DAT_00108018 = DAT_00108018 | 8, local_58 != 0)))) &&
((iVar6 = FUN_001045a0(pFVar20,99,"compact-output",pcVar13), iVar6 == 0 ||
(local_5c = local_5c & 0xfffff8be, local_58 != 0)))) &&
((iVar6 = FUN_001045a0(pFVar20,0x43,"color-output",pcVar13), iVar6 == 0 ||
(DAT_00108018 = DAT_00108018 | 0x40, local_58 != 0)))) &&
(((iVar6 = FUN_001045a0(pFVar20,0x4d,"monochrome-output",pcVar13), iVar6 == 0 ||
(DAT_00108018 = DAT_00108018 | 0x80, local_58 != 0)) &&
((iVar6 = FUN_001045a0(pFVar20,0x61,"ascii-output",pcVar13), iVar6 == 0 ||
(DAT_00108018 = DAT_00108018 | 0x20, local_58 != 0)))))) {
iVar6 = FUN_001045a0(pFVar20,0,"unbuffered",pcVar13);I haven't tried it yet, but I can see it being useful if it's somewhat accurate (that's a big if), and quite different from what Ghidra gives you in pseudo code.
I could use this for a regular project for which I have the source.
https://github.com/JusticeRage/Gepetto/blob/main/gepetto.py#...
They mention that the AI's decompilation is about as good as Ghidra's (though of course less trustworthy).
The benefit is in explaining the decompiled code, they give an example where the prompt to the AI is something like "here is some code decompiled with ghidra, explain it in detail".
From the article:
"the paraphrase of disassembled or decompiled code into high-level commentary, can be assisted by automated tooling as well.
And this is just what the G-3PO Ghidra script does.
The output of such a tool, of course, would have to be carefully checked. Taking its soundness for granted would be a mistake, just as it would be a mistake to put too much faith in the decompiler. We should trust such a tool, backed as it is by an opaque LLM, far less than we trust decompilers, in fact. Fortunately reverse engineering is the sort of domain where we don’t need to trust much at all. It’s an essentially skeptical craft."
All these GPT-based plugins are just toys. There's more serious research like this[1][2][3]
[1] https://keenlab.tencent.com/zh/2019/12/10/Tencent-Keen-Secur...
[2] https://keenlab.tencent.com/zh/2020/11/03/neurips-2020-camer...
[3] https://keenlab.tencent.com/zh/2021/08/11/2021-binaryai-publ...
You could probably also take the explanations from the LLM, convert those into embeddings, and then do semantic search over all functions in a binary. For example, searching for "get process handle and inject dll" and getting a list of prospects. It's less useful in an obfuscated binary, but for things like modding games or extending end-of-life software it could be very useful.
Maybe this is the next generation of automated code scanning.
Next feature on github: "Our LLM has scanned your code and found a potential buffer overflow. Please mark as a bug or a false report"
It feels like it'd be difficult to acquire a large corpus of vulnerabilities to train on.
ShiftLeft
Oversecured
etc.
There are a few companies doing this.
But there is still the issue of whether companies doing this kind of RE want to send their code to OpenAI's servers. If you're reverse engineering in order to determine whether you should sue another company for copyright infringement, you are probably cautious about sharing code in the first place.
For this use case there will always be a need for some kind of alternative that you can self-host without sharing your data with another party.
Here's the public github repo: <https://github.com/tenable/ghidra_tools> Please feel free to fork it and send me PRs with any features or fixes you think might be useful.
Explain the following C++ code:
int sum(int n) {
int result = 0;
for (int i = 0; i < n; i++) {
if (true || false) {
return 42;
}
}
for (int i = 0; i < n * 2; i++) {
i++;
}
return n;
}
Then ask it what is the result of that function when n is 5, and then when n is -1.Note that the function looks trivial but it has certain tricks to confuse it. e.g.: "result" is never returned, instead, either 42 or n is returned.
Explain the following C++ code:
int sum(int n) {
int result = 0;
for (int i = 0; i < n; i++) {
if (true || false) {
return 42;
}
}
for (int i = 0; i < n * 2; i++) {
i++;
}
return n;
}
This C++ function is called sum and it takes an integer argument n. It has a local variable result which is initialized to 0 but is not used in the function.The function has two for loops. The first for loop iterates over the range 0 to n-1. On each iteration, it checks if true || false is true. Since true || false is always true, it immediately returns the value 42. This means that the first for loop will only execute once.
The second for loop iterates over the range 0 to 2n-1. On each iteration, it increments i by 2 (since i++ increments i by 1 and i is incremented again on the next line). This means that the second for loop will only execute n times.
After the second for loop, the function returns the value of n.
What is the result when n=5?
If n is 5, the function will immediately return 42 when the first for loop is executed, so the function will not reach the second for loop or the return statement at the end. Therefore, the result of calling sum(5) would be 42.
What is the result when n=-1?
If n is negative, the first for loop will not be executed at all because the loop condition i < n will be false from the start. The second for loop will also not be executed because the loop variable i will never be less than 2n, which is a negative number in this case.
Therefore, the result of calling sum(-1) would be -1, which is the value of n that was returned at the end of the function.
The second for loop iterates over the values of i starting at 0 and going up to, but not including, n * 2. Within this loop is a statement that increments i by 1. This statement has no effect on the loop because i is already being incremented by the loop itself.
which is incorrect.
In my case, it gave me wrong results.
A new age of programming voodoo is upon us again. When knobs become IP ..
https://ai.stackexchange.com/questions/32477/what-is-the-tem...
Edit: tech-priests. But technopriests [1] also looks fun.