Dogbolt Decompiler Explorer
dogbolt.org
dogbolt.org
As recently as last week we had some horrible performance problems but it looks like the queue (https://dogbolt.org/queue) is mostly still fine! Other than the long pole of a few of the decompilers being backed up, things are humming along quite smoothly! Josh + Glenn have done some great work on it! (https://github.com/decompiler-explorer/decompiler-explorer/c...)
Decompiler Explorer - https://news.ycombinator.com/item?id=32079227 - July 2022 (82 comments)
I ditched Ghidra in my experiments in favor of angr early on because Ghidra did not play nicely with multiprocessing and I had a lot of data to process. Well maybe it does but it was much easier for me to achieve the same thing with angr.
Love the name! Although I feel compelled to point out that Compiler Explorer is the name of the project and Godbolt is its author's last name, but I suppose if people are to the point of using Godbolt as a verb the ship has sailed.
My work focuses on recognizing known functions in obfuscated binaries, but there are some papers you might want to check out related to deobfuscation, if not necessarily using ML for deobfuscation or decompilation.
My take is that ML can soundly defeat the "easy" and more static obfuscation types (encodings, control flow flattening, splitting functions). It's low hanging fruit, and it's what I worked on most, but adoption is slow. On the other hand, "hard" obfuscations like virtualized functions or programs which embed JIT compilers to obfuscate at runtime... as far as I know, those are still unsolved problems.
This is a good overview of the subject, but pretty old and doesn't cover "hard" obfuscations: https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=1566145.
https://www.jinyier.me/papers/DATE19_Obf.pdf uses deobfuscation for RTL logic (FGPA/ASIC domain) with SAT solvers. Might be useful for a point of view from a fairly different domain.
https://advising.cs.arizona.edu/~debray/Publications/generic... uses "semantics-preserving transformations" to shed obfuscation. I think this approach is the way to go, especially when combined with dynamic/symbolic analysis to mitigate virt/jit types of transformations.
I'll mention this one as a cautionary tale: https://dl.acm.org/doi/pdf/10.1145/2886012 has some good general info but glosses over the machine learning approach. It considers Hex-rays' FLIRT to be "machine learning", but FLIRT just hashes signatures, can be spoofed (i.e. https://siliconpr0n.org/uv/issues_with_flirt_aware_malware.p...), and is useless against obfuscation.
Eventually I think SBOM tools like Black Duck[1] and SLSA[2] will incorporate ML to improve the accuracy of even figuring out what dependencies a piece of software actually has.
[1]: https://www.synopsys.com/software-integrity/software-composi...
[2]: https://slsa.dev/
> My take is that ML can soundly defeat the "easy" and more static obfuscation types (encodings, control flow flattening, splitting functions). It's low hanging fruit, and it's what I worked on most, but adoption is slow.
If I wanted to implement my own toy HexRays-like decompiler using a few of these techniques to decompile x86-64 binaries is there any high quality up-to-date paper/resource you would recommend?
Or do you think that "A Generic Approach to Automatic Deobfuscation of Executable Code" paper is a good enough start?
Also, what do you think about https://tigress.wtf/ ?
Might also be worth considering an approach integrating LLMs for summarizing code. Maybe you could fine-tune a pretrained model that already "understands" source code to associate sources with generated code? If going this route I would still probably use a disassembler to preprocess, and maybe also extract basic blocks to use as my "target" domain for fine-tuning.
As for Tigress, I used it extensively and found it to be really great most of the time. There are some limitations to be aware of: it only works with C code, and you have to turn your multi-file projects into a single file with a main() function. Also, its C parser (CIL) has some limitations (e.g. doesn't recognize the word static in "struct foo x[static 1]") so you might need to translate your C code first. I translated manually because it was a really rare issue for the code I started with. I also had mixed results using Virtualize and JIT. Sometimes they would emit invalid code, so I ended up just throwing out that data.
In my view, the up-and-coming Tigress challenger is obfuscator-llvm. I think it is very promising for future work because it inherently supports more languages than only C. But currently obfuscator-llvm is much more limited (~3 transformations compared to ~48). So if you're using C, today I would pick Tigress.
As far as I know, BNIL (https://docs.binary.ninja/dev/bnil-overview.html) is the only one that is designed to be readable and it still wouldn't make sense to include it in an IL comparison such as the one done here for decompilation in my opinion.
Though for me the big selling point on Binja is the Intermediate Languages (ILs). HIgh-level IL is the decompiler but you also get Low-level and Medium-level ILs as steps between assembly and source. If the decompiler is a bit funky you can look at the ILs to get a better idea of what is happening. the ILs are also just much nicer to read than plain assembly so I tend to use them a lot.
Its a feature that isn't really matched on any other platform. Ghidra and IDA both have a single IL that is more machine readable compared to Binja's human-readable ones.
https://hex-rays.com/ida-free/
The only thing you lose out on is Python scripting, which is kind of big, but for a free tool you really can't complain.
You probably want to use both IDA and Ghidra since they have different strengths/weaknesses and community plugins.
[1] https://docs.binary.ninja/dev/bnil-overview.html [2] https://riverloopsecurity.com/blog/2019/05/pcode/ (an example)
- Anyone care to give some pointers?
Theoretically, a fixed point should be reached.
angrily writes a letter to his congressman who won't understand a word of it
> Vector 35 and Hex-Rays jointly sponsor the hosting on Digital Ocean as a community service.
Of note: HexRays is not only cleaner, but right now their queue is mostly empty while others are backed up.
And it's no conspiracy theory or intentional sandbagging, you can see the implementation: https://github.com/decompiler-explorer/decompiler-explorer
and if anyone can improve the other tools performance we'd be happy to accept it. We reached out to the Ghidra devs: https://github.com/NationalSecurityAgency/ghidra/issues/5228 but they didn't have any silver bullets for us either.
Impossible, sadly.
oof
> In short: your source code is stored in plaintext for the minimum time feasible to be able to process your request. After that, it is discarded and is inaccessible. In very rare cases your code may be kept for a little longer (at most a week) to help debug issues in Compiler Explorer.
1. If a third-party does their link-shortening, which gets the program text, then - it doesn't matter how nice they are. And if that party is Google then, well...
2. The language you quoted still allows them to keep effectively all information through mining aspects of it rather than keeping the entire code as a stretch of plain text.
3. If GodBolt or its servers are subject to US law, then there might be National Security Letters which compel it to pass information on to the US government, and keep that secret. And this is not a conspiracy theory, this what Snowden has exposed about Google, Apple, Microsoft, Yahoo etc.
So - I respect and like the GodBolt'ers, but you don't have a good guarantee of your data being kept private.
One was a CP/M-86 "small model" executable, the other an object file (16 bit intel OMF object file) - i.e. compiler output.
Boomerang looked like it'd have most chance of getting somewhere since it mentioned having a DOS .EXE analyser.
I'm surprised that "Hex Rays" (i.e. IDA-Pro) got nowhere...
Dogpile was an automated way to search all of the search engines at the same time with one query.
https://web.archive.org/web/19990429194414/http://dogpile.co...
> It's meant to be the reverse of the amazing Compiler Explorer.
With a link to https://godbolt.org/
It’s very obvious that Dogbolt Decompiler Explorer is primarily named after Godbolt Compiler Explorer.
I thought for a longtime it was some joke I wasn't getting related to deities smithing people.
I thought it was the sibling part to the Jesus Nut. https://en.wikipedia.org/wiki/Jesus_nut
To a far far lesser degree, I’ve experienced many examples of “you named it X but everyone at work calls it Y and now you have to live with that.” It used to really irk me for some reason.
As a data point: Search on stack overflow yields "500" hits. https://stackoverflow.com/search?q=godbolt
OTOH the godbolt domain is at least not actively used for a number of other TLDs getting one of those might be an easier option.
I always call it the compiler explorer but the url, as a sibling comment says, is memorable.
That's "deities smiting people.", but I really like the idea of deities smithing people :)
Now that reminds me of a verse from a song I heard on the radio as a teenager:
Had a meeting with my maker
The superhuman baker
He popped me in the oven
And set the dial to lovin'That's St IGNUcius to you.