I own a copy of IDA (legally). It was an absolute pain to purchase and it seems that a large portion of their margins are dedicated to piracy control. I won't detail the process...but it seems unusually personal.
If I had to guess they will expand their decompilers (the actual flagship project). It will be years before Ghidra + a community catch up to them and Binary Ninja (I also own a copy of it) may never. The disassembler is just a familiar tool. Their decompilers are way, way far ahead.
Can you elaborate? I would really like to see what the HexRay decompiler does better (or worse) than Ghidra, but I am too poor to buy it (and dogbolt.org is not interactive, so I cannot edit function signatures to "help" the decompiler etc.).
Is it better in general, for a specific programming language or platform (e.g. C++, Windows), or for a specific use case (e.g. obfuscated code)? I heard about their cloud-based stuff, although I don't know what they are exactly doing there. Maybe ML trained on source code? Function signatures of the latest malware?
After several hundred hours with Ghidra, I think it certainly would need some polishing, in particular:
- UI. Too many frequently-used dialogs are not optimized for keyboard usage.
- Decompiler too stubborn sometimes, ignoring user input (e.g. manually specified types).
- Decompiler needs better heuristics for the treatment of some common cases (e.g., often doesn't recognize for-loops and array accesses)
- Quite dangerous: Sometimes the decompiler gets lost, especially if a function contains handwritten assembly code with unusual control flows. Okay, can happen. But instead of displaying a warning it just shows you the part it could decompile and you have to figure out by yourself that something is missing.
But most of the above issues are fixable. Instead, I would be interested in learning about more fundamental differences between the two decompilers.This is the thing that sticks out the most IMO. IDA decompiler is quite a bit more flexible than Ghidra's. When you assert a type, it will usually not ignore it. It may sometime get a bit lost if you give conflicting types to dependant variables, but otherwise, it's pretty good at this.
One of the annoying bits of ghidra (though it may have improved, it's been around a year since I last used it) is that there's no way to "split" a variable. Sometimes, ghidra will have some code that looks roughly like:
int x;
x = 0;
doSomething(x);
x = 1;
doSomething2(x);
(Obviously, really simplified).The problem is, sometimes, x needs to be an int for the first function, and a bitflag structure for the second. But Ghidra has no way to say, "hey, from this assignment on, treat `x` as another variable", so you have to either generate a union (ugly) or deal with sending the wrong type (also ugly).
IDA tends to be much better at this. All variables start out as "split" as it can make it (almost in static single assignment form). Then, the user can tell it "Those two variables are actually the same, please treat them as one". I find this flow works really, really well.
Right click -> “Split Out As New Variable”, but it seems like this doesn't work for stack reuse yet (just registers, or more generally, simple varnodes).
https://github.com/NationalSecurityAgency/ghidra/issues/2573
Also can I haz offset pointers pretty please? It's not in the last released version I tried, and the last Git version I tried had offset pointers but they didn't affect decompiler output so multiple inheritance and container_of linked lists still came out broken.
One issue with Ghidra's that I keep hitting is its poor support for amd64 SIMD. There's a good example at <https://github.com/NationalSecurityAgency/ghidra/issues/249>:
0000000000000000 <intrinsics>:
0: f3 0f 1e fa endbr64
4: 0f c6 ca 1b shufps $0x1b,%xmm2,%xmm1
8: 0f 58 c1 addps %xmm1,%xmm0
b: c3 retq
IDA produces great decompilation: __m128 __fastcall intrinsics(__m128 a1, __m128 a2, __m128 a3) {
return _mm_add_ps(a1, _mm_shuffle_ps(a2, a3, 27));
}
The Ghidra output is a mishmash of CONCAT pseudo-macros.Due to these heuristics IDA produces real actionable code quicker. My experience with Ghidra (less than yours) is that it produces mostly garbage on a lot of different things and it requires a lot of prep work to make truly usable. This might not be noticeable on small or simple binaries but on larger binaries it actually becomes a real measurable problem. While Hex Rays isn't perfect, it's about as close as we can come to it right now and it generates very human-looking code with smart optimization removal. One thing I remember with Ghidra not long ago was a common optimization like using SSE registers for arrays would produce a page worth of non-sense for something simple. Additionally, detection of standard libraries still isn't good so you end up wasting your reversing time on re-reversing a different compilers version of strlen than actually doing the work you need to do. If you could use FLIRT signatures in Ghidra legally I'd imagine Ghidra would be vastly improved.
I'm not a reverse engineer, but I have been working on this with the assumption that it would be a good time-saver for a reverse engineer. Feel free to DM me if you'd like to discuss.
I know what you mean. I tried to purchase it and got this email:
Dear Sir/Madam,
Thank you for your order. Please could you send a copy of your passport and fill out the attached form? Our compliance policy now requires this.
Many thanks
Hex-Rays SA
I used a cracked copy after that.I don't have formal data, but in my narrow slice of the world that's what I'm seeing as well.
IDA's price is so high that it's easy to justify using Ghidra instead. Heck, if Ghidra does something less well, it might be cheaper to pay to improve Ghidra (and then you can use those improvements forever). I encourage organizations who are thinking of using Ghidra to contribute back to it; if those improvements get integrated back in, then those improvements will continue into the future along with other improvements.
I guess there are a lot of people that just click around in the UI but for a keyboard based workflow, Ghidra has a lot of catching up to do even to IDA 4.x.
I like Ghidra's interface better, but I'm also not a keyboard-first type of user. I work with a lot of different OSes and software. I stopped trying to remember most keyboard shortcuts 10+ years ago, because there were too many variations, and the consequences of using the wrong one can be dire.
Releasing and open-sourcing Ghidra was a truly magnificent gift by the NSA, and I can't thank them enough for it.
[1] I'd love to see a Ghidra equivalent of Lumina, for example.
Arguments of keyboard/mouse efficiency are dead to me. The limiting factor of any reverse engineering is definitely not how efficiently you can keyboard navigate a UI. People are too in love with efficiency of the wrong things and not what the real limiting factors of reversing are, which is comprehending the program. I don't love clicking around but it has never once slowed me down on a project.
I'm guessing that I'm on the same pirate blacklist that a lot of people landed. I guess I should have expensed it instead of having the Company buy it. I was the only person doing RE work, it was a single user person license and I was also the only linux user out of that 300 people org.
I think there are a lot of people that would like to pay for IDA but can't get it.
On the other hand I really don't like the Ghidra user experience.
Is there something similar to FLIRT in r2 or ghidra?
Just to clarify, if I wasn't concerned with budget and had to pick a product today, I'd still go with IDA, but the mindshare among new hires and interns is swinging pretty rapidly towards Ghidra. If they are able to continue adding features (with US government funding, and open source development), IDA is right to be worried about whether they will still be the best choice in 5 years.
beware that these are not the latest versions.
Many years ago we wanted to buy IDA Pro license and were quickly declined because we had whois privacy protection enabled on our domain.
You can create your own signature databases, although i have been unsuccessful in my brief attempts.
The algorithm is different from flirt though.
Its roughly, mask a bunch of stuff (relocations, etc), and hash whats left. There is some parent/child analysis for ambiguous calls i believe.
My recollection is flirt also masks its version of a bunch of stuff, but builds a trie of the first xx bytes with some provision for needed values past that, and some parent child analysis. They have a paper on the algorithm you can google.
Ida ships with more signatures out of the box, In my experience. Although i havent seen them listed anywhere.
I think it is safe to assume that offering cheaper products will be the last thing they are going to do.
But then again those take time to show their benefit, and they probably make piracy easier.