De-anonymizing programmers from executable binaries
blog.acolyer.org
blog.acolyer.org
I found it was easy to get mind-blowing results on a single dataset - e.g. 99% accuracy for working out who wrote which emails in the Enron [0] dataset.
The moment the test set was even slightly cross-genre, results became bad.
So this is great in theory, but I would love to see a test set from the same authors outside the context of the Google Code Jam. I'd be surprised if the results were anything like as good.
There are very few real-world cases where you have a large amount of same-genre data for the questioned author. (No one writes a thousand suicide notes / ransom demands / scams)
[0] A large email database often used for authorship tasks https://www.cs.cmu.edu/~enron/
Incorrect back testing is absolutely standard with stock/bond/forex/etc trading. There are a million hustlers who sell their back-tested systems. In essence, without forward testing all you have is an untested hypothesis.
Symbols and strings are free-form text that would seem to leave a lot more of a stylistic fingerprint. If you could get this level of accuracy from machine code alone, I would be very surprised.
The article says that stripping symbols reduced accuracy by 24%. This is a far greater drop than from optimizing the binary. It leads me to believe that a large part of the success in classifying is coming from symbols and strings.
Everyone has their own distinctive personal library/template/boilerplate that they are pasting for every problem since includes, typedefs, defines, library helpers, IO, etc are always the same.
You can see some examples here to see what I mean: https://www.go-hero.net/jam/17/solutions
In practice, it might be interesting to give a fuzzier output - the esimated probability that an individual wrote the code for a certain binary. This would help when two programmers (say, Alice and Bob) write code that yields binaries with similar properties. In the current system, either Alice or Bob is picked -- there is no way to express that it might be either of them.
In this case I would expect that we see a multitude of code styles in the source code of the very first publicly released version of the Bitcoin client. But I think nobody has yet pointed this out. This to me at least implies that if there were multiple programmers involved, there was at least someone who "ironed over" the complete source code to make the code style much more homogenous. But under this hypothesis we would expect that the code was audited at least somewhat thoroughly. But in fact the first versions of the Bitcoin client were rather "put together hastily" and contained lots of bugs that were fixed afterwards. Look for example at the following comment:
> https://github.com/bitcoin/bitcoin/blob/595a7bab23bc21049526...
So first versions of the Bitcoin client rather look to me like some solo person who can program, but at least whose job is not to program in C++ every day, put together an experiment (the Bitcoin client) hastily.
Now let us assume that there was a huge team behind, but only 1-2 (at most 3) programmers with very similar code signatures. But of course the other team members also had things to do (to be worth their salary), like working out the math etc. But if it were a huge team that contains people of such special skills, why the hell did they hire such a bad programmer to implement it?
In this sense I believe there is strong evidence to drop the "huge team hypothesis".
The internet is a vast network where the individual components don't understand the whole, just like a brain is. How do we know the internet itself hasn't turned into a super-intelligent AI already?
Al dente, baby.
Would that be the case if they used software that enforced a common, coding style on the team? And on top of the team putting in at least a little effort to keep consistent? There's quite a bit of that software already used in big organizations, including defense sector.
My bet would rather be on Dave Kleiman
> https://en.wikipedia.org/wiki/Dave_Kleiman
His death 2013 fits well with the public disappearance of Satoshi Nakamoto. Also Dave Kleiman worked in computer forensics and was from very early on involved in the Bitcoin community.
It is well-known that Dave Kleiman and the "self-claimed Satoshi Nakamoto" Craig Steven Wright knew each other (Craig Steven Wright claimed that they collaborated when creating Bitcoin). Also the fact that Craig Steven Wright could convince people that are deeply involved in the Bitcoin community in a one-on-one interview that he is Satoshi Nakamoto, provides strong evidence that Craig Steven Wright knew lots of obscure things about Satoshi Nakamoto "that only Satoshi Nakamoto should be able to know". If he knew this information from Dave Kleinman, this can be plausibly explained.
http://lesswrong.com/lw/jgz/aalwa_ask_any_lesswronger_anythi...
Specifically this comment http://lesswrong.com/lw/jgz/aalwa_ask_any_lesswronger_anythi...
I have always thought the speculation about SN's identity suffered alarmingly from availability bias. Everyone wants to believe he's someone they know about. Someone who's extremely well-known for working on or writing about things extremely similar to Bitcoin. There are far more smart nobodies tinkering with things than there are top names working on those things.
Also suffers from a bias I'll call "Shakespeare didn't write Shakespeare bias." The idea is that a notable thing must have been written by a high-status and well-credentialed person, and could not have been done by a talented no-name.
I think your reasoning is interesting. Nevertheless consider that if we hunt for SN, it is very likely that it is at least a person who knows more than a little bit about various topics that were necessary to implement Bitcoin (e.g. finance, cryptographic hashed data structures, existing protocols for crypto cash, ...). I agree that it is quite possible that this could also be a talented no-name, but I think this property provides a strong criterion to check if someone claimed that some specific person (who is a talented non-name) they have in mind is SN.
Also one would have to provide an explanation why such a person is not more well-known. It can be that he/she has a difficult personality. It can also be that he/she lives in a country where such a person will have a lot less opportunities etc. But these constraint again limits the circle of talented no-names that could be SN.
TLDR: Your reasoning is interesting, but your hypothesis has strong consequences.
I don't think much has changed.
"Yeah, it's like a hacker wrote it."
"That's ill, man. It's incomplete!"
http://listserv.linguistlist.org/pipermail/ads-l/2005-April/...
https://web.archive.org/web/20100820000534/https://www.wnyc.... (Search for 'fist')