Anonymous programmers can be identified by analyzing coding style
freedom-to-tinker.com
freedom-to-tinker.com
I remember teaching a Matlab class for engineers and scientists that was about 50/50 male/female, and the women tended to have much neater code. Code written by males often had comments all jumbled up, inconsistent number of spaces between braces/operators/etc, incoherent variable names, worse names for functions, and so on.
During the telegraph era, boys often hung around outside the office, as they could earn a bit of coin delivering telegrams.
With the introduction of the phone, they tried using those boys as switchboard operators. Problem was that they would prank the callers by either cross wiring calls or unplugging them mid-call.
So instead women were hired, as they didn't do that. Instead they would listen in on calls, and so became a source of gossip.
This was seen as an acceptable trade-off by the companies, as at least the calls where properly and reliably routed.
Because of that, it seems like Code Jam is an artificially easy test case for this sort of identification -- I'm pretty sure a human could look at my solutions and conclude they were obviously all written by the same person.
I'd really love to see someone like grugq weigh in with some thoughts here. The idea of having tools to parse and rewrite my code to be as generic as possible came to me years ago.
I feel like awk could do this quite handily, as far as homogenizing spacing, cases, underscoring, etc.
You basically want an obfuscator that replaces all the names of things with generics, and then randomly permutes blocks of code without changing the code paths possible in the final binary. (Perhaps some optimized-for-performance version of this, but that might identify the tool you use.)
It sounds relatively easy to write if you stick to certain coding guidelines (like using techniques amenable to static analysis).
However, this still won't work in some cases, because you'd need more advanced tools to handle profiling of what sized functions and such you ended up writing.
It would be interesting to try and write a tool which defeated any analysis of author patterns in the code, but would require understanding the program across the boundary of function calls, which is a difficult problem. (You probably couldn't write Turing complete code, for example.)
Here's a Schneier post about a Concordia University study about identifying e-mail authors. https://www.schneier.com/blog/archives/2011/08/identifying_p...
> We used a combination of lexical features (e.g., variable name choices), layout features (e.g., spacing), and syntactic features (i.e., grammatical structure of source code)
The AST stuff is super interesting. The other signals are somewhat superficial. But comparing ASTs? That is deep.
PS: 1] shows that using AST's does does not get you THAT much of entropy gain compared to other features.
1: http://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=6234420
What I understood from that - it worked quite well with code bases like from the Google Code Jam (large LoC, no style guides etc), but not that well with smaller amounts of code and I'm looking forward to some additional results e.g. with a codebase from a corporate development environment.
At one job I had we called uncommented, poorly formatted code "curtis code" for some reason.....
"We used a combination of lexical features (e.g., variable name choices), layout features (e.g., spacing), and syntactic features (i.e., grammatical structure of source code)"
In particular, "layout features" is a huge issue in some languages, and not at all in others. For instance, a language like Javascript, or PHP, give great flexibility about layout, so in those languages I can see each developer having a unique style (and I have been involved in style debates regarding those languages), however, a language like Python has a fairly fixed layout, since the whitespace is significant. And also, in Clojure, I think most programmers use Emacs and accept the Emacs clojure-mode indenting as the default.
Variable name choices is another where some environments encourage similarity, and others allow for unpredictability and unique styles. Within the Ruby On Rails framework, for instance, there are norms about the creation of variable names.
I would guess that syntactic features is perhaps the one characteristic that shows a great deal of uniqueness in every language. I am often surprised at the choices my fellow co-workers make, when it comes to how to solve a problem.
Off the top of my head:
* When / if single-line if / while / etc statements are used * How many blank lines are used between functions * How often blank lines are used in functions * How much indentation is used for initializing lists / etc. * If multiline strings are used.
Etc.
Some of these are covered by PEPs, yes, but enough people don't follow PEPs religiously that even those offer some information.
Good practice is to follow a very explicit coding style which makes code written by different developers indistinguishable - the more the better.
Go ahead, identify which developer wrote which part of Linux kernel or, god forbid, jdk/src/
Does anyone know of any such projects, or what they might be described as? None of the queries I've tried produce the intended results.
You could even take it one step further, if you can identify the author of source code, can you not then forge that signature to make it look like they wrote something they didn't?
I guess it's because you're looking for "obfuscators" while such programs are usually known as "automatic code formatters"... and any decent IDE is going to have the functionality to do this.
For something standalone, look at http://en.wikipedia.org/wiki/Indent_(Unix)
It's hard to describe, but you kind of get the knack of spotting trends after a while. The accuracy's kind of poor, though. There are a lot of programmers in the world; if someone uses and reuses very distinctive routines, fine, but otherwise you're only really going to get a general impression.
You probably won't find Satoshi this way.
And then you have to justify your closed-world assumption: how do you know Satoshi (under his real name) was even in your dataset? Maybe after Bitcoin he went back to closed-source work or commercial projects, and none of his source code other than Bitcoin appears in your dataset. Then the guy your analysis picked out isn't 'Satoshi' so much as 'the guy who looks the most like Satoshi (but actually isn't)'.
Color me unimpressed