Paper Mario PC ports beckon as coder completes decompilation of the N64 classic
pcgamer.com
pcgamer.com
How is this is even remotely legal. They have literally decompiled Nintendo's binary -- up to the point the decompiled sources build to a binary with the same checksum -- and are now redistributing it. There's absolutely no interoperability motivation or any other type of DMCA exemption that would apply here. It's like claiming that the art assets are subject to copyright but not the program.
And then some site like Kotaku picks it up and frames them sympathetically, so Nintendo becomes “draconian” when it is doing exactly what it needs to do: protect its IP and business.
What I’m fascinated by is just how much free work people throw into obviously doomed projects. And now they’re doing all kinds of acrobatics hoping to find some loophole. I think this suggests just how strongly emotionally attached people can become to games they grew up with. I can definitely sympathize.
But also… if you’re this talented… make the world a new game. Make A Hat in Time if you love Banjo Kazooie so much. Make Stardew Valley if you love Harvest Moon so much!
I have also seen it in non-game projects, e.g. retro OS reimplementations.
I am not even sure where this misconception comes: that you can just decompile the original software, rearrange the source files a bit and that's it! . Even otherwise smart people actually believe this. It's incredibly damaging and pervasive thinking; by now I have to mistrust all new "miracle" contributors that come up with significant amounts of code.
Nintendo of course, wants to "sell" you the same game more than once, but it's also obvious that this is not consistent with the intent of the law.
It simply does not, specially in the US where EULAs can force you to sign away rights you may otherwise have e.g. https://www.eff.org/issues/coders/reverse-engineering-faq
But this is not reverse engineering, this is outright decompiling the copyrighted code, and then distributing it.
I'm skeptical about this idea that these people are ignorant and that these projects are doomed. Where have you seen this?
"The game's assembly code is manually decompiled into C source code."
I remember the GTA 3 decompilation project (re3) guys posting videos of themselves working with the IDA output that they got from the GTA 3 binaries. Probably also a very crappy idea from a legal perspective.
Depends on how you reverse engineer. If you just throw a compiled binary into a decompiler, rename the variables and release that, you're probably breaking the copyright law in whichever country you reside.
And in many countries of the world (including the US), not even looking at the copyrighted code is necessarily allowed. See the entire clean room thing.
Copyright is about protecting the rights-holder's ability to determine where the exact copyrighted source material can be used. Copyright does not offer the rights-holder protection for all works with equivalent semantics (that'd be a patent.) For copyright to pertain, the work has to be "textually identical", or — if we want to apply considerations about "sampling" — has to have "textually identical" sections borrowed into the derivative work.
This means that, in both the "prose text" and "source code" cases, all you have to do to avoid copyright, is to retain all the same semantics, but replace all the individual words. Note that this falls under what academia would call "plagiarism" — but there's no law that protects an IP rights-holder from having their work plagiarized.
In prose, this process of plagiarizing a work by replacing all the words, is often euphemistically called "paraphrasing." As far as the law is concerned, it is no different than actual paraphrasing — of digesting the original into your mental model and then attempting to describe the same idea in your own words. Even though, very likely, the structure of the original has been retained in your plagiarized derivative.
90% of the news articles you read are this sort of plagiarized paraphrase of some other news article — the article that "breaks the story." There is nothing the original journalist who breaks the story can do about that, under the law, because the paraphrases aren't literal word-for-word renditions of their copyrighted text.
In code, this is called "renaming all the identifiers." Which is something that happens inherently in decompilation, because all the original names for the identifiers got thrown away when the original code was compiled.
---
Also, to be clear on two issues, because surprisingly nobody in the thread has clarified this yet:
1. In these decompiled game projects, the original game's copyrighted binary-asset representations — which you'd need to byte-for-byte reproduce the game, and which you can't really "rephrase" as above — are not stored in any encoded form in the decompilation's source-code repo. Nor are any representations of assets — characters, logos, etc — to which trademark applies. Rather, for both of these cases, the build framework for the decompiled project contains logic to extract these assets from an external ROM image of the original game that must be supplied at build time. Or you could just as well supply a different ROM image — your own "asset pack", per se — to get a compiled game that is free of such assets.
This means that the decompilation project forms the basis of a tool that can be used to create a work that is byte-for-byte identical to the original — provided you have... a ROM image of the original to feed into it. You know what else matches that description? A file copy utility. (The project can also produce binaries that are not identical to the original — and that's usually the goal of the project, making e.g. PC ports of the original game, which are for entirely different target ISAs and so entirely are different object code from the copyrighted released object code.)
2. The decompilation project never publishes anywhere, the thing that it builds when you run it. Doing so would violate copyright. Just like if Rifftrax actually combined their audio dub track with an original movie and distributed that, that'd be a copyright violation. But, in both cases, it's not a copyright violation to give the end-user a tool to create a derivative work from an original, and tell the user to apply the tool to the original themselves, to derive this combined work for their own private consumption.
That reason being "paranoid lawyers who were operating on a lack of case law to interpret." We since have had many real code-IP court cases, and now have the actual established case law to refer to — and that case-law points to code copyright working akin to prose copyright (where obvious plagiarism is still not copyright infringement), not akin to e.g. musical-arrangement copyright (which is almost patent-like in its "plays the same? same arrangement" semantics, the chilling effects of which are the reason there's no such thing as an open database of classical-music scores.)
The key thing you've got to realize is that publishing a decompilation to an HLL of a work, is not the same as publishing a disassembly of that work. Assembly languages are bijective mappings to object code for the ISA the assembler targets; so of course a disassembly of object code is a derivative work of that code. But decompilation, to a sufficiently abstract HLL (which even C qualifies as), is doing literally all the work needed to create a legal paraphrase. It's replacing the structure of the code with structures available in the HLL; it's replacing the bald memory addresses of globals and stack slots with named and typed identifiers; etc. The result is "a" decompilation, not "the" decompilation. It's not a bijective mapping. In fact, by default, it's not even a surjective mapping — we do not so far have any entirely-automated decompiler (for large projects containing binary assets, like games) where the result of the first-step automated decompilation is a project that can immediately be compiled to produce the original object code. A lot more manual work is needed to reach that step.
Maybe it would be clearer if I describe it this way: you can take a program that was written in C and compiled to x86-64 object code, and then decompile it into Rust or into Haskell or any other language. You don't have to decompile it into C. And, perhaps surprisingly, this alternate decompilation will still be just as likely as the C decompilation, to produce the original object code when compiled! (Which is to say: not very likely by default, but possibly after a good few weeks of tweaking.)
Or, I can put it this way: you can take a program that was originally written in raw assembly, and decompile it into C. Decompilers apply heuristics; those heuristics don't depend on the original object code having been generated by an HLL compiler. If you employed structured-programming techniques in your assembly, then you'll get if statements and functions out in the decompilation.
If you can tell me with a straight face that you believe that a Rust decompilation of a program originally written in C infringes that C program's copyright; or that a C decompilation of a program originally written in assembly, violates the assembly program's copyright — in a world where a news article that is a plagiarizing paraphrase of another has no copyright issues whatsoever — then I'd really like to hear your reasoning.
---
A bonus analogy, just for laughs: take a modern copyrighted prose fiction work, let's say Lord of the Rings. Rename all the characters and places and such (to avoid trademarks.) Now feed the text into ChatGPT and ask it, for each paragraph, to spit out a paragraph that communicates the same information in different words. Publish the result on Kindle Unlimited. Do you think https://en.wikipedia.org/wiki/Middle-earth_Enterprises has any standing to sue you? What would their legal argument be? What would their evidence be?
(And be careful here of providing an "argument that proves too much" — any argument that would also apply to "every work of Tolkienesque fantasy published since LotR" isn't going to work. You can't sue just because there is clearly a race with equivalent properties to those of Tolkien's elves, or a sword with powers equivalent to those of one of the swords in the books — there are thousands of published books that cribbed both, not to mention the entire base setting of Dungeons & Dragons. Assuming that those cases "don't cross the line", what do you use to legally prove that this work does "cross the line"?)
And while yes, you might point out that a translation of a work of fiction into another language, is explicitly noted as constituting a derivative work. But, again, how would the copyright-holder here prove that this published work is legally equivalent to a translation? Usually such a legal argument would cite something that must remain the same between translations, e.g. uses of words that are foreign loan-words to begin with. But we can tell ChatGPT to avoid those.
My point being: when you look at the very fine boundaries of a law, they're not about what the text of the law says, so much as in what situations it can be proven that the law applies. The obvious example is an "unenforceable law" where there is no possible mechanism to detect people breaking that law. Here, we have a law that can be enforced for central cases — but becomes unenforceable "at the edges." In practice, evaluated as case-law in a Bayesian "absense of evidence is evidence of absence" sense, this means that those edges can be said to fall outside the scope of the law. If you very clearly ask a lawyer, not the question "is this in violation of copyright", but rather the question "would anyone ever bother to sue me for this, knowing they'll win?" the lawyer will answer "no."
This is already wrong. Assembly is not a bijective mapping. Different dissassemblers will generate different representations and even for the same assembler there is usually more than one way to generate exactly the same machine code.
Therefore your completely artificial distinction of dissassembly and decompilation breaks down.
> You can take a program that was written in C and compiled to x86-64 object code, and then decompile it into Rust or into Haskell or any other language. You don't have to decompile it into C. And, perhaps surprisingly, this alternate decompilation will still be just as likely as the C decompilation, to produce the original object code when compiled! This is exactly the same
Your "Rust" program would still be a copy (or at best, a derivative work) of the copyrighted program and therefore a violation to distribute it without permission.
It is ridiculous to think that because I have machine translated a GPL C program into Zig/C++/Rust/whatever I can now freely distribute the translation under the terms of the "Do What The Fuck You Want To Public License". With such a ridiculous interpretation of copyright law, why even bother with the GPL? Quick, someone tell rms!
And obviously courts agree it is ridiculous, see e.g. Atari vs Nintendo 1992: "Even though the [10NES] program was written for a different chip and in a different programming language, these similarities suggested that the Atari program was not an independent creation, but an unauthorized copy."
Apologies, I don't know much about (disassemblers for) modern ISAs; I'm mostly aware of "dumb" assembly languages used in the 70s and 80s, before macro assemblers, instruction modifier prefixes, instructions with opcodes that change based on argument types, etc.
In those early assembly languages, the instructions themselves (ignoring labels) are literally just symbolic representations of the opcodes and arguments the machine uses, in a 1:1 mapping. There is precisely one canonical assembly syntax that can be assembled by the system assembler, and it's the same assembly you see in the "system monitor" when you jump to a text section in memory and LIST it. (And, in fact, in these old systems, the "system monitor" and assembler are usually related/sharing code, such that you can modify code in memory from the "system monitor" by specifying a patch in either object-code-bytes or canonical-assembly form.)
My argument about bijectivity might not apply to modern assemblers, but it definitely did to whatever assembly language the IBM bootloader was written in.
Re: ReactOS, and Re: 10NES... well, just read the Wikipedia abstract for the case you're citing: https://en.wikipedia.org/wiki/Atari_Games_Corp._v._Nintendo_....
> However, the United States Court of Appeals for the Federal Circuit differed from the district court on whether reverse engineering could hypothetically be allowed, declaring that "reverse engineering, untainted by the purloined copy of the 10NES program and necessary to understand 10NES, is a fair use." Thus, Atari was denied the fair use exception to copyright infringement, due to the illicit way they obtained Nintendo's source code.
> One month after the decision, a similar ruling in Sega v. Accolade determined that reverse engineering was fair use. Several legal scholars have concluded that the main difference between the cases was that Atari had lied to obtain an unauthorized copy of Nintendo's code.
What happened in the 10NES case, and the thing the ReactOS team are trying to avoid, is the team gaining access to a copy of the original IP-holder's source code (to which they have no legitimate right of access.) Doing that would "pollute" the project's legal status, because if such access can be proven, then any work created by the ReactOS team could be claimed to have been created with knowledge of that original Microsoft source code, and would therefore have a presumption of being a derivative work — unless it could be proven in rebuttal that clean-room reimplementation principles were followed.
From https://www.lumendatabase.org/topics/15#QID195:
> Courts determine whether or not copying occurred, rather that the independent creation of a program, by comparing the two programs for evidence of copyright infringement. The determination of copyright infringement is done through an analysis of whether there exists a "substantial similarity" between the initial work and the product of the reverse engineering effort. Making such a determination can be quite complicated in the software context since different parts of the computer code may be similar due to the industry standards of the overall structure and user interface of programs as well as their compatibility requirements. In order to prove a claim of copyright infringement, the burden is on the initial work's owner to show that the defendant had access to the original code.
The Wiki page goes on to say:
> Legal scholars have argued that reverse engineering has since been curtailed by the Digital Millennium Copyright Act of 2000, upsetting the balance established in the Atari and Accolade cases.
But to be specific, https://www.lumendatabase.org/topics/15#QID195:
> Circumvention, according to Section 1201(a)(3)(A), means "to descramble a scrambled work, to decrypt an encrypted work, or otherwise to avoid, bypass, remove, deactivate, or impair a technological measure, without the authority of the copyright owner." Reverse engineering, on the other hand, is the scientific method of taking something apart in order to figure out how it works. While not all acts of circumvention require the use of reverse engineering, the reverse engineering of works protected by technological mechanisms requires circumvention. The placement of digital protection systems on copyrighted works essentially fences in the information a reverse engineer seeks to discover about the way the product works.
There are no "technological measures" used to protect IP rights (i.e. DRM) in these older works; so no argument about circumvention of such mechanisms pertains to them. The object-code in the ROM is the program, in a readable form, clear as day, sold to the end-user. Only the older arguments about reverse-engineering pertain.
Further, guess what "reverse engineering" refers to in the above judgement? Very likely, it didn't refer to "black-box reverse engineering" (which was intentionally impossible in the 10NES case, so the judge here wouldn't have bothered to mention it); but rather something like "decapping the 10NES lockout chip, pulling the 10NES ROM from it, and disassembling that ROM."
In other words, the implication here is that not just black-box reverse engineering, but even white-box reverse engineering provides an effective "Chinese wall" to derivative-work status, as long as the original copyrighted work (which is the source code, not the object code) is not accessed by the reverse-engineering team.
This case-law has been widely misinterpreted ever since as requiring a higher standard to be met ("clean-room reverse engineering": using knowledge of the original only to implement a specification) — but this higher standard actually only pertains when the reimplementation was informed by knowledge of the original copyrighted work (which is, again, the source code, not the object code — unless the object code "is" the source code, as in the case of the old 1:1-bijective 'canonical assembly' code I mentioned above.)
---
Of course, these arguments are all about reverse engineering with a specific pretence: namely, to ensure compatibility for some later-produced novel work that does not itself directly incorporate the reverse-engineered material.
These game decompilation projects seemingly cannot be said to be targeting this purpose. Or can they?
These projects are clearly informational in nature. On their own, they produce no useful consumable product — they require the original game ROM image in order to output... the original game ROM image.
In other words, these projects function as nothing more than a documentation of the original work. Which is an expected and allowed byproduct of the reverse-engineering process.
Again, from https://www.lumendatabase.org/topics/15#QID195
> There have been many attempts by companies over the past two decades to bring claims against software developers for their reverse engineering efforts. Since reverse engineers must make intermediate copies of the original work through the disassembly or decompilation process, the copyright owners of the initial software program have claimed that such a procedure is not covered by Section 117. They have argued that reverse engineering should be considered copyright infringement since some of the retrieved technical data used in the development process includes copyrightable expression.
> In Sega v. Accolade, the case most often referred to discussing reverse engineering of computer software, the appellate court determined that reverse engineering is a fair use when "no alternative means of gaining an understanding of those ideas and functional concepts exists." The court considered Accolade's intermediate copying of parts of Sega's video game console during the reverse engineering process in order to make compatible games of minimal significance to the rights in Sega's copyrighted computer code. The court held that forbidding reverse engineering in this context would defeat "the fundamental purpose of the Copyright Act--to encourage the production of original works by protecting the expressive elements of those works while leaving the ideas, facts, and functional concepts in the public domain for others to build on."
Which I'll sum up as: you can't take e.g. the Mario 64 decompilation, put your own assets in it, and publish that as a work — that's a derivative work. But you can freely work with the Mario 64 decompilation as an "intermediate copy" to understand Mario 64 while implementing your own clean-room work.
The only real question, in my mind, is the legal gray area around how widely one can share/reproduce/distribute such an "intermediate copy for aiding in understanding during a reverse engineering process." If the reversing project is an open community effort, then the "intermediate copy" cannot be kept secret. As far as I'm aware, no case law specific to this subject exists. All we can go on, for now, is the fact that some very litigious companies (e.g. Nintendo) have not yet sued the people doing this.
(If you think the problem comes down entirely to distribution of the decompilation, there's a neat solution to that: rather than a repo containing the decompilation per se, one could instead produce and evolve a tool which, on consuming the original ROM image, produces said decompilation, using growingly-complex set of heuristics and edge-case overrides of those heuristics. It'd be a generic decompiler program, just specialized to faithfully decompile and symbolicate a single particular program.)
I was literally thinking about assemblers from the 80s, which are _anything_ but dumb. If anything, it is _todays_ assemblers which may be dumb. Just think how many 8086 assembler kits there were (and are) and whether their syntaxes resemble each other at all. Now single MASM itself and think of how many million ways you have to create a program that builds to the same object code, whether you use macros or not. (Did you use .if or CMP?)
> it definitely did to whatever assembly language the IBM bootloader was written in.
No. Just no. We have assembly listings of the IBM BIOS. Even the first line of it may be different depending on whether you use DW or DB and still compile to the same machine code.
> There are no "technological measures" used to protect IP rights (i.e. DRM) in these older works;
This is also wrong, because _there are_. We literally discussed 10NES which is one such method for a console two generations older, and you can see the ones from Paper Mario 64 right in the decompiled source: https://github.com/pmret/papermario/blob/main/src/general_he...
> Re: ReactOS, and Re: 10NES... [incredibly-long-wall-of-text-that-you-keep-editing-and-moving-from-topic-to-topic-so-that-no-one-has-a-chance-to-interrupt-it-is-really-quite-hypnotic]
No one is claiming that reverse engineering is illegal. What is claimed is that decompiling a protected work in full and then distributing that decompilation is definitely illegal.
You can definitely decompile for reverse engineering. I have done it as part of my RL work, at least in the EU (which definitely clarifies that you have a right to decompile and _even copy_ the parts that are strictly required for interoperability. Not the entire game). But it is ridiculous to think (and I emphasize: ridiculous) that just because I apply a simple transformation to a protected work I can end up with a non-protected work that is legal for me to distribute without authorization.
It would mean the end of copyright for computer programs as we know it.
In fact, it is so ridiculous, that even since the act of _compilation_ is also a non-revertible transformation (or in your terms, non-bijective), the actual object code would NOT be subject to copyright ! (Because, despite what you're saying, what is copyrighted is the actual source code, with the object code being a derivate work).
I am not going to pronounce myself on it being or not a copyright violation in any and all countries of the world. As most of the people having strong opinions here, I’m not an expert nor that knowledgeable on the question.
Impressive feat of decompilation anyway.
If to do that you would had to dissasembly the ROM of the Nintendo 64, or of some protection routine in the game, or you were writing an emulator (i.e. another program to interoperate with), it _may_ have been fair game. However looking at the code of the very game you intend to create clones of is not going to fly in any jurisdiction I know of. Much less literally creating the cloned games by directly copying the original game's code in its entirety. There is just no way to justify this.
I am eagerly awaiting the jurisprudence supporting that. Because if you have nothing to quote, giving your opinion very forcefully doesn’t magically make it more than your opinion.
Armchair layers are in full swing today. I am very curious to know how all of you know so well where the border between what is and isn’t acceptable is.
Yes, sorry if I tend to at least initially trust the opinion that does not _completely_ destroy all software licenses in place today. This is a forum where software developers comprise a large % of the audience.
> "I am eagerly awaiting the jurisprudence supporting that."
I say that if I castrate people without their consent it is OK, as long as I only do it using properly sterilized scalpels in accordance to current medical practice.
You say I can't. I am eagerly awaiting the jurisprudence supporting that.
You have nothing. I am not aware of this precise subject having been put to the court ever and the law is far from clear. Contrary to you, I'm not pretending to be a subject expert on things I know nothing about.
> I say that if I castrate people without their consent it is OK, as long as I only do it using properly sterilized scalpels in accordance to current medical practice. You say I can't. I am eagerly awaiting the jurisprudence supporting that.
I'm not even going to entertain this level of ridiculousness.
That seems fair to everyone but that doesn’t necessarily say anything about if it’s legal
This is like not being allowed to burn CDs on your PC because you can potentially use them to burn a copyrighted .ISO. I think that media companies tried to take that away in the VCR days but failed.
Another similar situation: ID Software open-sourced their DOOM code, but not the art assets. So you have a bunch of DOOM engines, including enhanced ones like GZDoom, BOOM, etc. Which you can download, but you only get the engine and none of the WAD files which contain the levels, sound and graphics.
So you can then use freely available assets like Freedoom or other packs that are out there, or use ones from your existing legal copy of DOOM.
Especially given the fact that we can generate and infinite training set:
High level code -> compiler -> assembly code
All the AI's got to do is learn the inverse map.
All you have to do is grab all of github, feed it to random compilers, grab the assembly output.
Boom : instant training set for a decompiling AI (and one that can recreate original code from assembly, hopefully complete with comments and meaningful variable and function names)>
Fresh from HN:
Instead of behaving like an unproductive ass, how about you do something useful and provide a link to folks working on the problem?
Do they guess it based on what they end up thinking the function/variable does in the code?
Soon: Freedows 12. "Your Honor, I just looked at the compiled code, so it must be considered clean from copyright claims!".
ReactOS and Wine exists and are reimplementations of copyrighted code.
Good luck trying to prove to any judge in the world that you came up with an implementation that happens to be exactly the same as the original game code, down to the same checksum, by pure chance.
What is this, "the ChatGPT defense"? Do you really think that code has intrinsically less right to copyright than art, or what?
>ReactOS and Wine exists and are reimplementations of copyrighted code.
Clean room reimplementations that don't even dare to look at the copyrighted code.
> Good luck trying to prove to any judge in the world that you came up with an implementation that happens to be exactly the same as the original game code, down to the same checksum, by pure chance.
Not impossible if you reverse engineer the software and look how the binary code is behaving, emulators does this all the time.
A reimplementation of code that leads to the same binary output isn't covered by the copyright, as long as you don't share copyrighted assets or the actual binaries. What people do with the alternative source code is their business, you can't hold responsible that person for others' action.
We are not talking about a program that "produces the same binary output" when executed. We are talking about a program that _is exactly the same_ as the original one, bit-by-bit.
It is ridiculously impossible to "independently" discover a program that is _exactly the same machine code_ as the original program by just looking at the inputs and outputs of the original program. (Barring trivial programs, of course.)
This is not reverse engineering. This is just dissasembling and decompiling.
Do you think that I can just gzip gplprogram.c and then redistribute gplprogram.c.gz under the "Do What The Fuck You Want To Public License"?
As if that would make a difference in how big of a copyright violation all of this is.
I’m sure people have done this, I have the leaked code somewhere so I could do it if you want
If you start at the entry point of the game and trace every memory access, and you run the game in a profiling and debugging emulator, you can figure out what every variable does. It’s a long and painstaking process. Definitely one you probably don’t embark on unless you really care about the result (or you’re well paid).
Some variables are easily guessed because they refer to a string that's stored in plain text. "%d:%s" is likely to be a printf format string, so you can follow that lead and find out which arguments correspond to %d and %s. Sometimes you get lucky and the debug message tells you directly something like "Invalid value of the coin counter: %d". This is a dead giveaway for what the argument corresponding to %d is.
Other kinds of variables can be guessed by using a debugger and looking at their values, or patching code to modify them and watching what happens.
Trained on enough code, couldn't it figure out reasonable enough names from the structure of the code?
On the other hand, it might not work at all if short functions tend to be incredibly generic, and all the real meaning is tied up in the nonlocal organization of it all.
And then of course you need context to know what the nouns and verbs are at all. I can imagine similar structures involving Luigis and Koopas, or invoices and payments.
I used to be somewhat active on a ROM hacking message board, and there were rules about how users could share their modifications. Users had to share a patch. Uploading or linking to prepatched or original binaries was prohibited, and moderators would routinely intervene and remove them. Having been exposed to such measures to avoid action from copyright holders in the past, I just don't understand why it is so popular to just post the code openly.
Yes, it’s not really “emulation,” but the UX is basically the same.
I also imagine emulating the game was never really a big problem. So this was more just for fun/passion?
I’m thinking about how SMW romhacks basically inspired (but competed with) Mario Maker.