Technical Time Travel: On Vintage Programming Books
tiemoko.com
tiemoko.com
> "Clean-room design" was an underhanded way to legally reverse engineer and clone a competitor's product. It works like this: engineer A produces a specification after studying the competing product, a lawyer signs off on the spec not including copyrighted material, and engineer B re-implements the product from the spec A created. A and B have the same employer, but since they're not the same person there's technically no copyright infringement. This technique was used during the fiercely-competitive market rush of early personal computing.
What a weird way to position this.
There's nothing "underhanded" about it. There is literally no copyright infringement in this case. It's pure reverse engineering as is done in many other industries and fields.
And the outcome is increased interoperability and improved competition by eliminating artificial barriers of entry in the market. Without it, the computing world would be a very different place, and I doubt anyone would like it.
In fact, the market was far from "fiercely competitive" prior to the IBM BIOS being reverse engineered. Before that work was done, IBM basically had a lock on the PC market. It wasn't until the clone rush of the 80s that prices came down and computing became accessible to everyone. Hell, IBM valued that near-monopoly so much they introduced the MCA bus in the hopes of locking the PC back down. Fortunately their competitors succeeded in establishing open alternatives, including PCI and so forth, and the rest is history.
I wonder how this author feels about the Google v Oracle lawsuit regarding Google's reimplementation of a bunch of core Java APIs...
https://en.wikipedia.org/wiki/Extended_Industry_Standard_Arc...
As a former field engineer I was glad to see the back of MCA.
IIRC it was considered fashionable back in the day to install a card, wait for it not to work and then make a joke about "plug 'n pray" before finding an older, ISA card and getting it to work by setting the jumpers correctly.
I assume "clean room" BIOS implementations were important to avoid copyright infringement claims since the BIOS source code was actually provided by IBM in the system's technical reference manual (similarly to other systems of the time):
https://retrocomputing.stackexchange.com/questions/12018/why...
Shame that modern vendors don't provide their system firmware source code.
I don't know if this is a real thing or not (I can see it being real in smaller companies, but incompetent when it comes to these things so what do I know), but that's the case described.
So long as no actual code was copied, no, there isn't.
An independent re-implementation of an API/interface/whatever for the purpose of interoperability is perfectly legal. This explains the clean-room approach, as it ensures that the developer re-implementing the functionality has absolutely no access to the original code, thus making it impossible for them to engage in copyright infringement (BTW, if you want to watch a fantastic dramatization of this, go watch Halt and Catch Fire. I love the show for a lot of reasons, but their dramatization of the reverse engineering of the PC BIOS was just... so good).
This is why, for example, LibreOffice can go ahead and implement their own Word doc implementation without worrying about Microsoft suing the pants off of them.
It's also why Wine is perfectly legal (though not without controversy when developers have tried to take shortcuts).
Heck, we wouldn't have AMD if this type of thing was illegal! AMD got into the microprocessor market by, yup, reverse engineering the Intel 8080 and cloning it:
https://en.wikipedia.org/wiki/AMD_Am9080
> It was originally produced without license as a clone of the Intel 8080, reverse-engineered by Ashawna Hailey, Kim Hailey and Jay Kumar by photographing an early Intel chip and developing a schematic and logic diagrams from the images.
Ah, now we're getting into a completely different forms of IP protection.
The Ford Mustang badge is a protected trademark, so you're not allowed to use it without their permission.
The car itself might also contain patented technical or design elements that you may not be allowed to duplicate without a license.
Broadly there's four classes of IP protection out there: copyrights, patents, trademarks, and trade secrets. This discussion vis a vis reverse engineering is primarily centered around copyright protections.
As for the rest of your comment, check out my other reply where I hopefully clarified things.
https://en.wikipedia.org/wiki/Software_patent#United_States
One has to wonder if, today, the reverse engineering of the IBM BIOS would be possible, given the likelihood that IBM would've patented their implementation.
Software developers independently come up with the same solution to common problems all the time! We even have names for it: algorithms, design patterns, etc.
If the same person both reverse engineers an existing implementation, and then writes a new one, then that kind of incidental duplication could look like copyright infringement, and the only defense is the developer saying they didn't do it. That's pretty tough to prove.
A clean-room approach solves this problem. Parallel re-invention of the same solution can easily be proven because the developer genuinely never saw the original implementation in the first place, and a lawyer examined the reverse engineered specification to ensure no copyrighted material was contained therein.
Suppose I've developed an API that has a few function calls:
void foo(int)
int bar(char **)
char *baz(float)
I have then written a bunch of code that defines my specific implementation of 'foo', 'bar', and 'baz'.
That specific implementation--my specific code--is subject to copyright and no one is allowed to copy my code and use it without my express permission. And that includes obvious things like copying the code and obfuscating what you did via renaming of variables and so forth.
But suppose you want to implement your own compatible version of that API so that someone else can use your library instead of mine.
To create your version you decompile the code and you see that the API is composed of those three functions 'foo', 'bar', and baz'. You then read the decompiled versions of those functions and you see what they're logically doing.
You absolutely can then go away and write your own version of this API! As long as you don't literally take copies of my code and just go write 'foo', 'bar', and 'baz' in a way that semantically does the same thing, then you're safe!
However, suppose 'bar' is a simple in-place sort, and you and I both implement a standard quicksort to do the job.
Sure, you wrote yours totally independently of me, but I could still go to a judge and claim that, no, you copied my version! Your copy of 'bar' is in fact a copyright violation because you just stole my code!
How would you prove otherwise?
So instead, what we do is get a third party. Their job is to read the decompiled versions of 'foo', 'bar', and 'baz', and then write down a totally independent specification that describes how those functions work, but doesn't contain any of the code.
To be extra safe, we even get a lawyer to read the resulting specification and certify that, indeed, no copyrighted code is present.
Then we hand you the specification, and you use that specification to implement 'foo', 'bar', and 'baz'.
Now, the specification might say 'bar is an in-place sorting function', and so you go ahead and implement your own version of quicksort.
But now, when I claim you copied my code, you can go to the judge and say "Au contraire! I worked strictly from this specification, here, that my lawyer has verified contains no copyrighted code. The only person that actually read the code is that guy over there points dramatically to the audience and he didn't write any of the code."
This provides a much much stronger defense that your implementation cannot possibly contain illegally duplicated, copyrighted code, and that any similarities are entirely incidental.
Yes, this whole thing probably seems like a crazy dog and pony show, but when you're a little upstart company like Compaq going up against the behemoth that is IBM, you can be damn sure you're gonna dot all your i's and cross all your t's, because they will send an army of lawyers your way, and it won't be pleasant.
This is just a popular method to increase the likelihood of success in the event of lawsuit. It's not actually standard, required or anything of the sort. It certainly doesn't prevent competitors from suing you anyway and burning your time and money in court. They can also be awarded an injuction that stops you from making money until the courts decide who's right.
Sony vs Connectix is an example of a company that directly reverse engineered firmware and still won in court but still lost in the market due to an injunction.
https://en.wikipedia.org/wiki/Sony_Computer_Entertainment,_I....
> the PlayStation firmware fell under a lowered degree of copyright protection because it contained unprotected parts (functional elements) that could not be examined without copying.
> While Connectix did disassemble and copy the Sony BIOS repeatedly over the course of reverse engineering, the final product of the Virtual Game Station contained no infringing material.
In the BIOS cases, this means that the implémenter can produce the exact same assembly code as the original BIOS and it’s not copyright infringement, but only if the implementer never saw the original code.
Yes.
> but only if the accused did not have access to the original work.
No, but the accused having access to the original work makes it less likely that a trier of fact (jury or judge depending on the kind of trial) will conclude that the creation was independent rather than copying.
I don't know of any actual cases that hinged on this, but lawyers tend to be belt-and-suspenders types. One remnant is the advice from the GNU folks on building clones of Unix tools---you could even be tainted by looking at the Unix source, as long as the clone was architecturally very different. (Which led to a lot of better implementations of Unix tools...)
If the same person reviews the product being cloned, writes a spec, has it reviewed, and does the implementation, the other side can argue that knowledge that was not in the spec was used in the development. They might or might not succeed in this argument, but cautious companies firewall off those who have seen the original product and those who develop the clone, especially if the competing company has a litigious reputation.
Even in asm books they would show flowcharts of the logic first and then try to implement those flowcharts while many modern things assume you have that global idea and jump right into the details.
I never sweat the details until I need to which seems to make me a child of those times, which I am of course.
On the flip side, these old books also discuss far more low level stuff, some to the gates level.
I was impressed with the step-wise manner in which the book proceeds from simple observations of soap bubbles and how they behave in order to develop more complex theories as to why they behave that way and the physics behind it.
Logic can feel so refreshing sometimes.
1) https://www.arvindguptatoys.com/arvindgupta/soap-bubbles-boy...
The film itself cost money. You had only a certain amount of it on-set.
You then had to wait for it to be developed to see how it turned out.
You had to synch audio to it, so you had to really think about what you wanted to say and for how long.
Finally, the distributed product had a standard length for distribution. (The reels sent out.)
I picked up a series of books on vacuum tube circuits from the 1950's created by and for U.S. military personnel and it is probably the easiest and most digestible explanation on how vacuum tube circuits work that I have come across. Wonderfully illustrated with diagrams as well.
EDIT: Here you go (maybe just Vol 1): https://worldradiohistory.com/Archive-Rider/BASIC%20ELECTRON...
Also (all 5 volumes): https://archive.org/details/BasicElectronicsVolumes151955
Here's an old video about celestial navigation:
It feels like we've forgotten how to make videos like these.
These things would not have been filmed at all in the past, or they would have been much shorter. There would be a film crew. Someone doing the microphone setup and telling the subject to enunciate etc.
And us corporates would have watched it in the cafeteria before lunch on the 16mm projector.
The low barrier to entry today is wonderful, but it also means a lot of stuff made is just fluffy and not very necessary at all. I guess it's also easier to skip most of it, so there's that.
I find this to be true and very relevant in modern education. You don't need to gamify mathematics. You need to clearly communicate just how awesome and interesting mathematics itself is.
This is just that generation. This is a good way to describe my grandfather or any of his friends.
What makes it even more amusing is that people on this forum, even if they don't know FORTRAN would find it ludicrously elementary considering it was actually the textbook for an MIT class at the time. (It was basically a programming course for non-CS/EE engineers and the assumption was basically that they had never touched a computer before.)
I subsequently had a quite granular mental map of how higher-level software interacts with the hardware through BIOS and DOS interrupts, and knew a lot more than I generally needed to about where the MBR is stored, disk cluster size, etc.
It was difficult times to learn about computers - no friends used them, there was no internet, so any knowledge acquired was hard-earned, and tended to really stick.
I wore my Vic 20 manual out copying out programs from it while learning not only BASIC but some pretty fundamental stuff about computer architecture. It was so significant to me I had to buy a copy from eBAY decades later to keep on my shelf.
Osborne & Associates went on to sell a portable CP/M machine that was very bulky and heavy but earned the right to be called portable.
[1] https://boingboing.net/2016/02/07/usborne-releases-free-pdfs...
Worse, the only programming books I could get running back then were those "Learn Java in 24 Hours" style books that came with a cd that included a compiler, but no IDE. I remember giving up after a lesson or two because all I had was notepad to type in the code, and the errors thrown were too complex to understand. (Probably a missed semicolon.)
I seem to remember it being there, but it didn't include any of the built in examples or games. Maybe I tried programming things from the books and either couldn't get them to work, or they would run too fast to be playable.
Either way, it wasn't predominately featured and without knowing anyone with programming experience, I didn't have any one to ask.
https://stackoverflow.com/questions/562303/the-definitive-c-...
https://www.amazon.com/Head-First-Design-Patterns-Brain-Frie...
I'd say Knuth is squarely in the latter category
Another way of counting is that the book "The MMIX supplement", which contains the MMIX equivalents of every single page/section/program in Volumes 1–3 that is affected by the details of MIX (the book is not by Knuth but the preface indicates Knuth reviewed it very thoroughly pre-publication), is 224 pages long.
Volumes 1, 2, 3 are together 672 + 784 + 800 = 2256 pages long (and Volume 4A is 912 pages). So roughly, less than 10% of TAOCP deals with either MIX specifically, or assembly language in general.
⸻
1. https://www-cs-faculty.stanford.edu/~knuth/mmix.html
2. For people targeting a platform like JVM or CLR (dot net), understanding how those machines are implemented might also be useful, although I've found my ancient memories of 6502 and 370 assembler are sufficient for having a mental model of how the code works.
>Many readers are no doubt thinking, ``Why does Knuth replace MIX by another machine instead of just sticking to a high-level programming language? Hardly anybody uses assemblers these days.''
>Such people are entitled to their opinions, and they need not bother reading the machine-language parts of my books. But the reasons for machine language that I gave in the preface to Volume 1, written in the early 1960s, remain valid today:
>• One of the principal goals of my books is to show how high-level constructions are actually implemented in machines, not simply to show how they are applied. I explain coroutine linkage, tree structures, random number generation, high-precision arithmetic, radix conversion, packing of data, combinatorial searching, recursion, etc., from the ground up.
>• The programs needed in my books are generally so short that their main points can be grasped easily.
>• People who are more than casually interested in computers should have at least some idea of what the underlying hardware is like. Otherwise the programs they write will be pretty weird.
>• Machine language is necessary in any case, as output of many of the software programs I describe.
>• Expressing basic methods like algorithms for sorting and searching in machine language makes it possible to carry out meaningful studies of the effects of cache and RAM size and other hardware characteristics (memory speed, pipelining, multiple issue, lookaside buffers, the size of cache blocks, etc.) when comparing different schemes.
Also, the underlying hardware he's describing isn't universal. Look at the assembly language for a vector machine (like a Cray). It's nothing like MIX.
These two statements are not contradictory. We've stopped teaching CS using assembly language a long time ago, but we still do teach assembly in OS or computer architecture contexts, which is totally fine and should continue.
And the fact that it's fictional it doesn't mean it won't age. Because he didn't pull the language out of thin air and he wasn't working in a vacuum, he was inspired by the trend at the time of its invention [2]. I bet if it was written today, it would look quite different, as proof of his language update for the recent editions [3]. So, it's anything but timeless, as said by even the author.
> won't age because it never existed to begin with.
Same as Da Vinci's helicopter? I think that one also aged quite poorly.
[1] https://en.wikipedia.org/wiki/MMIX
[2] https://esolangs.org/wiki/MIX_(Knuth)
[3] "In my books The Art of Computer Programming, it replaces MIX, the 1960s-style machine that formerly played such a role"
Although I guess Volumes 1-3 haven't yet had their MIX code replaced with MMIX, so perhaps even the current ones are vintage.
However, I think the raw BIOS functions don't work in extended mode, but I may be wrong.
So even if there are exploits lurking, they wouldn't affect modern OSes which are all at least 32 bit now.
This is specifically to the x86 family, it's possible that some other chip still has original code ready to go even in 32 or 64 bit mode.