What Happens When You Mix Java with a 1960 IBM Mainframe
thenewstack.io
thenewstack.io
My current struggle is with one line of code:
(defun getenode (l) (cadr l))
That ought to be simple enough. But it's being applied not to a list, but a "hunk". A "hunk" is an obsolete MacLISP concept.[1]. It's a block of memory which has N contiguous LISP cells, each with two pointers. This is the memory object underlying structures and arrays in MacLISP. Macros were used to create the illusion of structure data objects, with hunks underneath. However, you could still access a "hunk" with car, cdr, cxr, etc.I'm converting this to Common LISP, which has real structures, but not hunks. That, with some new macro support, works for the regular structure operations. So far, so good.
But which element of the structure does (cadr l), which usually means the same thing as "(car (cdr l))", access? (cadr (list 0 1 2 4)) returns 1, so you'd think it would be field 1 of the structure. But no. It's more complicated and depends on how hunks are laid out in memory.
The Franz LISP manual from 1983 [2] says "Although hunks are not list cells, you can still access the first two hunk elements with cdr and car and you can access any hunk element with cxr†." At footnote "†", "In a hunk, the function cdr references the first element and car the second." This is backwards from the way lists behave.
A blog posting from 2008 about MacLISP says "A Maclisp hunk was a structure like a cons cell that could hold an arbitrary number of pointers, up to total of 512. Each of these slots in a hunk was referred to as a numbered cxr, with a numbering scheme that went like this: ( cxr-1 cxr-2 cxr-3 ... cxr-n cxr-0 ). No matter how many slots were in the hunk, car was equivalent to (cxr 1 hunk) and cdr was equivalent to (cxr 0 hunk)." Note that element 0 is at the end, which is even stranger. The documentation is silent about what "cadr" would do. Does it get element 2, or get element 0 and then apply "car" to it?
The original code [3] contains no relevant comments. I'm trying to figure out from the context what the original author, Greg Nelson, had in mind. He died in 2015.[4]
[1] http://www.mschaef.com/blog/tech/lisp/car-cdr.html [2] http://www.softwarepreservation.org/projects/LISP/franz/Fran... [3] https://github.com/John-Nagle/pasv/blob/master/src/CPC4/z.li... [4] https://en.wikipedia.org/wiki/Greg_Nelson_(computer_scientis...
Based on †, it sounds like (cadr (hunk (hunk 1 2) 3)) should return 2.
Is the old LISP code available online somewhere? I'm curious to see it.
I put all the code on Github. The oldest version of each file is exactly what ran in 1986.
The code is delicate. It's a theorem prover, and there's much manipulation of complex data structures, with little explanation of what's going on. The overall theory is documented; this is the original Oppen-Nelson simplifier and there are published papers. But the code has few comments.
(getenode znode)
i.e. getenode is only called on znodes. The definition of makeznode is: (defun makeznode (node)
(prog (l znode)
(setq l (list (list '(1 . 1) node)))
(xzfield node (list l node nil))
(setq znode (tellz l node))
(or (null znode) (eq znode node) (break makeznode))
(return l)))
So it seems like getenode is called on a regular list whose structure looks like (list (list '(1 . 1) node))
Assuming "node" is an enode, the way to access it would be (cadar znode), not (cdar znode). Try changing the definition of getenode to (defun getenode (l) (cadar l))
and see if it runs. (or (setq znode (tellz l node)) (return t))
(or (eq node (getenode znode)) (zmerge node (getenode znode)))))
and (tellz) has taken the path which ends (return (baserowz* i)). "baserow*" is an array which contains links to "node" items, not list cells. What "z.lisp" and "ze.lisp" are doing, by the way, is solving systems of linear inequalities by using linear programming on a sparse matrix.Also, see
(defun isznode (x) (and (hunkp x) (= (hunksize x) 8))) ; original
(defun isznode (x) (equal (type-of x) 'node)) ; CL version
which indicate that znodes are hunks/structures, not lists. This is inconsistent, yet somehow it used to work.(If you want to talk privately about this, I'm at "nagle@animats.com". Too much detail for HN.)
(unless (setq znode (tellz l node))
(return t))You should probably also test some code in Maclisp.
The order of display of hunk slots is historical in nature. For better or worse, the elements of a hunk display in order except that the 0th element is last, not first. e.g., for a hunk of a length n+1, (cxr1 . cxr2 . ... . cxrn . cxr0 .)
It could still make sense to have the layout in memory being sequential, like (cxr-0 == CDR, cxr-1 == CAR, ...others ...).
Note also that CAR extracts the leftmost element of a hunk, just as it addresses the leftmost element of a cons. Similarly, CDR extracts the rightmost element of hunks and conses.
It seems more logical that CADR is just the combination of CAR with CDR. It don't think the designers would try to transpose the fact that it means "second", with proper lists, for hunks. It just seems unlikely, but I have no proof.
Also:
(Note that the operation CAR is undefined on hunk-1's, but CDR is not.) This means that if you want to make a plist for a hunk of your own, you can use its cdr as a hunk; it does not mean that you can blindly assume that any hunk wants its CDR treated that way. The exact use of the slots of a hunk is up to the creator; it's a good idea to mark your hunks (e.g., by placing a distinctive object in their cxr-1 slot) so that you can tell them from hunks created by other programs.
My guess is that there is some metadata associated with a hunk, stored in CRX-0, a.k.a. CDR.
Actually, this blog post makes the story clearer:
http://nikhilism.com/post/2016/systems-we-love/
It isn't a physical IBM 7074.
When it came time to migrate from 7074 to S/360, rather than rewriting their 7074 software, they just wrote a 7074 emulator for S/360. And, it sounds like, they are still running their 7074 software, under their 7074 emulator, most likely on a recent z/Architecture mainframe.
The article makes it sound like people still use "1960s mainframes" when I very much doubt anyone is still running 1960s hardware in production. People use modern machines–modern IBM mainframes, which are multicore 64-bit processors–or other mainframe vendors such as Unisys or Fujitsu use mainly I believe x86-64 running Linux running a software emulator for the old mainframe CPU.
A lot of legacy, sure, but I think this article makes it sound even more legacy than it really is.
Indeed. The point of the talk was that 1) legacy is often assumed to be bad not for any real technical reasons but just because it is legacy and 2) a lot of what was being presented as legacy wasn't even legacy. Their OS 2200 version was actually newer than the Oracle DB they were using on the "modern" side of the stack.
http://www.pcworld.com/article/249951/computers/if-it-aint-b...
You've probably seen plenty of crazy stuff in legacy systems but I'm hoping at least one surprises you. Maybe the first one. :)
Still, even the blog post you link to doesn't make it absolutely clear that java was running on an emulated s/370. It says that the decision was made to emulate the older architecutre rather than rewrite the old programs, but then it goes on to say "These are still operational". Does it mean the old programs? Or the old machines? It's hard to say.
As to how unlikely it is to see a very old machine still in use, instead of one made in more recent times, last year I talked to an engineer who claimed he had seen a PDP still in operation in some transport company if memory serves.
IBM offered 7074 emulation as a standard IBM System/360 product.[1] On an S/360, it required some special hardware support. In 1972, IBM gave users a free IBM 7074 emulator, software only, for System/370 machines.[2] They may still be running that program on a Z-series mainframe.
[1] http://bitsavers.trailing-edge.com/pdf/ibm/370/compatibility...
At my last job as a consultant someone made https://github.com/manheim/antimony for testing IBM TN5250 mainframe screens in Ruby - it's like Selenium but even more brittle. :>
And the one they actually accessed and received query results in 6 milliseconds was introduced in 2008, not so old and slow:
https://www.app5.unisys.com/offerings/ClearPathConnection/st...
"Single image performance range of 300 MIPS at the entry level and maximum single image performance of approximately 5,700 MIPS (32 processor system)."
"An expanded memory subsystem supports larger memory capabilities and offers memory configurations that include the ability to expand up to 4GW per cell and up to 32GW for a maximum eight-cell system."
Everything old is new again. I love our profession.
"How are you going to get hacked if I talk about your mainframe? It's not connected to the public internet, is it?"
"No. Well... we don't know... but ... hackers! Hackers are really smart Marianne."
Part of the compromise was that I promised I would only use information that was already available publicly through government reports and news articles. I went back through my talk and documented where each fact was already published somewhere else until they were comfortable with it. So the ambiguity on whether the 7074 was the actual machine or an emulator was deliberate... there were certain things I could not find a public comment on and therefore agreed to avoid making direct statements about.
This all seems super annoying, but it makes sense when you realize how heavily scrutinized public servants are. In the end they are only trying to protect me, my organization and Obama's legacy. Three things that are really important to me. So I can't exactly blame them for it. I was happy to be able to find a middle ground where they felt comfortable, the organizers weren't too badly inconvenienced and I got to give the talk I wanted to.
A.k.a. "Them as not knows their history are doomed to repeat it". But I suppose it's more the case of business imperatives than real ignorance that drive this mad race to make new stuff that works worse than the old stuff only so we can then go back to the old stuff with a different name.
Btw, that lady is my new tech hero:
“The systems that I love are really the systems that other engineers hate,” Bellotti told the audience — “the messy, archaic, half chewing gum and duct tape systems that are sort of patched together.
<3 <3 <3
Thick clients were being pushed in mid 90s. But then Clipper Chip plans went south ...
I have friends who have been looking after legacy applications for an airline running on Unisys. The core apps for reservation, Cargo booking and weight/balance were written in FORTRAN. In recent times, the front end was written in Java to give web access. They tried to rewrite the core apps but it was impossible to do so and get the performance.
That's a common sentiment. I wish I could find the quote by someone who made the transition; it was about how happy they were to be able to compile so much quicker, and how getting immediate feedback made them so much more productive.
I compile/assemble COBOL and IBM's assembly language on a z13 daily and it's pretty much instantaneous.
Oh, I was talking more about older ways to organize the data centre: batch vs timeshare processing.
Well, Cobol is a bit like the C of mainframes - you can manipulate memory directly and so on. You can't really do that sort of thing with Java.
b) if the whole thing was indeed running in an emulator, the emulation overhead would have negated all direct memory access advantages
Some (or all?) of the latest Unisys mainframes run on Intel Xeons, but with custom chips for translating the machine code of their old architectures.
I don't work in this area. I just like reading about it. Though, unfortunately, it's difficult to find clear specifications and descriptions on how these architectures work.
For example, Unisys' 2200 ClearPath architecture is one of the (if not _the_) last architectures still sold that uses signed-magnitude representation, as well as having an odd-sized integer width of 39 bits. (INT_MAX is 549755813887 and INT_MIN is -549755813887, and the compiler has to "emulate" the modulo arithmetic semantics required of unsigned types in C. ClearPath MCP is also a POSIX system, and has to emulate an 8-bit char type.) I discovered that you could download the specification for their C compiler for free online, which was useful when discussing the relevancy of undefined behavior in C. But AFAIU (and this is where finding concrete details is more difficult) the latest models of the ClearPath line use Xeons with custom chips bolted on to help run the machine code of the older architecture. In any event, the point is that while the old architecture is arguably emulated, it's not a pure software emulation that you might assume, and the resulting performance is better than the previous models of those mainframes, which were still being built at least until a few years ago. In other words, direct memory access isn't ruled out because the I/O systems may have been intentionally designed to work efficiently in a backwards compatible manner.
There's a reason why people often resort to emulators.
There's this whole thing about how they are getting data from mainframes where "the data was being returned in between one and six milliseconds".
But then: "harvest that data from the magnetic tape and load them up into more traditional databases. That Java application was extracting the data from the databases"
But then: "But the data from the mainframes was actually arriving (from its new home in the database) in less than six milliseconds. The bottleneck was — of course — the Java application."
So of course it is entirely possibly to write slow Java applications. But then the story seems to end! So what happened? Did they fix the application?
I think, though without any proof, that overreliance of 100s of mixed quality libraries, combined with 'best practices' of enterprise development and heavy application of design patterns creates a very large surface area for change. This makes reasonable translation of functionality to Java almost impossible.
I do think Java the language has a long way to go, but it is catching up a bit. I see it as the least common denominator language for most companies. Large open source java projects have been successful (everything on hadoop, hbase, cassandra etc) and a lot less enterprisey.
> but everything lives in my butt (err, butt).
I had to open the comment in porn (er, incognito) mode in order to parse it :D.
I didn't think magnetic-core memory and CRTs were interchangeable...
While magnetic core memory was dominant from 1955 until around 1976, when silicon went mainstream. Machines with MCM were in general use for a long time after that too.
It's possible the interviewer confused the widespread use of CRT in pre-flat screen monitors. (Shrug, maybe CRT was very briefly mentioned? who knows?!)
Williams Tube memory is pretty interesting BTW. Worth taking a look at Selectron memory too. You know, for fun.
Interesting problems not terribly well described in the post, would love to read more.