Legacy Train Control System Stabilisation (2012) [pdf]
webinfo.uk
webinfo.uk
Metro Trains which is a consortium took over the train network at this point and was contractually obligated by the state to upgrade all core and edge technology systems.
I worked at MTM circa 2016/2017 and all the truly legacy operation control systems had been replaced.
there was a heap of funding allocated for fun things like maintaining OCMS/IT systems and updating customer facing/passenger information systems.
> that emulated PDP-11 system was replaced in 2014-15 by an off-the-shelf “Westrol” system from vendor Invensys.
> Due to the core software limitation (no source code available) we were compelled to integrate some original PDP-11 computer cards into the final product and this resulted in a hybrid PC platform. There are two very distinct hardware technologies in use in each complete system. The PC environment which is the host platform and the DEC environment which is the ‘legacy’ system. The connection is made using a proprietary Unibus adapter that links the Osprey co-processor to the short Unibus. This Unibus which is totally contained inside the special PC housing allows us to retain some key legacy DEC PDP-11 cards that in turn deceives the software ‘to think’ it is still running in a full DEC system in terms of timing. A marriage of ‘old’ and ‘new’ was borne
One of God's own prototypes. A high-powered mutant of some kind never even considered for mass production. Too weird to live, and too rare to die.
But why did they choose such a card in the first place?
A particularly nice feature of the bus is that it was memory-mapped, and everything (memory as well as custom hardware) would send an ACK signal back to the processor when addressed; addressing a non-existent memory location would cause a bus trap.
Edit: The bus is simple enough that creating an arduino, say, interface to drive old DEC cards woould not be hard. I'll keep that thought for the next time someone needs their nuclear power station updating [0].
Meditech, however, uses its own network stack on the OS, replacing all Windows networking components with a Meditech MAGIC driver, and Meditech's own software then controls TCP/IP inside its emulation that runs on top of 2000 in usermode. It's amazing how well put together the stack is that this could be (fairly easily) migrated up to a Server 2012-era box* and still work the exact same way.
*(that I know of - I assume they have newer support by now)
https://www.berkshirehathaway.com/
I think neither business needs to convince anybody of anything. If you need what Strobe is selling and can afford it, they probably made the sale.
Ok the web page looks old but maybe not that old
Here's the 'Chairman's Letters To The Shareholders (Historical Compilation)' from the 1997 version of their site - https://web.archive.org/web/19970530212045/http://www.berksh... , with yearly publication.
In terms of expressiveness, digital logic can't "do" anything that software can not (including drive real electrical / analog IOs provided the software and the digital logic have interfaces for). Whether you implement the pdp11 in software or logic, either way you still have to create it to take into account all the timing details of the original hardware.
It's just some operations can be done much faster in hardware, but I doubt any part of the pdp11 can't easily be emulated in software on modern CPUs with a lot of timing margin to achieve real-time cycle accuracy. I could be wrong because pdp11 I don't know much about, but for an early simple RISC pipeline of similar performance it is certainly possible.
I assume this is just a solution that works and is supported and they've been using for a long time. It would
In simulation it's sufficient to compute accurate cycle information, yet slower than real time, or with jitter.
In contrast to that "cycle accurate modelling" the goal here was to replace a pdp11 in operation.
It seems like a super fast system would be easy to do this with, but reality is different.
Sure, my point with anecdotes is that the hardware system can be modeled in software to cycle accuracy, and it can be emulated with software. Therefore it can be emulated with cycle accuracy (provided you have the CPU power to meet timing).
> and if your emulator does a garbage collect, or some other thing, you'll miss it.
Well if you do anything that blows your timing budget then you miss your deadline by definition on any real time system. Obviously you can use real time garbage collection (or not use it) and control other latencies.
I wasn't suggesting you could just write some code and make a real time cycle accurate emulator without paying attention to real time deadlines.
EDIT: From the FPGA vendor:
> Double+ speed Osprey Co-Processor. Occupies one ISA/EISA PC slot. Writable control store, FPGA implementation of PDP-11 architecture with 4 MBytes of tightly coupled, zero wait-state memory. Includes FPJ11® compatible hardware floating point. Performance equivalent twice the DCJ-11® and 100%; of FPJ-11. On-card x86 microprocessor rapidly processes "virtual" I/O instructions. This card is used for Osprey/PC applications where all applicable I/O devices are emulated on the PC by Osprey software. Includes standard environment support for MS-DOS® V6.2.2. Windows/NT® environment support available.
So, highly unlikely that such a system using standard x86 PCs with all their SMBIOS traps and other garbage is able to emulate IO could possibly drive external IOs with better precision and accuracy than a proper real time SOC.
Cycle accuracy may be a requirement in the processor such that legacy software is guaranteed to run the same, but they almost certainly are not doing cycle accurate IOs if they're going across an ISA bus to a DOS or NT PC to be emulated there.
We had a punched-tape and microfiche of the xxdp diagnostic test suite to pass: http://bitsavers.org/pdf/dec/pdp11/microfiche/ftp.j-hoppe.de...
Several nuances that we replicated were probably never used by the production software. For example the order of execution of a floating point instruction trap followed by a stack pointer trap due to the pipelining of the fpu instructions, even though we weren’t executing anything out-of-order. I also produced the correct partial result for the division algorithm in the event of overflow condition.
Amazing. I've never been anywhere that lost their source code. Would love to know how it happens.
1. Contractor relationship gone sour, turned out they had withheld source code from us. This lead to a dispute regarding the code ownership. Meanwhile we were running the corresponding binaries in the wild. We ended up reverse engineering the missing parts.
2. Source code is centrally managed on a server, hard disk failure, restore not working or backups wholly missing, don't remember. This was the pre-git era, probably Visual SourceSafe, which was entirely centralized.
For me personally, let's just say I feasted my eyes and learned a lot. But I wager that there is still a lot of companies vulnerable in very similar ways today.
There were version control software around in the 70’s but I do not think that the adoption rate was very high. I think it’s reasonable to assume that they didn’t use it in this case.
> The Programmer's Workbench has proven to be very popular with both management and programmers, and is now used by almost all software projects at the author's installation.
The followup paper https://dl.acm.org/doi/10.1145/800283.811111 at https://dl.acm.org/doi/pdf/10.1145/800283.811111 comments:
> SCCS is used extensively within Bell Laboratories. For example, at one installation more than 3 million lines of source code are controlled with sccs. These 3 million lines represent 5000 files and the work of approximately 500 programmers. Additionally, sccs is used at many installations outside the Bell System.
So, SCCS runs under Unix. The software package we are talking about here is Ericsson JZA715, written mostly in Pascal with a bit of PDP-11 assembler, running under RSX-11M, it was first put into operation in Oslo in 1979. I don't know what development platform Ericsson was using, but I doubt it was Unix, seems far more likely to have been RSX-11. What version control systems would Ericsson have been using on RSX-11 in the 1970s? I doubt it was SCCS – could it have been something else? Anything? RSX-11 has a versioned file system, could they have just used that? Or did people build more advanced version control systems layered on top of that base?
The original version was for OS/MVT and implemented in SNOBOL4 using the SPITBOL compiler. https://apps.dtic.mil/sti/pdfs/ADA084326.pdf says that SNOBOL4 was available for the RSX-11D by 1979.
I have no clue what Ericcson was using. My inference from the literature is that apparently when software developers have access to a version control system, they used it, even back in the 1970s.
FWIW, I regard versioned file systems as a form of version control.
I did come across https://books.google.com/books?id=Gf4YAQAAIAAJ&pg=PA89&dq=%2... from 1978 ("Software Designers Workbench", Paul A. Scheffer, Martin Marietta Aerospace
> One of the more innovative of recent software engineering activities is the concept of a Programmer's Workbench (PWB)
(PWB is the term for the product which included SCCS.)
> A simple generalization of the PWB concept results in the idea of a total software engineering facility, i.e., a Designer's Workbench (DWV) ...
> Using a PDP 11/70 architecture base, and system software which includes RSX-11M and the UNIX operating systems, standard PWB tools, and INGRES data base capabilities ...
Since it looks like UNIX supported the RSX-11M in 1978, why couldn't Ericsson be using Unix on an RSX-11M at the same time in Norway?
When I worked on RSX-11M, and later VMS, we did not rely on SCCS because it had no added value. We relied on the file versioning provided by the file system and rarely kept more than one old version around because disk space was always at a premium. Changes to production code required filling out a paper form and passing it to the test department, who would submit the appropriate batch job to rebuild everything, then copy your changed files into the production directories after the tests pass. It was just like modern CI systems, except for the fact that it was mostly manual and a build-and-test cycle was at least 2 days.
It's no wonder everything is emulated/old/never-replaced.
Edit: Apparently it did eventually get replaced https://news.ycombinator.com/item?id=29158255
Edit: Also worth pointing out that formal verification was likely not a component of building the original software, and it likely has bugs it's kept since inception. Updating the API to a language more conducive to formal verification might be an option that improves safety over the legacy system
Edit: Apparently they did replace it all eventually
...which was then presumably lost/reformatted some time after said dev left (quit? laid off? fired?)
Unable to build said source code, a subsequent developer added a python script to postprocess the output of said binary for their own needs. A hack, but it worked.
Then I got my hands on it. Well, I needed data that was discarded by the binary, so hacking on the python script was a non-starter. Instead, I reverse engineered what the binary was doing by stubbing out enough of the source code to get it building - then compared hex dumps of the outputs and tweaked said source code until my modified source code gave the same output as the checked-in binaries. Then I got rid of the python script by moving the logic into the binary's source code, where it really belonged. Then I made my own changes.
Didn't bother setting up CI though. Was porting a project in a heavily compressed (<1 month!) timeframe, and nobody else needed to touch said binary, and the codebase wasn't going to be used for any future projects (it was pretty terrible, and much time had been spent creating a much better codebase for future projects.) I at least made sure I checked everything in, though!
I have heard this many times, especially for the "single use executable for a specific use case". Many years later, it had made its way into production code and is on its 10th iteration.
Does Ericsson still have the source code for a PDP-11 software package they sold in the 1970s? Maybe they do, maybe they don't.
Did they ever even ask Ericsson for the source? Possibly, they may have thought that modifying PDP-11 software was too hard–people with skills in doing it are hard to find these days, it is unlikely to be written in C, probably either assembly or some obscure language few have heard of–so even if Ericsson still had the source and was happy to hand it over (free of charge or for a reasonable fee), they might not have thought it worth their while to acquire it, and hence never have asked Ericsson for it.
EDIT: Actually, turns out JZA715 is mostly written in Pascal (so not that obscure a language after all) with some modules in PDP-11 assembly – see page 18 of http://srsv.org.au/wp-content/uploads/2017/02/S-38-1-Jan.pdf – the software ran under RSX-11M
I'd be interested in learning how the signalling tech has evolved over the years. There has always been a push by PTV (the government body responsible for public transport in Victoria) to establish a "turn up and go" timetable where trains depart often enough that you shouldn't need to plan ahead. One of the restrictions of the old signalling system was a fairly large minimum distance between trains that prevented frequent service. This caused some congestion, especially in the hub and spoke model of the Melbourne network where many lines converge to share tracks as they approach the CBD.
Yeah, is this factoid - central Melbourne train signals being controlled by an emulated PDP-11 - still true in 2021? I don’t know. Definitely this is known to have been true some years ago but can anyone confirm if it remains true today? The construction of the Melbourne Metro requires new signals, which I assume would be controlled by a brand new control system not the old PDP-11-based system. But if you are going to introduce a new system for one line, wouldn’t it make sense to try to migrate the existing lines to it as well?
> Train Describer replaced with TCMS in 2014-2016.
Replacing train management systems like this with new ones seems to be all but impossible, my brother in law worked on one in the UK that was designed, built and implemented and the controllers basically said "no thanks". It only got implemented when they rebuilt the UI and control panels to look like the old one.
Oh. So that's the real problem. They lost the source.
"Not fun". Kinda glad we didn't get the contract.
I assume this will be similar to the European Train Control System Level 3 [2] (which does not exist in a cross-vendor spec yet afaik) - it's being built by Rail Systems Alliance, a consortium made up of CPB Contractors, Alstom and Metro Trains Melbourne.
[1]: https://bigbuild.vic.gov.au/projects/metro-tunnel/about/tech...
[2]: https://en.wikipedia.org/wiki/European_Train_Control_System
This video was great thanks!
My naive assumption was that as trains operate in a relatively controlled environment a headway much smaller than 120s would've been easy to achieve. Cars regularly drive sub one second from each other! Is it just about minimizing the chance of catastrophe or is there more to it?
Emergency braking distance of a TGV from 300 kmph (186 mph) is something like 3,500 m [2], carrying ~450 people.
Moving block signalling systems take into account the current positions and velocities of both trains, and generate a maximum speed and braking curve. See [3] for the details, it's pretty fascinating systems design!
[1]: https://en.wikipedia.org/wiki/List_of_countries_by_traffic-r...
[2]: https://railroads.dot.gov/sites/fra.dot.gov/files/fra_net/16...
[3]: https://en.wikipedia.org/wiki/European_Train_Control_System
Trains are held to higher safety standards, and so even modern train control system are usually based on absolute braking distance (and even when you want to space trains closer together regardless, as soon you need to throw a set of points between trains at a junction you're back to absolute braking distances, because as long as they aren't set safely towards one direction or the other, a set of points is effectively a stationary obstacle), i.e. there's always enough free space in front of the train to safely brake to a stop.
This limits how closely trains can follow each other, and on urban railways it's normally the combination of dwell times (doors opening, passengers alighting and boarding, doors closing) and platform reoccupation times (wheels starting turning on the first train leaving the station to wheels stopping turning on the next train arriving at the station) that determine your minimum headway. (And for actually usable – as opposed to merely theoretical – capacity you also need to add at least a small amount of extra margin to compensate for possible dwell time variations and other day-to-day occurrences)
The Jubilee line is 36km long, and each train is 125m [1] so that means there is on average just over a kilometer between each train at 30 trains per hour. That might sound like a lot, but stopping distances of trains are much greater than cars. First of all the trains weigh around 200 tons plus another 70 tons for passengers, and metal on metal doesn't have as much traction as car tires. Even if you could decelerate as fast as a car, everyone standing up is going to end up in a pretty bad shape after that so you need to do it slower. According to someone on StackOverflow passenger trains are designed to decelerate at 1.2m/s2 [2], which at 100km/h gives a stopping distance of 350m.
[1] https://en.wikipedia.org/wiki/London_Underground_1996_Stock
[2] https://www.quora.com/How-long-does-it-take-for-a-train-to-c...
I don't believe that is true. The Melbourne city loop was using the Ericsson JZA 715 system. According to page 9 of [0], Melbourne's implementation of JZA 715 went live in 1982, but Oslo's JZA 715 implementation went live in 1979. Furthermore, JZA 715 was an evolution of earlier Ericsson computer-based rail traffic control systems, going back to the JZA 410 which went live in Stockholm in 1971 and Copenhagen in 1972. So, if these Ericsson systems were in any sense a "world-first" (don't know enough about the history of this technology to say), it looks like the "world-first" happened in Scandinavia, and Melbourne may have simply been "Australia-first" (or "Southern Hemisphere-first" or maybe even "not-Scandinavia-first")
[0] https://www.jonroma.net/media/signaling/articles/ericsson/L....
> Valuing Navy personnel highly, the Navy elected to field test with one contractor and nine Marines.
You won't be able to use them in a system at the current time though.
> You won't be able to use them in a system at the current time though.
What do you mean? Why?
I did not realise just how much of the network is Dark -- for instance, everything beyond _Burnley_ (just 3 stations away from Flinders St) is Dark!