How the Singapore Circle Line rogue train was caught with data
blog.data.gov.sg
blog.data.gov.sg
One thing that I think would make the collection even better is spoiler free titles to maximize the suspense.
For example, the screen saver story lost much excitement when I knew the cause from the title.
Edit: Or, you know, hardware failures that nobody expects anymore.
This story is a beautiful example of such a story. Great read, great bug, great analysis!
We seem to condition ourselves to quickly forget the excitement because even the most epic bughunt will tend make you lonely if accidentally used for conversation in polite society.
Years ago I worked on a stock control system for a popular clothes retailer. Where previously store managers had to spend hours on the phone ringing the depots to see what was in stock and reordering popular lines, now they could fill in orders an they would be emailed to the depots in machine readable format, with the order balanced across multiple depots depending on availability of stock.
After the system had been in operation for a year or so, managers started to complain that some orders weren’t turning up. After a bit of digging it only seemed to affect one particular depot, but I couldn’t initially see a common factor between the orders. It had nothing to do with the size of the mail, or the originating store. I could tell from the logs the emails were being formatted and supposedly sent to the depot, yet they never translated into orders. The system would silently fail after we’d handed the email to the depot with no obvious reason why.
So I wrote a quick script to collect all the orders for a day from all the stores and compile a list of all product lines that were present in the orders that went missing, but not present in the orders that succeeded, just to see if there was a pattern. This turned up a few dozen items and after a quick visual scan of the product descriptions one stood out slightly, so I phoned the depot on a hunch.
The depot had an anti-spam filter on incoming emails that silently dropped messages with pornographic words in the body or any attachments. The retailer had just received a new seasonal product line, which was selling well, and the product description was “bondage tights”.
After getting the depot to whitelist emails from our system, orders for this product line started working.
PayPal refused to admit to us that this was the cause and insisted they couldn't shed any further light on the issue despite the fact we could trigger it 100% of the time.
In the last few weeks it's started randomly choking on the word "Pharaoh", God knows why, and God knows why it only does so about half the time, but I'm not going to bother asking PayPal about it.
(last seen in Hungary, likely being involved with the prime minister; there was some scandal about it)
But yeah, it sucks that their system can't learn to give you an exception after it has been reviewed once or twice.
It wouldn't boot if you tightened the screws on the case.
Solution: Don't tighten the screws on the case.
Solution: Move the HDD such that tightening the screws doesn't bend the SATA cable out of making good contact!
The worse version is broken in-board signal paths.
Both ultimately result from mechanical distortion.
In one of my subroutines I changed the value of a passed in variable. In my code there was a point where I called that subroutine with the value '0' for the variable.
The Fortran compiler was pass-by-reference, so the compiler accommodated me by creating a reference to the constant 0 and passing it to my code. In my subroutine I changed the value of that passed in 'variable'. Since the compiler was using that "constant" elsewhere any conditional that looked for 0 would fail after that point since the value of 0 had changed.
My professor and I only found it after we printed out the machine code generated by the compiler.
How does that get through testing? "Passing a number to a subroutine can result in the value of the number being changed"?
A subroutine receives values from the caller through
the parameter list; likewise, the subroutine returns
values by modifying the variables in the parameter
list. Subroutines are invoked in a FORTRAN program
as a statement, not an expression. The syntax used
is,
CALL name(actualParamList)
See you modify the parameter to send a return value. Really ancient stuff. And if you're working on hardware that is challenged in not having a "load immediate" instruction for floating point, well you get "optimizations" where the compiler creates what is essentially a pre-initialized variable and uses the address of that when a constant is used. int increment(int x) {
x = x + 1;
return x;
}
int y = increment(2);
Would blow up your program's concept of 2, and I can't imagine that being missed before sending it out to users. But I've never used Fortran, so I may be seeing false equivalences to structures and practices in modern languages.void foo(Integer i) {
// Putz with reflection, change value stored in i to something else
}
main() {
foo(150); // Nothing special, you just created an autoboxed integer and then alter its value
foo(4); // Java pulls you an Integer from the autoboxing cache, and you do indeed accidentally alter the value of 4.
}
I suspect that someone flipped the switch away from "More Magic" :)
[1] https://twitter.com/stansted_exp/status/527011030728441856
For reference, 1kV arcs across about 1cm of air, so 25kV needs a full foot of space in air (obviously less with a dielectric, but you get the orders of magnitude).
Source: I design trains for a living, though on the mechanical side
Arguably, based on the preliminary explanation, it seems like PV46 might be over-represented (it had 4 interruptions) -- after all, it never travels in the opposite direction to itself.
> ... faulty train signalling hardware on PV46 was emitting erroneous signals in addition to the ones it was supposed to emit. These erroneous signals occasionally prevented trains travelling in the vicinity of PV46, including in the opposite direction, from properly communicating with the trackside signalling system. This loss of communications led to the activation of the trains’ emergency brakes.
[1]: https://www.lta.gov.sg/apps/news/page.aspx?c=2&id=89defe38-4...
While the press release says these were "in addition to the ones it was supposed to emit", it's not hard to imagine that a similar error could happen where the wrong signals were sent in-lieu-of the intended signals, and thus leave trains rolling that should have detected an emergency-stop situation.
Not all problems can be troubleshot and there are some things that remain mysteries forever.
From the bottom of the linked press release, which summarizes the overall investigation:
"In particular, I thank the engineers and data scientists from DSTA and GovTech respectively, without whom we would not have been able to theorise the possibility of a faulty train, and identify PV46."
A "steely eyed missile man" moment. https://en.wikipedia.org/wiki/John_Aaron
It's truly wonderful to see how they went about identifying the problem. This is fantastic work and the team deserves major kudos for nailing this.
EDIT: I am corrected by "Rifu". Activity data of all traincars was not available in the first instance, so a broad association analysis would probably have yielded nothing. The type of analysis the authors use seems warranted. Although I'd still recommend doing broad association analysis on any dataset, before moving into complicated visualization techniques.
It was done in a single day, OK, but the train usage data is not "extra data" it is something that is "ordinary data" that you have typically for two to four weeks in advance (as a planning) and daily or at the most weekly afterwards, let alone - in November - the actual usage data for August-October.
You have to manage a time table, the personnel and the maintenance of both the rails (and station and power lines, etc. i.e. the "static" parts of the infrastructure) and of the "moving" parts (the trains).
In order to do so you have "allocation" tables, usually made in advance two or four weeks at least, stating "who" (personnel) and "what" (train) are "where".
These "programs" are monitored and changes to it (think of people not showing up at work, a train having a malfunction, etc.) modified to allow for these impredictable events.
So at the end of the day (each day) you had a printed piece of paper with modifications - if any - scribbled on it.
Nowadays the same thing is done on computers, possibly in a much more detailed way, but the basic data "which train was in service on which line at what time" has been available on paper since day one (or two) of any railway in the world.
While I can (barely) understand how these basic data for the 5th of November was not available on the 5th night, there is no real reason why data since August and until - say - the 4th of november was not available.
As said, the data analysis carried on the data provided by SMRT has been carried on in a clever way, as finding the culprit from that set of data would have not been at all easy through other simpler analysis methods, but somehow SMRT (and LTA) provided the "wrong" set of data.
Also remember that until they begun the analysis it wasn't clear that a single train was the cause. I imagine they were investigating trackside issues as well as issues with those trains that stopped, rather than an independent train that wasn't experiencing problems.
TLDR; It's not a case of looking at pre-defined timetables.
1. http://www.rioranchomathcamp.com/Topology/SubwayNamedMobius....
OT: I am a huge fan of Python, so whenever I see Python helps to solve a thing I always say myself "yay Python!"
http://www.straitstimes.com/singapore/transport/mystery-of-c...
For the next few days after, they shut down the telco signals in the train tunnels to determine if it caused the signal interference.
Probably the handoff to data.sg was probably done because it doesn't necessarily make sense for them to maintain a separate data science division within the company.
They are much, much better than that. They hire the best grads and pay lots, but expect a lot too:
One thing that stands out in Singapore is the quality of its civil service. Unlike the egalitarian Western public sector, Singapore follows an elitist model, paying those at the top $2m a year or more. It spots talented youngsters early, lures them with scholarships and keeps investing in them. People who don't make the grade are pushed out quickly.
Back in Canada, on the other hand...
[1] https://en.wikipedia.org/wiki/Senior_Wrangler_(University_of...
I wonder why the rail managers allowed the defects to inconvenience passengers for so long instead of rethinking their maintenance procedures.
Not obvious.
I just yesterday attended a presentation by a ex NSA guy (who's of course pitching his company Adatos, so grain of salt) who claims you could feed the raw data in and the pattern will be found.
As a side note, I find alarming the fact that the blog of the Data Science Division of a government (any gov) is hosted on a third party provider.
These guys look like they know what they're doing though, so I guess it was a weighted decision.
EDIT: Great article and good problem solving skills
I propose:
TL:DR - They had a strange hard to identify bug, but used a limited data set and interesting techniques to quickly find the esoteric cause of the problem.
so to add in your point, to his TL;DR I would put in "the intial dataset included only the trains that had suffered the fault, but as the fault was caused by a functioning train, a more comprehensive dataset was necessary to find the problem; had it been provided initially, less detective work might have been necessary"
They shut off the mobile networks for a while to try and isolate the issue. It's interesting to read what was going on, though.