An IT migration corrupted 1.3B customer records
increment.com
increment.com
https://www.tsb.co.uk/news-releases/slaughter-and-may/slaugh... - the report itself
Interesting notes from the summary:
* Functional testing took 17 months, and began very late due to project delays.
* The board was told that there were only 800 bugs, and yet the independent review found that there were more like 2000
* Non-functional testing seems to be rushed
* the Sabadell COO reported SABIS, one of its child companies and major IT contractor in the migration doesn't have the capacity to respond and solve incidents after go-live
And the TSB board decided to Go Live anyway.
Looks like a classic IT failure story - project is late, save costs on operations, get tired of testing and fixing and decide that it is good enough to be released. But also some very weird incentives - SABIS was apparently not the best vendor for the job, but they are a part of Sabadell, the bank that bough TSB.
Looking at the "Integrated Master Plan" Gantt chart of the project on page 44 is Waterfall as hell and they still failed to plan for Non-functional testing.
These some first impressions, it looks like a interesting read.
For startups, the myth is that success comes from a "move fast and break things" culture, so they wear IT disasters with a badge of honour. In finance, it's the opposite - with enough planning and governance, huge budgets, and long timelines, they believe you can guarantee success on any project.
Modern technology companies make tons of changes all of the time. As a result they become good at them.
Looking at the reports they learned EXACTLY the wrong lessons. It's not to make fewer changes, it's to become good at changes.
Really? That hasn't been my experience. Nearly every time some moderately large update to any modern software goes out it seems to break a bunch things. What it doesn't break it usually just ditches entirely.
Move a bit at time.
I think modern developers just make a whole lot of bullshit excuses for why they are crap at their job.
My experience is those is that those lessons (e.g. the need for extensive automatic testing) are learned the hard way.
Yup, those updates that change the operating system that is run on billion of devices in extremely varying hardware and configurations. Very simple.
"Developers can not touch the database"
Umm... someone needs to, and we're a small team.
"Developers can not touch the database"
"Why not?"
"They can exfiltrate data".
"So could someone else who has access to the database"
"Developers can not touch the database"
"If we give a non-developer access to the database, by definition, they have access, and therefore the ability to exfiltrate data"
"Developers can not touch the database"
"Can you provide some examples of how other companies who've passed your inspection are making any sort of database changes/upgrades without someone having access to the database?"
"Developers can not touch the database"
And that ended it apparently.
This is pretty basic information hygiene, and if you don’t have enough care to implement even that then why should they partner with you? Sounds like a liability waiting to happen.
They would not give us any indication of what they expected.
> I believe what they are saying is that you should not be able to access sensitive customer data by virtue of “just” being a developer, and you should audit who has access to such data and why.
It should be pretty straightforward for them to provide a "here's our initial base minimum requirements guidelines" and work from there.
The parameters they laid out were similar to "anyone who would be capable of running any code on the database should not have access to any data". Every clarifying question about "what about X? what about Y?" was met with basically, "no, that person could exfiltrate the data". Which, yes, anyone who has access to the data could, theoretically, exfiltrate the data.
I showed a demonstration (from laptop) and the guy got really upset that I had 'customer data' on the screen and in my database. It took several minutes of back and forth to explain to him that it was "faked" data.
They shouldn’t have to spell out how to protect data to you, and simply by asking “what about X” as if you are the only one to ever say this to them or notice that there are of course ways data could be exfiltrated betrayed your ignorance, not theirs.
Did you expect them to say “my god! Of course! Why didn’t I see this before? This regulation we are legally and contractually bound to obey is not perfect, so hey why don’t you just ignore it.”?
And to top it off you started showing realistic looking data without being upfront and clear to them that it was synthetic? Come on.
We had plenty of secure process in place. The moment he heard the word "developer" and "database", he basically shut down and refused any meaningful engagement.
It is not his job to educate on secure practices. It was his job to inspect what we had in place, and he stopped doing that the moment he heard those two words.
"Would having a non-developer administrator connect and process the changes?"
"Developers can not have any access to any production data".
That's not a meaningful response to the question.
> And to top it off you started showing realistic looking data without being upfront and clear to them that it was synthetic? Come on.
I'm not sure how much more 'up front' you can be besides saying "we develop with faked sample data, this is faked and unrelated to or connected with the production system" at the start of that walkthrough. Perhaps if we'd opened up a Northwind database it might have looked more plausible?
No one was expecting them to ignore anything. We were expecting someone to ask "are you doing X? how are you doing Y?", then followup with "you need to remediate these aspects of your system before we proceed further". That's what I would expect.
Also when showing anything that remotely looks like customer PII, it's your job to clarify that this is not real data. Given the precedent, I can see why they thought you were being careless and showing real production data.
It sounds a lot more like the smaller company was interfacing with someone who was not technical or at the very least did not have a clear understanding of the internal processes of the big company, and expected the smaller company to have prior experience of this thing and to guide them through it.
Most likely, the process would have involved the samaller company communicating with someone inside the larger company, like a DBA, and asking the large company DBA to make the changes they needed. In large financial corps, company developers are not allowed to touch the database either, in fact where I worked, we were not even allowed to create a database for testing an internal-facing project. Processes are that anal.
The entire anecdote was about them telling the representative that they DO in fact have access to the production DB. Why would you expect him to act as if they didn't have access after being told they do?
In small companies, a "developer" is, by comparison anyway, a jack of all trades, who is probably expected and required to know how to do every possible job, including database management. So when the large company said "no developer" they possibly meant that, internally, their developers could not touch databases and only DBAs were allowed to, and they expected the smaller company to have similar processes in place. Which is quite unrealistic.
At the end of the day, paranoid as that may sound, the big company person was probably getting stuck on the role description of "developer" and would not have freaked out if the smaller company had used the term "DBA".
Now if you mean "run code" at all that's another story.
Corporate IT is an especially bad place. You can only lose. When things go well they start cutting costs by letting people go and offshoring jobs. And when things go badly you get the blame. The IT department at my company is hard to work with and often drive me crazy. But when I look at their structure it’s obvious that it doesn’t allow for people to a good job.
If there were only 2000 bugs in a full bank migration, they were doing pretty well.
Our fairly small single page app is going on 300 now after some 5 weeks of testing.
The real problem at these banks is the multitude of systems and datafeeds that impact systems downstream a hundred systems long.
The real solution to these big project failures is, in my experience, the phase deployments. With the first step being retiring old systems and consolidating data repositories to reduce knock-on effects.
It gets a bit murkier with executive positions, in my opinion, since the unlimited upside does start coming in to play.
Also IT skills are very transferable between industries so your best employees would avoid sectors with high risk without increased compensation
I’ve done a lot of work in banks and this story sounds very, very familiar to me. Change the details of the system and it could be any number of projects that I’ve personally worked on. I’ve been brought into a number of projects like this at the ‘near completion’ stage, and each time I’ve reported on what I thought the risks were, suggested how they should address them, and advised them to delay delivery until they do. Some of those projects worked out well, some of them completely bombed, but I’ve never been put in a situation where I was made even partially accountable for somebody else’s poor decision making. So if your question is “how do I protect my own interests when my boss wants to be an idiot?”, then that’s the serious answer for how you do it.
As developers and engineers we’re responsible for letting management know if a project is not on track. What they choose to do with that information is their problem.
Discussion: https://news.ycombinator.com/item?id=21849977 (but not many folks discussing the layoffs).
As it says there, in the short term productivity drops, as you'd expect - in the long run, quality and productivity increase.
And even for structural and mechanical engineers, not everything requires reporting. Even when the project is subject to it, any number of institutional and social factors can make that essentially impossible.
What items might we see on such a list? Oh, I don't know... 100% test coverage, perhaps?
In the UK this is called Governance. Fancy word for what is, effectively, liability avoidance.
Sure, here’s another option: hire good engineers and pay them the market rate for good engineering.
We all know what really happened here: difficult engineering work was outsourced to a cost leader.
Even if they had better “regression testing” (as the author laughably says was missing) this project was doomed with this leadership.
There are countless stories of large scale failures involving non-consultants as well as very well paid consultants.
Software is hard and most developers aren't very good. As the complexity of software increases, the number of developers and teams that can manage it becomes vanishingly small; to a point where no amount of money can help.
The best chance for a project like this to succeed is to break it down and migrate over the course of years (which might not always be possible).
this reminds me of https://www.theregister.co.uk/2019/04/23/hertz_accenture_law...
The idea was to have tickets for everything and "golden" ones vichy cost a fortune but are super high priority.
IT was mad at being outsourced and played the game. "super important" thing coming? Ticket. "just a cable, man"? Ticket.
What the idiots in management discovered, is how THEIR company actually works. How important the human relationship is.
One of my colleagues in IT had the pleasure to tell a board member who came to be serviced right oway to fuck off (his words) and to open a ticket, if he knows what this is Then he picked up the phone as someone was calling and turned away.
This is called "work to rule", and anybody who doesn't know the risk and repercussions of such an event happening, should not be in business management in the XXI century. Unfortunately, a lot of people try to pretend that trade unions never existed, and in so doing they lose an excellent occasion to learn how workers actually think and act.
This is great. Usually the board member would get preferential treatment and never experience how things really are.
There are countless stories of large scale failures involving non-consultants as well as very well paid consultants.
I wouldn't call it a straw man. There are more complex mechanics in outsourcing situations. Often outsourcing is not complete. They still keep a handful of inside workers supervising consultants. And often those inside workers are woefully incompetent, ruining the work of "well paid consultants". That happens because the outsourcing only saves those with good connections: the idiot relative of someone with power.
I've also heard many managers in the consulting firms telling us to never go out of our way to fix problems. If the customers want it fixed, they must pay more hours. Except of course when the customers get very upset and demand unpaid overtime to reach some arbitrary deadline.
https://news.ycombinator.com/item?id=15831784
Every time someone I work with casually raises the possibility of outsourcing, I say I'm open to the idea. Then I share that link and ask them to review it before we meet to discuss our options.
Sure but there are undoubtedly more failures involving cost leaders doing the same job as highly paid consultants. It is foolish to suggest otherwise or that there isn't a correlation. Your diminishing returns example might be true, but also doesn't discount the original statement.
Software is very hard, many engineers are not sufficiently experienced at what they are doing and its hard even for those that are. Further complicating this is a cost issue, quite often large scale applications utilize resources that cost a lot of money on the production side and so to save money, utilize a considerably smaller data and infrastructure set and on the testing and development side. Queries are optimized against this smaller data set. The application goes live and suddenly an application tested against a cherry picked data set of x records is trying to handle 10,000 times the data. Then bad things happen.
The lesson is to get good at changes, not to avoid them. Making fewer changes makes you bad at them over the long term.
I’m super curious how customers ended up seeing other people’s accounts. That seems like a massive major flaw in logic somewhere.
Sadly, the article goes into “general history and commentary”-mode way too quick.
I don’t know anything about this particular case, but the usual cause of this is over-zealous caching (usually by an intermediate load balancer or the like) that returns cached pages while ignoring the session data. The result is that one person logs in and gets their account info, and then everyone else who tries to view the same page gets that person’s info instead of their own. For example, this happened to Steam in 2015: https://arstechnica.com/gaming/2015/12/valve-explains-ddos-i...
Of course, there are also other possibilities, but this is the most likely. Alternatives that come to mind (though I’ve never seen any of these actually happen):
- A bad RNG creating the same session token for multiple users
- Concurrency bug in the web server returning results to the wrong connection
- Messed-up migration causing account IDs to be associated with the wrong credentials (maybe IDs from the old and new systems got mixed up)
I would put money on this one.
New architecture, new platform, implementation takes too long, migration and testing insufficiently planned and when push comes to shove even the too low non functional testing goals not reached but still gone live.
With big bang approaches one has to keep in mind that complexity grows non-linear and at a certain size it becomes very hard. This migration was at a size that was getting hard but the approach taken and the way it was executed was clearly (see report) not on the level required. Big bang on smaller scale is the fastest and cheapest. Big bang on larger scale starts having big risks.
A more incremental approach may have been better and cost effective with much lower risks. A honest cost comparison may have shown that - provided the big bang approach would have been properly planed with sufficient large scale integration, migration and proofing milestones and buffers for risks inherent in this one-shot approach.
>All complex systems that work evolved from simpler systems that worked
Being sent to prison for buggy code... that's quite extreme. Are people even fired for buggy code? If we want to get more extreme than firing, how about losing 1 year's compensation in addition to being fired.
...and moreover, what about the usage of third party software? (worse if it's free software...)
Our local Post Office is seriously lacking man power. They pay roughly minimum wage. On top of that, since the postman also delivers money and other special mail (such as serving legal notices), they would be legally liable if the job is not done right.
Is it any wonder that they have been trying to find another person for years!?
Any industry, including ours, has best practices.
Yes and no. Those people are all responsible for doing their own jobs to a certain standard. If a building falls down it is extremely unlikely that it was due to poor workmanship by an individual bricklayer. The root cause will be in the design and the bricklayer isn’t in a position to challenge the design.
It’s like any contract negotiation, you’re always able to push for more favorable terms, but eventually you’re just weeding out people who aren’t smart enough to walk away.
Or perhaps it was the usual incompetent planning of the owners of TSB - banks splitting up is not a new or impossible thing, you know. And complexity of IT systems is the most absurd argument against splitting up LloydsTSB - why would any regulator consider this at all?
Also, that's the whole reason LLCs exist as one of the pillars of our economy. Having so much risk would make any endeavor so risky as to not be worth it.
Criminal negligence also exists, but the way it was phrased didnt seem to imply it.
Responsibility in an organisation is the orgabisation's problem, and people can be fired/sued.
I'm not advocating for that model in software, but it does provide a very transparent way of determining who is responsible. As the EOR, you may not design every member or do every calculation, but ultimately it's your ass on the line. I suspect it also changes your behavior when you know you are the one taking responsibility.
In an ideal case, the IAR was identified at the outset of a piece of work and was then involved throughout the project life cycle (waterfall or agile). Review points included bid release and contract acceptance as well as design and test readiness reviews.
– Thomas John Watson Sr., IBM
People, often contractors, do get fired all the time for buggy code. Companies get sued for buggy code. However, the issue here isn't really the code, it's the management of the project. Bugs happen all the time, and at every stage, during development which is why large migrations like this need to be done very carefully with ridiculous amounts of testing and rollback procedures. Even when you think you've got everything covered, it's amazing the things that can go wrong on go-live.
Buggy code is not the problem, but a buggy process is.
Plausibly you can trash all your customers’ data without affecting the company’s accounts, so I guess there is room for improvement.
We all know how that looks in practice - Experian’s senior management pointing the finger at a single lowly sysadmin. Or VW blaming their emissions scandal on a single engineer.
There is one individual responsible: the CEO. That’s why he or she is paid the big bucks. Maybe make all the C-level jointly responsible. They can share cells afterwards.
No more software = computers are bricks. You go back to manual operations. Emails? Replaced by letters.
Writing "I fail to see any downsides to this" is pretty much equivalent to "of course we should ban cars, that way we stop car accidents; I fail to see any downsides to this".
The culprit was not insufficient testing!
The culprit was insufficient testing, plus lacks of restorable backups, on multiple levels.
Last level of backup: there's 5 million records, use a few dozen reams of paper and print out the account totals for every single account.
For all of the complex systems you can think of (planes, living organisms, etc), the reason things usually go so well is that there are multiple levels of checks and balances. Everything is usually veering towards entropy, and fail-safe systems try to ensure when things go wrong that they fail safely.
You might restore from backup, but the other bank won't.
The issues here are far more complex than just keeping track of the current balance: a first year CS student could do that part.
I don’t see why it shouldn’t be possible to feed data to both the new system and the old system for a while...
Manager: “Perfect. Let’s sell it!”
Scientist: “Well, we’ll need to test it first. We should run some trials —“
Manager: “We definitely need to test more often! But I already agreed to have it out by next Tuesday, so maybe we can just release it and aim for more testing next time.”
In reality, the scientist will say "let's publish! who cares about testing, publish or perish!".
The exact opposite is true. When a process happens infrequently, it's more likely that the people involved will make mistakes and overlook steps. The correct answer is to make changes more often, and to develop robust processes and automation around those changes.
Robust IT needs decoupling and small incremental changes. This will outperform any coordinated release scheduling in terms of reliability, if implemented correctly.
Enterprise tend to favor big bang migrations on a specific date because somebody higher on up set a particular date and everything falls into place with a Gantt chart running waterfall. The reality is that it falls onto a few technical folks to triage a large amount of teams, including the ones from the company they're trying to break away from (which introduces friction). This includes significant risk to the project.
"TSB chose April 22 for the migration because it was a quiet Sunday evening in mid-spring."
This might've gone better if TSB chose months prior to April 22nd for the long duration migration and testing to be completed, and a period of weeks or months for going live post migration. The F5 load balancer (hardware commodity) could've slowly cut over the traffic 10% at a time to the new migration site to get a feel for user experience. Coordination with the TSB network team would be necessary to accomplish that.
It is a tough spot though, I hope the team learned something from that.
I doubt an F5 load balancer would work in this specific case. But there definitely should've been a software router-adapter that routed requests to two systems and converted their replies to a single format. This would've let them migrate their customers in batches instead of a big bang cutover.
That's what I've always done when migrating data to a new banking system.
The API layer made requests to both the old DB and a new DB that had been populated during a small window of scheduled downtime.
We spent a couple weeks/months running checks in production that the old DB and new DB were returning identical results, though still returning the old DB's results as source of truth. Eventually, we flipped the source of truth to the new DB, and some time later decomissioned the old DB.
0 - https://www.goodreads.com/book/show/17255186-the-phoenix-pro...
I don't understand the requirement of doing the "nuclear button" migration, except maybe a shortsighted way of trying to save costs.
We generally aim to leave the old system toggled off for a release or two (allowing us to switch back in the case of a serious defect) and then rip all that old code out in the subsequent release.
I've migrated a lot of "stuff" from one place to another, from filesystems, to databases, to backups, and there is usually some perl running in the background.
Of course the scale is very different. The biggest thing I've personally migrated was in the low Tb of data. Certainly nothing so large/necessary as banking stuff.
Backups!
The risk however is evaluating consistency between the systems before doing the rollback which in this case would probably require a more advanced testing capability than what was available...
[1] - https://www.oracle.com/au/middleware/technologies/goldengate...
I saw https://www.masterywithsql.com/ posted on HN a few months ago, and I'm planning to finally start it over break. But what do I use after that?
The trick is not in the application of highly complex SQL. The trick is in the PROCESS — making it robust, testable, traceable, and reversible for every object despite the complex dependencies. And then indeed validating every step went well, with the validation depth depending on the objects criticality (OK return code vs check of record counts vs check of totals on fields vs check of totals per category etc)
If you‘re really into it, consider the use of automation and tracing IDs for individual loads.
If you‘re worried about the time it takes despite automation: work on the bulk of the data, but use DB triggers to record delta from your snapshot, then treat that „sidecar“ accordingly.
Then: do at least one end-to-end test.
Learn from the mistakes at every single stage, and act on what you‘ve learned.
- learn how to backup the database (I suggest via the official documentation)
- restore that backup (on a different environment)
- validate the restore is complete (compare the data)
- if the backup and restore are good, now you can start learning how to migrate without the stress that you'll lose historical data
- database restore - database migration (your case) - data migration to a different platform and presumably different data model (the bank‘s case, or my answer here)
Not sure which case the question was reffering to, though...
At this kind of scale + uptime requirement, you'd want to do a migration like this gradually - hydrating the new system with data and keeping it up-to-date with changes for a period of weeks/months while doing testing + validation (and ideally also doing a gradual cutover, although that might not be realistic given banking infrastructure/application design).
Probably the relevant approaches to look into after reviewing basic backup/restore are things like log-shipping or change-data-capture, although choosing the right approach would be highly dependent on the underlying technology, architecture, and requirements...
Commonly-held factoid statistic says they have a 50% chance of going out-of-business within 6 months since they lost so much vital customer data. I hope they made backups and didn’t just rely on snapshots, database transactions or journaling.
IT professional prime directive #0: don’t lose vital data.
Alas instead it's seen as an opportunity to push the risk on to developers.
With this kind of response and IR35 coming up I can foresee more TSBs happening.
But I suspect money never reached the developers. Until the late panic stages when it was too late to correct course.
You don't say... (Insert Nic Cage meme here)
Does that seem like a lot of effort for 5 million accounts? Maybe worth it since they were already paying EUR 100 million per year to license the old Lloyds system.
I’d bet they moved from one big honking RDBMS to another. Curious if old and new were different DB vendors, if anyone here has insight.
It's… close to a person-hour per account… At that point you might as well migrate each account by hand.
I don't get how small-value transactions could turn into large-value transactions on migration, though. Hollerith cards? Wrong column numbers in the COBOL program? Decimal point / comma confusion? Whisky Tango Foxtrot?
I think the date was "00YYMMDD" or something like that, time zone offsets were represented as a float-respresented-as-string, but to indicate a negative offset they added 80000 or so...
So if that's what happens to systems designed today, imagine how legacy cruft from decades ago must look. I would not be surprised if the answer to "Hollerith cards? Wrong column numbers in the COBOL program? Decimal point / comma confusion?" was "all of the above, and then some".
I started out working in that era (no COBOL, but all the rest of it). At least some of us were suspicious of data-in-a-few-characters and character column-number based records (ummm, punchcards). Some of us checked all kinds of things on input to try to reject garbage. Spaces in the middle of numbers? BOOM. Something unexpected in the "record type" letter? BOOM. The card reader actually had a diversion output tray where we could spit out the rejected cards.
But, still, lots of bad stuff got through.
Then, in 1967, the world’s first automated teller machine (ATM) was installed outside a bank in north London
I wonder how much rework was needed when decimalisation came a few years later.