The 74,000 numbers of Barclays Bank
shkspr.mobi
shkspr.mobi
What I don't understand is how that comes to be. Are just people who are not skilled enough applying to work at a bank? Do they have some weird selection criteria that tends to discard all the good people?
You'd think that they'd hire some good people by accident and those people would be more than capable of running the whole operation, but apparently they don't.
And if you don't save it you'll get a bollocking for showing initiative. Because, apparently, our fetishistic obsession with WIP-limiting means we can't walk and chew gum at the same time.
I've never worked for a bank but this all sounds very similar to a (mercifully brief) contract I did for a multinational in a different sector many years ago.
Most people dont like that kind of bureaucracy.
Add the usual outsourcing issues, managing communications across 3 countries and 3 opposing project managers and little happens.
Banking does have good developers but they mostly aggregate in investment banks and hedge funds, that are inhouse and have more means. Retail banking is unfortunately not a great industry financially or technically.
maybe a server-side lookup would have been a little bit nicer than shipping the whole list to the client side, but a static client-side app can be stored in the CDN layer, doesn't require provisioning any server resources, and sidesteps any security concerns about servers processing user input.
it would have been nice if they'd cleaned up the user input before querying the dataset, but overall this looks like a clever solution to release a feature with the smallest risk and maintenance profile possible.
Downloading a big JSON file locally is underrated. I implemented the exact same strategy on https://overframe.gg, a Warframe database. It allows a fantastic user experience ... the tradeoff being the slow loading time (which is invisible to most people on good connections).
Overframe was a SPA though and had to lose the approach on most pages because of SEO. The usage were too different. Before that every single page on the site was clientside and available immediately after initial load.
For a tool though, especially one that relies on user input (and thus will load in the background before your user starts interacting with it), it's a great approach as long as your data is somewhat bounded.
Back in the day I’d even do super fast geo proximity locally by including a json of the locations of first three digits of zip codes and truncating the zip code of the user to make the comparison.
edit: checked the comments of the original post and they apparently used perl to produce a 11kb regex.
A little regex to remove out anything non-digit would leave you with just the numbers to compare to the numbers in the list.
The fact that there isn't any user input sanitization makes me inclined to believe that this implementation decision may not have been well-thought-out (which I'm guessing was the author's point exactly)
most likely, this page started out as a list of all 74k phone numbers, and at some point somebody built a lookup tool on top of it. there's no reason to assume that keeping the full list secret was ever a design goal.
>why not POST the number to a service which can be updated?
Making life easy for a web developer, rather than saving time and money for users, is unethical behaviour. The developers are putting their (and their company's) comfort over that of the people using their system.
Perhaps that's OK on a day-to-day basis. But this is a system designed to help people in a stressful situation trying to ascertain if a caller is genuine. That's not a situation where you want to needlessly pipe a couple of hundred KB over a dodgy 3G connection.
The bigger problem with this approach is that it exposes a canonical list of phone numbers that are guaranteed to not be flagged as fake by the tool itself. That isn't likely to be a _massive_ concern here specifically, since it will always be easy to find a valid customer service number without that access. But I'd expect a team looking at dealing with fraud or scams to be pretty conservative about information disclosure generally – it would be a good argument to implement this as a backend service.
I fully agree that the criticism in the article is not warranted
It's still lazy programming...
So, even with the code to turn it back into a list, handle the leading zeros on many numbers, etc, it would be under 10k total, much smaller than the crazy regex.
Not general purpose, and a bit silly, but you can get the list pretty damn small.
[1] Like: "1234, 1235, 1435: becomes "1234, 1, 200"
[2] The < 4k gz file that you can reproduce most of the list from: https://github.com/tyingq/barclays/blob/main/out.c.gz?raw=tr... Gunzip, then read each line, then add the prior read line to it, prepend a leading zero and print. "Prior read line" is 0 for the first line. There's a small amount of the list you would have to hardcode or handle another way...the numbers without leading zeros, total of 63 numbers.
That seems like a nightmare to do searches on. I'd probably compress it with a prefix tree or similar. In terms of off the shelf solutions, sorting the list, saving as a newline delimited text file, and then compressing with LZMA2 gets you down to 22KB.
https://www.barclays.co.uk/content/dam/json-files/TelephoneN...
content-encoding: gzip
content-length: 186136
Don't be misled by the fact that your browser transparently hides the compression from you.
Also, there are a fair number of duplicates in Barclay's original list, so it's even more lazy an effort than that :)
Does this really need any more UI than the browser’s chrome?
Not every web user knows “ctrl+f” or “find in page” on their phone, but those that don’t, do still know how to scroll 74,000 lines of ordered numbers (which is surprisingly easy and fast - i just tried scrolling 74k lines in excel on mobile).
The one use case i’m thinking might be confounded is a screen reader for the visually impaired or blind but even then, i’m thinking this use case is probably covered?
It's not technologically hard to implement a search, so I don't see a reason to only show a big list.
Having all the numbers visible in the page would certainly help it show up in web-search results for those phone numbers, though.
Overassessment of general literacy (the reading kind, not computer or maths literacy) is exceedingly common in technical fields. Averag computer skills are abysmal.
About 5% of computer users have "advanced" literacy, defined as "Some navigation across pages and applications is required to solve the problem. The use of tools (e.g. a sort function) is required to make progress towards the solution. The task may involve multiple steps and operators. The goal of the problem may have to be defined by the respondent, and the criteria to be met may or may not be explicit"
Scheduling a meeting room, or determining "what percentage of the emails sent by John Smith last month were about sustainability" are examples of level-3 tasks.
A quarter of the adult population cannot use computers at all, 14% are at "below level-1" skills, and 30% can only perform very basic level-1 tasks, for a total of 70% of the population which has only very basic skills ... or less.
It's easy to over-estimate the general literacy and numeracy of the population, especially if you yourself are college-educated and work in and/or with information technology.
The United States performs one of the most comprehensive assessments of adult literacy. *The key lesson for me is just how limited it is.*
https://nces.ed.gov/pubs2019/2019179/index.asp
The findings correspond highly to a study of adult computer literacy amongst 20 countries by the OECD:
"Skills Matter: Further Results from the Survey of Adult Skills" http://dx.doi.org/10.1787/9789264258051-en
Computer usability expert Jacob Nielsen has a discussion of this as well: https://www.nngroup.com/articles/computer-skill-levels/
I've discussed this as "The Tyranny of the Minimum Viable User", which both notes that much of the population has very basic skills, and that this also hampers the very small minority who do.
https://old.reddit.com/r/dredmorbius/comments/69wk8y/the_tyr...
For a text file? Why?? Wouldn't the worst case be the size of the file?
The text file also gets transformed into a html document with DOM (inspect one in Firefox to see), which likely has its own caching tricks for performance.
It surprises me that you need to ask that
While other comments have already mentioned that it's actually "only" ~190kB which get transferred thanks to gzip compression, there is still a lot of potential for a smaller transfer size by simply using brotli instead of gzip for compression. With brotli the size of this file is <60KB, so only less than a third of the file size of the gzipped data.
Using brotli would be an easy win and improve performance of loading all other assets as well, without the need to change anything in the actual implementation of the website, however Barclays Bank is probably still running a webserver stack which doesn't support brotli yet.
In your scenario (“ just tell the customer to always call them back”), which number should the customer call back?
It can take ages to get through back to an agent and it may not even be the same one or even in the same department.
187k seemed quite expensive, the file actually drops to 49k with Brotli and removing whitespace.
Sorry if this put a backend team out of a job, but I definitely approve of this design. The other number normalization defects could just as easily be fixed client-side
I stumbled on a class that was very short and simple. If you know enterprise code, that should immediately make alarm bells ring.
And yes, indeed, all this code did was provide a (world-visible) API endpoint which returned the result of "SELECT id FROM <table with millions of rows>" as a JSON array of numbers.
IIRC it was used for something theoretically reasonable like "figure out if this search index contains all rows from that table".
But I'm sure there were less hacky ways to achieve the same result. Maybe even put permissions or firewalls on it.
We couldn't immediately disable or firewall it as it had embedded its roots deep, and a lot of things calling it didn't know they were calling it.
e.g. changing 02071160299 to 82071160299 still appears to pass the regexp.
The "correct" regex is 11KB and can be found here https://gist.github.com/jes/e678e4300d1cfcbcc12b46aaa7e58e30
I agree with another comment here that the 1.3MB JSON is a better solution though
Also something really weird: when I view zimpenfish's comment in the post it shows as 54 minutes old which makes sense, when I click it to view direct or reply it shows as 2 days old which is really wonky. Any idea what is going on there? https://i.imgur.com/DvBDcIL.png
Might be something that got resubmitted from the second chance pool and then automatical rewrite only did half the work it should or something is cached or something similar?
There is a (set of) protocol for signing telco connections that was developed:
* https://en.wikipedia.org/wiki/STIR/SHAKEN
* https://www.fcc.gov/call-authentication
Anyone in the telco space that can comment on how much this is implemented in their geographic area (US, EU, APAC, etc)?
As far as UK banks go, they're one of the more secure ones - giving you out a pinsentry 2FA device for accessing your account, setting up new payee's etc.
You'll see the same thing on different sites with credit card numbers. See, every credit card issuer uses a scheme called an "issuer identification number" to tell you what type of card it is. And, there's a list of them [0]. So, for instance, I can tell that any valid credit card number that starts with 4 must be a Visa. So, why ask me what type of card it is, when you can figure it out yourself?
Those rare sites where you actually want to put your Social Security number in have these sorts of issues, too. Some of them require dashes, some of them don't. Some of them will only allow you to enter 9 characters, but won't tell you not to use dashes. It's infuriating, really. And, any site where you enter an account number can potentially have these types of issues.
It's not like being just a little more flexible with input formatting is a major feature with huge implementation costs and security risks, either. We're talking about failing to take a few minutes' thought in order to be a little more kind to the user, and that's just inexcusable, IMO.
---
[0]: https://stevemorse.org/ssn/List_of_Bank_Identification_Numbe...
Just in case somebody considers implementing phone number parsing and normalizing: Just don't and use Googles libphonenumber, which offers solid, ready to use implementations in various languages (there also exist ports for other languages too): https://github.com/google/libphonenumber/
+49 (0)30 23125 123
Your naive approach to just replace non-number characters would lead to an invalid number in such a case. Using libphonenumber you'd get a correctly normalized phone number while the user doesn't get bothered for inputting a technically invalid number: https://libphonenumber.appspot.com/phonenumberparser?number=...People should NOT be trusting calls simply on the basis of the number presented and should ALWAYS validate the caller by another means.
Spoofing of numbers is a reality on communications networks worldwide.
This applies even in places in the US where they have STIR/SHAKEN tokens. STIR/SHAKEN tokens are mostly about the reactive tracing the origin of the call, not necessarily about proactive blocking spoofing.
The UK, like 99.999% of other telecoms networks in the world doesn't even have the equivalent to STIR/SHAKEN and is reliant on carriers implementing (or, more often not) various limited and mostly informal possibilities.
You would have thought a company such as Barclays would be more knowledgeable than to go around saying that you can trust a CLI if it matches one on a list we publish.
It was based on some completely Byzantine web framework and was incredibly slow. I checked the web requests it was making once while trying to add a user to a group and found that the UI was loading every group from our directory server (~22000 groups) into some JS object that then loaded them all into an html list element. Then it had a separate request to the web server for field validation for some reason.
It also didn’t have an API, much to the chagrin of every attempt to automate any infrastructure system.
Banks never did. This is par for course. Just be lucky they didn’t send it in FIX, because you know a bald guy clinging to their 2010 Blackberry “because of the keyboard and emails” really wanted to send they payload over the FIX protocol
And if for some reason you can't - please don't return your data in a proprietary format based on a clever hack. Just use gzip.
I can see the point of being efficient but this doesn't seem that bad all things considered. At least the list is also human readable unlike a regex or custom compression.