The most remarkable legacy system I have seen
thehftguy.com
thehftguy.com
In 2004 an acquaintance asked me for help with sharing an Internet connection to the residents of his condominium after he had failed getting a common router/switch solution to work. The router was not playing ball unless all clients presented themselves in one and the same subnet, which prompted the unmanaged switch to pass traffic directly between the ports, and that was a no-no for historic reasons relating to Windows and its malware of the time. So I repurposed an old anno 1999 ATX motherboard with a mix of Ethernet cards - the board offered 6 PCI ports and the condominium has 5 residents - 256 MB of RAM, and a low-power passively cooled Pentium III to act as router and switch in one, running OpenBSD with some dhcpd(8) and pf(4) to manage clients and traffic.
16 years later this set-up is still in use 24/7 due to ISPs claiming logistic problems preventing them from installing gateway equipment in the building. It serves a VDSL2 line of 60/20 Mbit/s coming in over the POTS. At its peak the computer had 6 years of uptime. It is still running OpenBSD 3.5 due to poor #OPSEC on my behalf.
Was "asinine" the word you meant there? It doesn't make sense to me in context.
It depends where you live and / or you might not notice them if they happen at night.
Even in France which is supposed to be a first world country I experience a couple power cuts every month (twice at night this week-end, it killed my tile generation and corrupted my filesystem on my OSM server).
I probably should invest in a UPS instead of complaining on HN !
Over the years it got more reliable but I'm not sure if that was universal or just my specific area getting a lot more built up.
Every month sounds very bad, even when I lived in a countryside in the least wealthy EU province (Lubelskie, Poland) it was never this bad. Maybe once a year (and usually with a warning that it will happen).
I live in France and my power cuts are less than 10 in +20 years.
Had +5 years of uninterrupted power at my previous place and at the current flat I live, it's currently at +1,5 years.
I now live in Dublin and I've had one in 6 years.
If I had to guess, I would think it's wildly variable based on location and the factors of that location. A more remote community with one main power line through forested area might experience power loss more often, and a larger city with higher density probably has its own problems with load spikes which may cause problems. Maybe a suburb or medium sized city with semi-recent infrastructure is actually the sweet spot for power reliability?
I've got solar, though, and a house battery is very much on the list once my budget allows it. I don't see local system stability getting better any time soon, and likely worse now that PG&E has more leeway to shut down power during wind/heat events and the like.
Power is a funny thing for computer guys.
I envy all of you.
Where I live, short (usually less than 5 minutes) power outages are common throughout the day. Half a dozen such cuts every single day.
And everytime after a cut and the power comes back, there was a 60 second delay for the internet to reconnect, and about 20 seconds more for VPN to reconnect. Annoying especially when working in teams.
It is unusual not to have a power outage during a 24 hour period.
Here's my current setup to stay ahead of the game:
* Solar panels, with automatic switching between the mains and solar when either of them goes down
* Voltage stabilizers
* Surge protectors
On top of all that I'm currently contemplating getting a UPS for my computer specifically, because there was this one time when the solar failed and it couldn't switch to mains because it was a power outage.
One thing I miss from my student days in the US is always-on electricity. During my 3 year period there, there was only 1 power outage, and that lasted less than a minute.
I live in a block with a bad circuit, and we have brief <5min outages at least once a week.
I dont like it because it makes my microwave's clock unusable.
Neighbours one block away don't have that many outages.
I would have though having a second router in front of the ISP router would be the way to go.
OpenBSD has become a little bit more RAM intensive on boot recently due to relinking shared libraries and the kernel to randomize them (Think .o level ASLR), but 256MB IIRC can handle it. Boots will be slow, but if you've dealt with 6 year uptimes in the past that probably won't lead to much of an issue.
We use to deploy those mini tower eMachines to the staff, but we'd rip out the WinModems and replace them with a NIC. So we had a box of about 100 WinModems collecting dust in the corner. The newer machines came with 20+GB HDDs so we'd swap those out with smaller 4-8GB HDDs we had laying around.
At some point a 1800 number got decommissioned but the local number and internal distribution tree were still active in the PBX. I ended up taking a decomissioned IVR and swapping out the Retorrex boards for WinModems and configuring it as dialup internet service and used a bunch of those large HDDs to host an FTP Server on it.
I left sometime in 2005 but my makeshift dialup ISP was still active as late as 2008 and there was an extensive collection of DIVX movies still on the FTP server.
For those interested, I gave a talk on Athena at PyData UK 2018 called "Python at Massive Scale". 4500 developers making 20k commits a week. Codebase with 35m LOC.
The video is here: https://www.youtube.com/watch?v=ZYD9yyMh9Hk
It covers Athena's origins, what it is used for, application architecture, infrastructure, dev tooling and culture.
I understand the value of it, but as an experienced scientific developer and python dev, the culture shock was huge. I don't have fond memory working on it. It is very different from traditional programming environment, closer to a kind of reimplementation of smalltalk env w/ python. I believe an influence was actually an old system that used smalltalk in the 90ies at JPM.
One of the fancy thing was integrating reactive programming, which worked through ugly hacks, at least at that time, by parsing code to detect dependencies. IIRC, it could manage list comprehension but not loops. They also had their own python binary w/ both python2 and python3 in one process.
I remember my first few months unlearning normal Python and figuring out how to build the 'pixie graph', a lazily-evaluated Python DAG suited for calculating financial instruments. It took a while to get your head around this, but when you did it was a very powerful and productive way to building trading and risk management applications.
To get some sense in how this worked, here are two public projects on github with good introductions:
JPM's actual pixie code is proprietary, extremely performant after 12 years of pushing into bigger and bigger scale problems, and is definitely not on GitHub!
I've not heard the term "twiddle scopes" before; in Athena it is "tweaking".
For finance applications (risk) the whole concept works quite well, doesn’t it - combining the natural advantages of code and Excel.
[Disclosure: I was part of the Quartz Core team in 2011/2012.]
I suspect the Common Lisp influence was beneficial.
The weirdest thing, language-wise, that I noticed was scoping. It's neither dynamic nor static scoping, but weird scoping. (But there are work-around to get something like static scoping.)
Outside of the language, the whole CVS-based version control and review process was weird. But understandable as a product of the late 1990s, when review-before-going-into-permanent-history must have been way ahead of its time.
Not the spaces in variable names?
But I was used to different and exotic conventions from the obscure languages I played with over the years.
It was ubiquitous in GS by the mid-2000s, and was rather instrumental to GS navigating the crisis in 2008.
Being able to accurately and quickly compute risks across the entire firm's books, rather than manually merging across separate systems was key advantage, which encouraged JP, Citi and BofA to build a SecDB clone themselves.
https://www.goldmansachs.com/our-firm/history/moments/1993-s...
This would have had me running screaming from the building
AAAAaaaaaa
Imagine the benefit to the business had they just stuck with Smalltalk the whole time! It is remarkable how many businesses could have been quietly making money for years using a high-quality dynamic language such as Smalltalk or Lisp rather than going through the whole C/C++/Java/Python/Ruby/JavaScript/Go treadmill.
But I guess if they had done that then their programmers might not have gotten to type as many {}! Nor gotten to reinvent the wheel so often.
And Kapital was an interesting experiment but rightfully died. Smalltalk just wasn't the right environment for something that complex and scale. I'd expect most lisps to suffer in the same way. And TBH it was more the image based model rather than the dynamic language per se though that was an issue.
There's a point where dynamism becomes a liability and Kapital went way beyond that point. Python fares a bit better but still, a huge effort is needed to control its failings at scale. This isn't helped at all by sticking a DAG in the middle of it as was done with Athena and its ilk. This all can be made to work but in the "cos its Turing complete" sense rather than the language helping in any meaningful way.
Source: I suffered through it.
I had to LOL. You have very nicely put into words an experience I bet many have had.
Why? Eg Scheme or even the uglier Common Lisp are quite disciplined. And Clojure ain't too bad either.
As far as I can tell, Lispers don't usually go around monkey-patching things.
(I realize Smalltalk is great, but my point is it has its issues as well otherwise I recon sooner or later the advantages would be so clear that Amazon or Google or someone would be all over it since it would give them a competitive edge ...)
That was the official story anyway. A lot of things were claimed when I worked there, a lot of it BS.
The culture in banking was definitely not for me: generally fairly smart people, but very perverse culture. This was the only environment I ever worked where people who actually lie to my face to make sure I made mistake when coding and fail. Everything was custom, and sometimes that meant athena, sometimes that meant terrible systems such as a "distributed database" that was a single threaded wrapper on top of an C++ STL hashmap running on NFS. The guy behind it was reworking its own protocol based on UDP instead of TCP because said system was too slow...
There was 0 abstraction, so if you wanted something such as EUR/USD pair, you had to ask the one guy who knew which id it was, so that you could get the data from in code. To this day I am convinced this was done on purpose for job security.
That system you're thinking of was Kapital.
http://www.cincomsmalltalk.com/main/successes/financial-serv...
It's the only Smalltalk usage at JPM I believe, they could have doubled-down on it but instead went to Python and as you say, tried reinventing the wheel in it.
That's not reinventing the wheel, that's building a new wheel you hope performs better/different with different materials and maybe some different techniques, and has a long and storied history of both successes and failures.
Reinventing the wheel would be if they published this system and some other bank recreated it without using what they published, thus "inventing" something that already existed and was available.
If it was like the predecessor system I worked with, it could handle either. But in a list comprehension, it could figure out the individual dependencies of the individual list items.
So, suppose you had a a function that took in a list, and returned a list with some computations performed on each element. If done with list comprehension, and one input element changed, only the corresponding output element would need to be recomputed, and only that would be recomputed.
If done with a loop, and one input element changed, the entire output array would be recomputed (even if all but one element remained unchanged - it would be hard to guarantee that property by code inspection at the time the dependency graph is built).
My first thought was Athena as well, but it could be Quartz (it started in 08-09 at BaML, vs Athena starting in 2006).
Both Citi and Barcap also tried (albeit to less success than JPM/BaML) to do similar things.
With that said, For those unfamiliar - at least two major banks run insanely large production python installs across thousands of developers. Core skills (nowadays) are python and react - and yes, we’re all actively hiring :)
All of the large banks also have pretty good open source initiatives - just check out github repos for JPMorgan, BaML, or GS. We have also been active participants in Pycon (key sponsors and doing sessions, like Steve’s) since at least 2009.
I joined JPMorgan in 2010, as the first bank in London to offer me a Python role. I ended up staying for 8 years, almost all of it working in Athena: Commodities and FX trading; a bit of Equities; then with the Athena Core team working on the machine learning environment.
I learned a ton, worked with many hugely impressive people - on both the trading and technology sides of the business - and left with lots of good memories.
It is not a far-fetched argument that institutes like that are only a burden to society, existing only to widen socio-economical gaps. The point of view of an insider would be very interesting I think.
[0] "There was one remarkable legacy system when I was working in a bank sometime back."
Same team subsequently moved onto Bank of America to create Quartz, which is BoA's version of the same concept.
Then subsequently moved on to found Washington Square Technologies, where I think they're doing something similar, but not attached to a specific bank? Not 100% sure.
I've heard the concept was shopped around to a few other banks as well, but did not take root.
https://www.techrepublic.com/article/jpmorgans-athena-has-35...
The update isn't a job for one person, and I don't think it is finished yet, unless it is ahead of this alleged schedule...
Thankfully, both JPM and BAML seem to have started to ditch the UI frameworks, at least on the web side (in favor of React mostly). I don't know what the situation is on the desktop UI side. But I know what they had before was loathed by many.
Developers hated it because it was like working with one hand tied behind your back. Traders hated it because developers couldn't deliver the best product because they were working with one hand tied behind their back.
I presume the core teams responsible for building the framework hated it because trying to get support for the proprietary UI tools was painful.
I imagine hiring managers and HR hated it because they would blatantly lie about the job - that the role would be WinForms/WPF (desktop) or JavaScript (web) UI development, without mentioning that all of that was really wrapped up by the proprietary Python framework, and very rarely would you get to touch the underlying industry standard tech stacks.
c::a+b
means c is defined to be the sum of a and b and dependent on both. If either a or b changes, c will be marked stale, and the next time you use it, it will be recomputed. It's essentially as simple as that, and is unified with database views - e.g. total::select sum(balance) from accounts
means that any insert/update to the "accounts" table will mark "total" as stale, and the next use will recompute it, but otherwise it behaves like any other variable.I'm not sure if the K4 engine is smart enough to avoid recompute on updates to other columns.
If you want more fine grained control, you can set up triggers - e.g. "execute this code on update to variable a" - where your code gets the changed indices of a list/vector (or changed tuples of a table) as arguments; this is useful if a plain complete recompute is too costly, or you need to do something at the time of change, rather than at the next evaluation.
K2 and A+ had a similar system, each with different syntax.
The Athena team ran out of names and one major component of it is called Bob.
There really are some gems out there.
In some places the web GUI was precisely designed exceptionally well, items placed together precisely, all designed expertly around the specific domain of concern.
On the other hand, there were race conditions, corrupted data, sql injection, mountains of source files that re-implemented the same things, IE6-7 era compatibility issues, etc.
Then there was the mysterious parts of the program which I never understood, like the auto-email capabilities whose scripts could never be located, mysterious mirror servers that I would figure out existed by looking at the ip addresses of odd requests, little bits of domain logic which I would accidentally break and have to cobble back together.
In a lot of ways, many of the problems just stemmed from age, it was an early 90's application in an early-2010's world. There definitely were some not ideal software development practices that contributed to difficulties as well. But I still had a sense that the previous developer had crafted this beautiful, unique, intricately complex and inter-dependent little world.
Sometimes I was sad I had to brush all of that aside even when my changes made things more robust, reliable and compatible with the modern world.
It's fun to gradually discover the model that the programmer was intending to program, then to see where it broke down X months/years later and was fixed with a hack. Then, you get to design a new model that fits both cases elegantly, instead of just a subset.
You kind of feel like the scientists in Jurassic Park! :)
In real life I get "chaos -> order" highs from organizing my collections, tidying up house, doing dishes, etc.
If I look around code base or my house and there is chaos I get stress. If everything is "in it's spot", ordered, understood I literally feel wave of "ease" wash over me, I am able to relax.
I've had exactly that experience. It was an old Apache mod_perl CGI app that used labels and goto statements for flow control, and for well over a decade nobody wanted to touch it because of that. Once I dived in to change it I realized there was a method to the madness. Mod_perl essentially takes your code and wraps it in a sub (function), so if you have subs in your script and use any sort of globals, those turn into function variables, and their use in the subs causes closures. By emulating function calls with goto and labels this problem was side-stepped.
Not how I would have done it (and not something I kept as I revised the script), but what at first appeared to to be complete insanity eventually ended up being based on the environment and the time it was built in. It's a good reminder that often what appears to be stupidity of the prior programmer is actually your own misunderstanding of the problem (and probably poor documenting of the problem as well...)
"Legacy code is code without tests"
Meaning if a codebase has no tests, it's api and behaviours are extremely hard to determine and change without breaking things.
Super good book if you're dealing with legacy systems at work.
We could onboard "interface developers" in a week. As in first time seeing Java and deploying code to production. Normal onboarding time was 3-6 months. We figured we had 10 year backlog of work using the more traditional tools (eg BizTalk, SeeBeyond). Our firm awarded $20k bonus for hiring referrals, demand was so great.
So my stack was a serious competitive advantage.
Alas, it was too simple. Our startup got acquired by Quest Diagnostics. They loved themselves some InterSystmes Caché. Their PHBs (Directors and VPs opposed to the acquisition) just couldn't handle that my stuff ran circles around theirs. So of course it had to be killed in the crib.
FWIW, technically, an "interface" (what healthcare calls something that munges HL7 or equiv) was just self contained code. Straight up data processing. Input, transform, output. Pass in a stream and some context, stuff comes out. Compose these snippets just like in Unix. In development, you'd just run it from the command line (or IDE). In production, we had a spiffy process runner with a spiffy management UI. Built-in logging, metrics, dashboards, etc.
I did some AWS Lambda stuff at my last gig. I absolutely fucking hated it. The managed scaling is maybe nice, but table stakes these days. The programming model is just a kick in the berries.
PS- Word about InterSystems. Of course I tried to be a good soldier. My architecture was more important than the implementation language, persistence engine; I really didn't didn't care what tools we used. Oh boy. Caché might be the worst tool I've ever used. For example, one time a compiler error bricked my entire dev runtime, unrecoverable. The Caché partisan mocked me "What did you do?!" Um, a typo. "Duh, don't do that!!" Apparently Caché self immolation is normal. So keep regularly image snapshots. At the time there was no version control options; you'd export "source" and hope it'd reimport later. Ludicrous.
realtime distributed stream processing in the late 90s
complex event processing before it was called that
distributed application frameworks
- A (much more complicated) make clone
- A test running framework
- A remote session tool
- A test specification framework
- A preprocessing tool
- At least six domain-specific languages for specifying compilers, assemblers, linkers, simulators.
Of course, you also need to know Linux, bash, C, C++, Perl, and Python.Needless to say, it takes some time to get up to speed. On the other hand, you can run some very simple commands and have a bunch of servers run hundreds of thousands of tests on your code, on different OSes.
Not sure when the bank deprecated IE though - it was certainly in use for most internal sites when I left in 2018.
I also heard in the early days it was a bit resource heavy on a user's machine.
I you're happy there, why not. Unfortunately over the last few years I got the impression that loyal employees don't necessasirly see the appreciation they deserve by their employers.
Getting a raise often becomes increasingly hard once you're working there for long enough. Switching jobs or even companies is often the best option to get a raise or better benefits.
There's plenty of big pond companies with enough roles to keep someone learning and improving for decades.
That said, unless I completely change my role and/or specialization, I don't see any room to grow.
I'm a frontend engineer with some fullstack mixed in. Going backend would certainly be an interesting change, but it would be at the cost of atrophying frontend skill/knowledge, the end result being it would be harder to get a future FE (atrophy) or BE (beginner/mediocre skillset) job. Especially since companies these days seem to be creating FE specific interview tracks.
That would be the point, yes.
I definitely did not mean to imply that spending 6 years doing the same thing at a big company is in any way better than doing the same at a small company.
My point was that since a large company likely does allow you to completely change your role, domain and specialization without quitting, you should look at years-per-role and not years-per-company.
They prefer to bring in/hire more candidates from other countries and pay them less than what current employees ask for them, because it is cheaper.
It was old, like "Side A got stuck have to run to the DC to fix it" old and you would get errors accessing those files but not side B files. I'm guessing 15+ years. Weirdly you can grab a PDF off this thing in less than 2 seconds. This had to exist at other organizations but this was the only one I saw.
Now Azure and AWS provide the same service for pennies on the dollar.
* The core Python libraries and third-party modules, with weekly rolling releases of new/old versions for controlled deployment. These came from internal repos (the AIM system you mention) so that security scanning and licencing was checked.
* Application-level code, with Python imports taking source code from a database rather than the filesystem.
These are explained in detail in my PyData talk video, linked in my other comment.
Does this statistic match what others have seen? If so, that’s staggering, and unfortunate. I know there are more sites that do this than there used to be, but I wasn’t under the impression it was such a large proportion.
As a devout Firefox user it is frustrating (especially being told your web browser is no longer supported).
They are setting themselves up to eventually migrate to Edge so that one web browser can access both the old legacy sites and the new stuff.
A couple of years ago I inherited a software engineering department, near the bottom of a downward spiral shipping 25+ years of VB6 code to customers. Brought in external help, and we spent a year and a half rebuilding a modern version. During that project, I saw the most awful code on a daily basis. I still have nightmares about it!
A new "lambda" takes at about 1-60 seconds to propagate to all nodes.
There is spectacular technology developed at banks, you just don’t hear about it because of secrecy.
I don't know if natural logarithms make sense in this context, but if so: using base e 100→5,000 ms is 3.9 orders of magnitude.
Regardless, it is IMHO a reasonably big deal, a bigger deal than 'an order of magnitude faster' indicates.