Four DB2 Code Bases?
perspectives.mvdirona.com
perspectives.mvdirona.com
I love reading about biography like these. I own a couple books, for example, "Where Wizards Stay Up Late: The Origins Of The Internet" did keep me awake for a few days on an early train. This one, absolutely easy to digest, thank you for sharing.
1. So do the four code bases share code they reuse (think of a common repository they depend on)?
2. Are the query planners all different because they needed to be platform specific?
3. Is DB2 still on four code bases?
Would be awesome if we have one central catalog of stories like this one.
BTW, I didn't know much about the author, so I did some quick googling: http://mvdirona.com/jrh/work/
(I hope HN won't flood it to 404! Mr. James, is this on AWS!? It looks like it might be.)
Moody was able to get some pretty intimate insights into the development of a software product at Microsoft at the beginning of the 1990s (1992-1993, I think).
I was surprised to learn afterwards that Moody had (has?) a reputation for being very pro-Microsoft; this book does not embellish the problems the people on the project team face, nor their conflicts.
"The Chip: How Two Americans Invented the Microchip and Launched a Revolution"
How do you write and deploy software for these things? I am so used to just being able to, like, build a Go binary and run it on a Linux box. Are they better in some ways that I am missing out on? Just curious.
Most older programs are written in RPG but there are ways to write programs in other language like C or COBOL. You can use tools like IBM VisualAge to have a "modern" environment instead of hammering away at green screen text terminals.
It's a very odd environment, unlike both Windows and UNIX. There is a job scheduler, output queues, print devices, security levels, objects...
Check out the young i developers group if you want to learn more. They embrace modern tech like node and github
As for writing software, you basically have no choice but to use tooling IBM has support for. There is a emulated POSIX environment (PASE) but it’s mostly used by IBM to port things from AIX - the traditional IBM i languages are RPG plus C/C++ as a newer addition, but they have runtimes for Java, Python, Node and a few others. The traditional languages are the only way you have to make classic green screen apps, while the latter are there primarily to host web applications.
As somebody who has never been near an AS/400, what are those drawbacks? Besides the environment being weird in the sense that it is unlike anything else?
The only bad thing I have ever someone say about the AS/400 was a supervisor I worked for during my training who had at one point written RPG programs, and she said ... some pretty unfriendly things about RPG. Which is not say that's the only problem; like I said, I have no knowledge of the system beyond what Wikipedia tells me about it.
As for writing software, you basically have no choice but to use tooling IBM has support for.
Is that really the case now? Even back in the 1980s, other software publishers made great tools and compilers for OS/MVS mainframes. Boole and Babbage's Resolve did a lot of cool stuff, including giving you the mainframe equivalent of "su" right down to the TSO terminal level. And we didn't even use IBM compilers; Capex had optimized compilers with great system dump formatting.It was developed in PL/S, a PL/I dialect, back when C was only relevant inside AT&T walls.
The binaries use a bytecode format known as TIMI, and IBM i contains a kernel level JIT. So any application gets AOT compiled at installation time, or when you feel like refreshing their native image.
The filesystem is follows a database model known as catalog.
Main programming languages on the classical mainframe environment are RPG II, Fortran, Cobol, PL/S, C, C++ and Java.
Yes, IBM i Java makes use of TIMI bytecodes as well.
Then there is the POSIX emulation layer to help port UNIX applications into IBM i.
At old good mainframe fashion, these environments are each its own little world.
When the need to modernize IBM i came to be, IBM started porting PL/S code to C++.
If one wants to actually target native code and not TIMI, there is a C compiler called Metal C.
This is just about IBM i, IBM z and Unisys ClearPath are also quite interesting. All written in systems programming languages much safer than straight C, and with lots of nice features which UNIXes still haven't adopted.
I've been trying to find an emulator for Unisys for years, as I've come in contact with those systems twice in my career, but both times it was "oh, those are used by [hushed tones] the legacy team" which isn't good for learning about how they work.
Do you know of something akin to Hercules, or even eBay something, which would allow a sufficiently curious person to kick to the tires on a Unisys OS?
Also, for future comers-along, that link has all the same language as the download registration, minus the download registration link :-(
Thankfully, I was able to find the real one from the Unisys MCP Express youtube channel:
http://www.unisys.com/offerings/clearpath-forward/clearpath-...
Separately, the idea that one runs OS 2200 on Windows using (apparently) .NET 4.5 blows my mind. I can't wait to see what's up with it.
https://www.itjungle.com/2017/12/13/2017-ibm-year-review/
https://www.itjungle.com/2017/06/21/seven-bright-spots-ponde...
Also, if you read between the lines, there's a ludicrous amount of infighting in IBM, which is probably how they ended up with four code bases. Still happens to this day.
You might be surprised then, because that's not silly at all. DB2 may hang on in the Back Office, the least technically sophisticated bit of banking for various reasons for a while longer, but it's long gone from the Front Office and Postgres is one of the replacements. I don't think anyone is starting new projects on DB2 even in BO.
I wonder if anybody has done a write up on this oracle/db2 versus “others” in the super-high-value back office scenario.
The corollary to this, of course, is that anybody who isn't involved in projects of that scale, is going to be far better served with Postgres/MySql.
It will be interesting to see what the uptake on products that try and fit in the middle - like Amazon Aurora - High scale and affordable.
Given those experiences, I personally stopped optimizing for today's flush tech budget. I now budget for what I know is a realistically shitty future state of affairs where we have to keep the lights on [1] while spending as little as possible.
[1] The one worse thing you can tell a trader during a 2008-level crisis than "I need more money for databases" is "you can't trade, we don't have enough capacity to handle the additional market vol".
There are far fewer people however that do really need it, than believe that they do... The vast majority of Oracle shops could probably switch to Postgres for less than their annual licenses, then start saving the money from the second year on...
My rule of thumb is that Postgres gets you 50% of Oracle for 0% of the price. Most of what most people need is somewhere in that 50%. SQL Server gets you 90% for 25% of the cost. I think Oracle will lose a lot of seats to SQL Server once it's stable on Linux.
Can you elaborate on what kinds of things fill the 40% difference between PG and MSSQL? Having used both PostgreSQL and SQL Server, my experience has been pretty much the other way around. When I need to do something with data in a table, PG probably has a feature allowing me to do it easily, and MSSQL probably doesn't. The simplest questions asked on Stack Overflow are answered with functions spanning tens of lines, often with various caveats.
The one area where MSSQL is clearly better than PG is GUI tooling, but even there I'm not particularly impressed. More than once have I read that someone absolutely loved SQL Server Management Studio, but the one conclusion I can come to after using it is "it works". For example, "Select Top 1000 Rows" is a fine view with autocompletion and everything, but then "Edit Top 200 Rows" has an SQL panel hidden behind a context menu, without autocompletion, and when you execute your modified query it rewrites it and removes all comments.
SQL Server is well ahead of PG in query parallelisation and complex replication topologies, even as of PG 10.1. SQL Server is also ahead in change data capture, but PG 10 narrowed that gap. Managing a large user base is easier on SQL Server. If you have those use cases it is well worth the money.
If I need to do something clever with the contents of the table, SQL Server's embedded R is very good.
Having used both PostgreSQL and SQL Server,
Well indeed, most people only use a subset of their DB's full feature set, but it will be a different subset for each person :-) I don't really use GUI tools so I can't comment on that but I've heard that pgAdmin 4 is awful.
> I've heard that pgAdmin 4 is awful.
I've heard the same. I wasn't a fan of pgAdmin 3 either, so I've stuck with the psql command line.
I love that in the HN community we can have these conversations without it descending into a platform war :-)
I did alot of Oracle up to SQL Server 2012 then just didn't need it any more, SQL Server + PG 9 and now 10 hit all my use cases for a tiny fraction of the cost. Now I wonder why so many people still pay so much for it.
In case you haven't heard of it, I highly highly recommend pgcli[1] if anyone reading this is doing the same as you and me but hasn't highly customized their psql shell yet. It's one of those tools I wish I had known about earlier.
For one, when there is no space left for wals, your instance juste crashes. In sybase/mssql your processes would just be suspended untill you add some disk space.
Also, spreading your databases on several disks is way easier. On postrgres adding tablespaces makes pg_basebackup lose functionality.
Those are some several limitations that the developper might not see, but they impact availlability.
Also, reclaiming unused space online can be tedious, and failover from one instance to another is not easy. You have to use 3rd party tools such as repman to do so.
All of thoses missing features have caused several outages and should be addresses in my opinion to be able to meet the needs of the most demanding workloads.
Still, I view postgres as the best solution for most business cases.
Are these not equivalent from the point of view of external clients?
And PostgreSQL lacks tooling at the same level as APEX, SQL Developer, JDBC/.NET drivers (specially bulk operations and distributed transactions).
Yes, but maybe it's a good thing after all. pgSQL is small, and mostly clear and restrictive, there's no bunch of imperative features and dark corners. The main point of PostgreSQL always was extendability, so it supports a broad list of programming languages for stored procedures.
To address efficiency concerns, executor's JIT is on the way - https://postgrespro.com/roadmap/56509
i know a lot of companies that are still on 11g. and this is not in any way state-of-the art, heck it does not even have a limit clause. It's hard to actually trust your statement that 12c is so much better, when 11g was basically trash.
Because it was lacking a tiny piece of syntactic sugar? Just use ROWNUM.
1. https://github.com/django/django/blob/master/django/db/backe...
In any case your original quote was "Because it was lacking a tiny piece of syntactic sugar?", and I was merely pointing out that it is a bit more than some syntactic sugar.
Who are these people, really? Noone tries to write language agnostic code then gets upset when it won't compile as both C++ and Java. People who try to write code that will compile on both Unix and Windows accept that they must make compromises... Or they have lots of platform-specific code hidden in a framework somewhere. So that is a silly argument.
If you are investing in a platform - and I don't just mean licenses but TCO, incl hiring experienced people - then it makes sense to exploit all of its features to their utmost and extract the most value. If you aren't doing that then you are coding with one hand tied behind your back.
As a cheap, easy example of such quirks you can look at software that has to work around them, like Django.
Sure. But the point I was making is that if you really do need Oracle's higher-end features, there aren't many or sometimes any substitutes. How long have the PG crew been trying to get AS OF working? But Flashback's been in Oracle for years and if you need it, you need it. And if you don't, then Postgres is a pretty safe choice. But Oracle is very far from "basically trash" because some guy doesn't like some bit of syntax!
But licenses can also be a bit of a nightmare to manage. I've seen many cases where licenses were violated due to needs for testing/development environments. It can be difficult go through the purchase process in big companies, specially if the software in question is not wildly used. Handling licenses can also be hard technically, when your licenses are tied to hardware (a MAC address for example) or configuration (hostname or IP). When migrating some legacy hosts to VMs, I had numerous headaches when dealing with flexlm servers. It can also be really annoying when you don't have perpetual licences and the editor goes down or is not able to generate a new licence for the (obsolete) version you are using.
Basically, you must add specific processes and even additional infrastructure (inventory, CMDBs) just to manage it. It can also slow down development.
I understand that in many fields, no viable OSS alternatives exist, or even could exist, specially for specialised software.
But if I can avoid proprietary, I generally do.
I'm not against paying for software, but other for of payments, like a support contract, are generally far less annoying.
Other popular big databases are Ingres and Informix.
At a big airline I worked at, we had three different major databases on our Unix systems: Informix, Ingres and Oracle.
The mainframes had a few more.
Here's a HN poll on what dbs people use. https://news.ycombinator.com/item?id=2684620
Sybase were always very willing to customise their product for financial services customers. Someone I knew there told be about the changes they’d made to their DATE datatype to make it easy for one bank to handle bonds from the Napoleonic era for example.
Yes, Postgres is often used and in-house developed DBs are fairly popular too. On the commercial side, Sybase still exists too! And of course SQL Server, tho' I am not aware of anyone using it on Linux in production yet.
The inertia keeping DB2 in the BO is all the third party apps running on it - if they move they will need to persuade all their vendors to move en-masse to whatever comes next. FO software is a lot more bespoke and anything from outside tends to be chosen by people who know what they're doing.
I have to believe a lot of the pain coinbase is experiencing stems from their decision to use mongodb
Almost certainly. I would go in-house before Mongo any day of the week. That's if Postgres wasn't already a better JSON store than Mongo is...
Teradata!
And IMS, but that's not relational.
High-value transactional environments most often neither relational nor even running on operating system (in modern sense at least), like zTPF "because a real operating system is too high level and therefore too slow for real transaction processing needs". Current users of TPF include Sabre, VISA, American Airlines, American Express, HP SHARES (formerly EDS), Holiday Inn, Alitalia, KLM, Amtrak, Marriott International, Travelport, Citibank, Citifinancial, Air Canada, Delta Air Lines, Japan Airlines and many others.
'Robotics' solutions are what new projects are, which are simply layers on existing solutions. Clamp over existing things, rather than re-implementing them.
It's funny: a few years ago in the FO we heard a rumour that the bank was starting work on robotics, and were curious and some of my colleagues were keen to get involved. Robots! We didn't realise at the time that "robotics" is the word BO use to mean "any sort of automation, whatsoever". Even VBA. Banks are now trying to recruit graduates with the promise of working on "robotics", ho ho ho.
The access to quality tech BO has is terrible. Lack funding for people that can DIY and lack time and budget to experiment.
Fintech is not driven by tech but by regulation where banks are forced to make an investment in change, not just talk about it.
I spent a few years managing DB2 instances at a previous employer and I've seen it hold its own against mixed workloads (i.e. shitstorms of bad-behaved applications and people running ad-hoc queries) that I don't think even a well tuned PostgreSQL could handle today.
DB2 has some really annoying quirks(1) and that IBM'ish feel that's very hard to explain, but it's also very understandable (from a tuning perspective) for such a complex piece of software.
DB2 for Windows seemed mostly crap compared to Linux or AIX, though.
(1) For example, at least since I've last looked into it, DB2 had no decent MVCC(2) and it was a real pain to deal with all the lock-related problems that would arise from large write/update transactions running in parallel with small reads.
(2) At some point (9.5, IIRC) they tried to simulate MVCC by letting transactions read previous row versions from the transaction log. This seemed to come with its own set of significant drawbacks.
The best part is how resource light it is. I was using Informix on a dual 200mhz HP K class box until about 8 years ago. HP literally just gave us a RP3440 to get us off it...
Back in '86, DB2 lacked a mechanism to run ad hoc SQL from a file. They had SPUFI for terminal use but apparently nothing you could just wrap in JCL and fire off.
Informix developed a product called Batch/SPUFI to do exactly that, along with porting their Ace report writer to use on DB2.
We got it working, but the company abandoned the entire mainframe product line soon thereafter.
But even I was initially excited - the feature list looked cool - multimedia, rich/extensible data types (called blades?), OO/table inheritance, built-in dynamic web page generation, etc. But the performance was absolutely terrible and all the advanced features turned out to be less than useful or slightly broken. I remember my disapointment when I first tried out the "search for images LIKE a reference image" feature and it matched a butterfly with a racing car. We struggled for a few months with the general flakiness and horrible performance of Illustra and eventually convinced them to allow us to use Informix instead and went back to storing images as files served by HTTP and using PHP or Perl for the dynamic HTML generation. The Informix sales dudes seemed convinced that Illustra was the future.
The performance was much better but for the web apps we were writing, we could have just stuck with mSQL or MySQl and saved a lot of time and effort.
Yeah, DataBlades. I too was burned by them in the 90s.
Sadly, he never would have been brought in but for the financial disaster that was Informix's acquisition of Innovative Software; the resulting financial stress necessitated "new upper management with Wall Street credibility". Hence came Phil White and his minions, and the rest is a sad history.
Last time I used it, it was still an independent company and in all these years I thought it was gone.
Question was about how it was doing on the market, given that I never saw it again deployed.
My favorite line out of it: "so I ended up taking a long weekend and writing support for a primitive approach to supporting greater-than-2GB tables." Wow.
It's the esoteric platforms like mainframe or as/400 that need custom code bases
> At the time I was lead architect for the product and felt super-strongly that we needed to address the table size limitation of 2GB before we shipped. I was making that argument vociferously but the excellent counter argument was we were simply out of time. Any reasonable redesign would have delayed us significantly from our committed product ship dates. Estimates ranged from 9 to 12 months and many felt bigger slips were likely if we made changes of this magnitude to the storage engine.
>I still couldn’t live with the prospect of shipping a UNIX database product with this scaling limitation so I ended up taking a long weekend and writing support for a primitive approach to supporting greater-than-2GB tables.
It's amazing how often one person with a solid grasp of a problem can run circles around teams that are shackled by process and cya and politics.
This raises a question which is interesting to me (not knowing anything about mainframes).
How long is the life of the IBM mainframe going to be? Are there specific plans already in place to replace these systems on a known schedule? Or are they going to be around indefinitely?
I (kiddingly) poke at the people at work that I'm starting to teach my 6 year old grand child COBOL and CICS, they will have lifetime employment.
I think a typical mainframe customer is more likely to just cease to exist due to random chance/bad business decisions/market shifting/stuff than it is to move off mainframes to another platform. We're talking about very long time spans and very stable systems here.