Ditching a Language
blogs.perl.org
blogs.perl.org
I have done severals. Changin languages, database engines, architecture, styles... even do the same thing several times in the same project!
And each time, I see a lot of code reduction. Specially, if I can change the language!. In one of them, I reduce by the butloads some badly C# project to python. I meant, close to 1000 files to few dozens. Yep, if we are talking about spaguetti, then that could compresse that well ;)
In fact, I think the action of change language (or move to the most recent versions with the most modern libraries/dependencies possible) is the SIMPLEST way to reduce the load of the job.
I do this all the time. Each time obj-c give some new trick that cut code, I apply it as fast as possible, across all my codebase. I learn to do that after the most insane upgrade/rewrite from .NET 1 to 1.1 then 2.0 that kill us because the boss wait to much.
The BEST way is obviously not bring (again) the same mistake that create that monster in first place. THAT is what make hard/impossible the task in the average corporation, because are the cultural problems that cause the biggest mess.
Also, is necesary to keep old project alive, and (this is something that bite me once) truly have the most hyper-perfect data upgrade/syncronization possible to minimize downtime and have real data from the start... real but clean! This way I create a 3-tier version in visual foxpro+sql server of a fox/dos app that was deployed in +2000 places with non-tech people before the internet, sucesfully (bar the first couple of tries ;)).
The codebase he describes is an eerily accurate representation of where we started, with the added complication that it was built by a sole developer than wasn't using any version control at all.
There are roughly 600k lines of Perl code in the back-end alone, but because there was a lot of duplication in lieu of version control, I have no way of knowing how much of this was actually in use. I suspect roughly 100-150k.
Our approach was pulling the platform apart into distinct (Ruby) services and putting HTTP interfaces around some of the legacy services where possible. We've ended up with < 15k of Ruby, including the front-ends. It's not perfect but there haven't been many major issues in our pilot release, and the team is happy. Fingers crossed.
While a lot of the commenters here seem somehow outraged at the computations, 5.5 man years seemed like a very aggressive schedule for rewriting a 1MLOC system, even if the end result were a 100KLOC one. My initial guess would have been 10 man years, and a cost in the millions.
So here you have what was probably a 150KLOC original system. How many man-years would you guess the rewrite took?
The truth of it is, in our case it isn't (yet) a full rewrite, and there is still a lot of functionality tied up in the existing codebase. So it's not easy to answer the question about man years, but I would guess around 1.5.
The biggest wins for us in terms of lines of code were not the language (we did consider sticking with Perl), but re-assessing the business logic, ridding the codebase of legacy junk and using existing libraries instead of hand-rolled solutions.
However, I would mention that this doesn't match the alleged situation in article, a million lines in heavy use.
Your app sound like it's real complexity was considerably less. Given that at least complexity goes up with size of the code, your success still might not mean that diving into the situation described in the article would be a good idea.
That said, the code we are replacing is in constant use and powers all our online stores, the main source of revenue for our company.
It would seem like the obvious answer would be to not sit down and rewrite the whole thing from scratch, but start replacing pieces (with whatever language they think they'll be successful at).
The idea that they should somehow just be stuck forever with a shitty mess of a perl application seems incredibly defeatist.
If you spent money on something and that something isn't worth what you put into it but still is worth a lot, you want to protect that investment.
As far as incremental improvements go, rewriting some part in another language seem pretty bad. I mean, if you are being incremental, then you have to make changes that might interrupted in the middle and then you'd be saddling the system with two different languages.
You could just easily rewrite the worst parts to conform whatever existing or new standard you have. I'm not fan of Perl but I'm pretty sure you could at least create a subset that would conform standard object-oriented practices and not have the problem of now having a system written in two different languages.
I don't think the software has any intrinsic value; it wasn't constructed from precious metals, which could be re-smelted and sold for scrap. If continuing to develop it/maintain it is costing them money (either in real terms because of development costs, or by lost opportunity costs), than they should look at how that's trending. At what point does the thing cost more than it's worth to them? (Maybe it never does, I've had banking customers who spend millions of dollars a year keeping a thirty year old, p.o.s. Cobol application running because it's the backbone of their operations).
And I didn't mean to imply that they should rewrite parts of it in another language, just that they should start decomplecting pieces of it so that they can replace those pieces with better-designed ones. If they want to stick with Perl, because that's what their expertise is, than I think they should do that. I would agree that it's probably counterproductive to take on both rebuilding the application while at the same time switching to a new language.
But unless the entire application is passing around internal Perl data structures, it seems crazy that they can't identify edges to the application functionality, and start to peel those edges away and encapsulate that functionality in a better way.
It goes more like this:
We have this big honking system that nobody really understands. We need a system that does X (where X is the list of features we think the other system does that are critical. X is subject to change as we discover other features to add and features that actually weren't necessary.) Please implement X this way.
And that can easily become an exercise in extreme frustration.
They're wrong of course, but this is how it is done, at least in government contracting circles. I have literally seen developers required to get out a RULER and measure user interfaces as they appear on one screen so that they can exactly duplicate them on newer higher-dpi screens without changing the appearance or functionality whatsoever.
Everyone presumes the million-line system is one which could be implemented much more simply... there are some systems, however, that need those million lines. Rewriting one of those in a new language is a whole different level of nightmare. Especially when the requirements are 'make it work like the old system' and nothing else, and the requirements for the old system are 30 years old and not even close to portraying the system as it stands currently... But hey, to avoid 'being the next healthcare.gov', political types will push anything out the door and just shoot anyone who points out problems in the back.
And this is the same damn trap all neophyte developers fall into. "Let's rewrite!"
Once that first wave of business requests and demands comes along, your precious sandcastle will crumble. Because the business team is fickle. And they stick you with deadlines. And then, mid-deadline, they change their mind. Or are forced to go a different direction because some shit government law is passed that requires you to broadcast your service requests with encrypted messages tied to pigeons (because, realistically, that's how the government does APIs).
Here on the internet where everything is made up and the points don't matter, you can get away with rewrites. Agile not working for you? Let me introduce you to the CADT model (http://www.jwz.org/doc/cadt.html).
All rewrites become tomorrow's bug-infested legacy ghettos.
(Don't listen to JWZ, at least on business/strategy. His employers don't exactly have a great history of success on that front)
I cut a lot of stupid boilerplate and was willing to cut features nobody needed. I added a lot of new, good features (not so good, too).
It was still not perfect, these days I would be much more fierce at shaking old stuff.
Still, you just delete and fix and delete and fix. That's how you make a turbo plasma rifle out of Singer sewing machine.
Ah. And then a week later you hear that Janice in the accounting office in New Zealand depended on one of those features. That's when you learn there is a Janice in accounting. And that your company has an office in New Zealand.
There are times where a legacy technology has limitations that ultimately prevent progress and create maintenance nightmares. A good example is some of the older database technologies where a 1TB database machine could cost 100k (example: Sybase).
You could try to maintain a series of expensive databases, but between replication backups, dev boxes, team of DBAs etc suddenly your costs to keep the old technology are really high. And if these are databases storing expensive financial data, well, maybe the risk of hitting your 1TB limit is an expensive risk to have on your plate.
And it might take 10MM to migrate to a new technology, but in this case it would probably be worth a switch.
When technology is a commodity, and working poorly is still working, then maybe it's hard to justify a switch. But if your old technology starts to limit your performance and introduce risks or problems that detract from your competitive advantage, then you might not have a choice.
It's worth remembering that the Mozilla project ended up being a success in that it dealt the blow that finally dislodged IE from dominance. The original Netscape codebase could not do this because at the time IE4 was already way ahead in terms of CSS support and Netscape was hitting an architectural dead-end. Maybe they could have piled on some more hacks to get NS5 out quicker, but even if it had feature parity to IE5 they had to contend with Microsoft's bundling which was Netscape's real undoing.
Lots of absurd assumptions in here. Many of them were acknowledged, but the author seems to think that this `millions of lines of spaghetti code` with `little use of existing libraries` will be rewritten as millions of lines of spaghetti code with little use of existing libraries. The rewrite should decrease the workload to a fraction of what the author used for his fuzzy math.
If the original system is THAT bad, the language isn't the problem, it's the architecture, and you should probably refactor it into multiple components which could potentially be multiple different languages.
If the only spec is the old code base, then you are probably doomed.
http://www.joelonsoftware.com/articles/fog0000000069.html
The most important part:
"The idea that new code is better than old is patently absurd. Old code has been used. It has been tested. Lots of bugs have been found, and they've been fixed. "
There's value baked into that old code: lessons learned, bugs fixed, workaround put into place. These things can be lost during a rewrite (and sure, there are other times you don't need them in the new version because the original problem has a better solution/doesn't happen in the new language/whatever), and potentially losing/missing should be considered carefully.
> Very little use of existing libraries ("not invented here" syndrome)
The senior dev and PM could sit down over a few weeks and do a assessment of other languages that had good library coverage for lots of the existing system without writing a line of code.
I'd bet the project size would shrink considerably.
At an old startup I worked in, we had a legacy codebase of about 600k lines with 15 years of cruft in an ancient dialect of C++ with lots of not invented here syndrome.
By that point the system was so old and fragile that it simply had to be rewritten. The few libraries that we did use and didn't write were no longer supported, modern OSs wouldn't run the software correctly, vendors had simply gone out of business and so on.
A good 70% of the system functionality was rewritten in C# by just a couple guys part-time over the course of a year basically just gluing together existing libraries.
Rewrites: Firefox, IE, Word, Windows, MacOS, ...
Many of these have been rewritten multiple times and they are all orders of magnitude more complex than some random million-line hunk of perl.
Rewriting is hard and may absolutely or just take a long time, but failure to rewrite is pretty much guaranteed to fail.
I'm horrified by it daily when I have to use scripts written by older bioinformaticians. The best benefits I've heard are string processing speed (<3 you Ruby) and package management (hello Python, Ruby, R).
Please rank several other languages by the same criteria so we can judge your objectivity.
The answer is that your sample is flawed. Scientists can turn anything into a "disgusting mashup of a language". Perl is in its position for a reason.
Yes--people who have practical experience writing and maintaining code that has to be maintained tend to write maintainable code.
Scientists (and, in my experience, especially bioinformaticians) tend to make horrible, awful messes no matter how maintainable you think a language is. (You can hand them Inform 7 and it'll still end up looking like Fortran ate the csh manual and vomited all over an APL keyboard.)
You have made my day.
(Not speaking for myself here, BTW. A decade ago I tried Python & never looked back.)
In your case, part of the problem may not be due so much to the language, as to the authors of those scripts you mention. People who have not studied software development as their primary discipline have typically not been exposed to ideas about good design, writing maintainable code, etc.
In-progress code for my PhD is mostly on Github: http://github.com/Blahah
See also BioRuby, biogems.info, sequenceserver.com
sub tokenize {
while ($_[0] =~ m!
(?<whitespace> [\x20\x09]+ ) |
(?<lf> [\x0a] ) |
(?<cr> [\x0d] ) |
(?<ident> [A-Za-z_]+[A-Za-z0-9_]* ) |
(?<float> [0-9]*\.[0-9]+ ) |
(?<float> [0-9]+\.[0-9]* ) |
(?<int> [0-9]+ ) |
# ...
(?<unknown> . )
!gsx) {
my ($k, $v) = each %+;
# $k: token, $v: data
# pos($_[0]): current offset
# ...
}
}Of course, rewriting something just so you can say it's in a different language is silly anyway. Whoever set that goal is being overly simplistic. They need to step back and re-examine their actual needs and real problems.
But I understand Ovid's frustration with so many people successfully switching from Perl to things like Go and being happy about it in their blogs ;)
I believe we need much more powerful tools that help us in understanding large code bases. Tools that can help us visualize what's going on. Tools that can do testing for us. Tools that can rewrite code for us (think Resharper or other Jetbrains refactoring tools), but an order of magnitude better.
If he's from France, why not just stick to talking about the French working hours? Or just one of the two countries. Bringing up two countries like that was weird to read, but maybe that's just me.