How To Survive a Ground-Up Rewrite
onstartups.com
onstartups.com
On the other hand, if you're saying to yourself "This code really sucks, they wrote their own complex CSV parser, I can rewrite this to use the CSV library in the standard lib" and things like that, no matter how big the task may seem, it'll work out in the end.
Our job as engineers is to prevent those criminal routines from murdering one another by carefully confining them into their prison modules and never letting them get too rowdy.
I am the Code Warden, and within these bytes my word is law!
Another danger i ran across were moving goal posts. Long rewrites will typically be interrupted to extend the old system with new features, which increase the effort to rewrite.
The gross part is (I'm guessing, obviously) that now each write to this data store incurs an additional write and some cpu load on the data store servers (you mentioned it was done via a trigger). But now the value of finishing the migration that you started is just recovering this extra write and cpu.
Here's some ways to do a rewrite well:
Set up the rewrite as an alternate, competing product and let the two compete in the marketplace. This requires having the resources to support developing two such products at the same time using separate teams for each, which is not something many companies can afford. The advantage to doing it this way is that it gives the rewrite time to mature before it's fully replaced the old system. What most people tend to not realize is that mature software with fundamental flaws can often be superior to immature software on a better foundation. Real-world use is orders of magnitude more effective as a QA process than any sort of in-house testing, and always well be, products that have been tried by fire and made the cut should always be respected no matter how horrible their underlying code is.
If you look at a lot of the major switchovers in software you'll see this sort of trend fairly often. For example, consumer Windows took several generations and almost an entire decade to fully switch over to the NT core.
Another good way is to rewrite piece-wise, in place (refactoring). There are lots of different ways to do this regardless of architecture, ultimately you're essentially creating a bubble of new code that grows and grows until it completely replaces all the old code. The advantage of this is that it takes much less effort than a full rewrite and you can do it incrementally and slowly. Often it's easier to mature each new component in isolation than to try to mature an entire new product.
And, of course, there's the hybrid approach. Work component-wise in an "add + deprecate" fashion. Create the new component, have it work in parallel along with the old component, and then deprecate the old component and work toward moving client code to the new component.
This is definitely the ideal situation, but as you said most companies do not have the resources to greenfield a separate project. Additionally convincing the higher-ups that its worth it is a difficult task.
> Another good way is to rewrite piece-wise, in place (refactoring). There are lots of different ways to do this regardless of architecture, ultimately you're essentially creating a bubble of new code that grows and grows until it completely replaces all the old code.
> Work component-wise in an "add + deprecate" fashion.
Everywhere that I have worked this has been the route we've taken. We isolated the old behavior until we were able to fully ween clients off the product (or perform a data conversion and completely use the new code).
Well no, sometimes they have the resources but don't want to spend them, then they end up spending more overall because they botch the replacement.
Competent programmers can maintain other people's code, those who cannot want to re-write everything. Train wreck in progress. Run from these people, or fire them if you're in a position to do so.
But then again, sometimes, the problem is just that the legacy code is shit. In a project last year, for instance, I had to rewrite every class I touched; the whole thing was so fragile that it was impossible to change anything without the rest breaking. So, I did this “incrementally” rather than all at once, and factored this into the cost of the features I was implementing for the client.
Legacy systems have gone through so many fixes and shifts that it's almost impossible to recreate them without deep understanding of the business demands it meets.
The business part, I would say, is critical to being able to maintain (or re-write) a project.
I'm talking 40, 2000 LoC files that all called the database whenever they felt like it, all echoed out data and modifed each other in wierd ways.
I spent about two weeks trying to untangle the mess and threw the towerl in at day 15. We just rewrote it using a PHP framework that not only let us query the database using a sane ORM, but also gave the project STRUCTURE and added mantainability value.
Sometimes a rewrite is the best cure. ;)
Probably why most big companies have such a horrible time getting good people to work for them: who really wants to work in a place like that if you have options somewhere else?
Immediately brought to mind the first large shop I worked in out of school (about 1992) where they had a big chart on one wall titled "Punch Card Elimination" with a downward trending line. I was actually shocked that a large, multinational manufacturer was using punch cards for anything, particularly when my project involved building cutting-edge (at the time) client/server systems that included wireless handheld devices on the client side. My introduction to the counterintuitive world of enterprise IT: you will often see opposite ends of the technology spectrum in use, simply because the value proposition of eliminating the old was never high enough before.
What about when what you're rewriting is basically an intensely-overgrown prototype that was never meant to run in production in the first place?
You're supposed to completely rewrite those, aren't you?
I am wondering what sorts of info I need to go looking for to get this finished. I have thought about it before, contemplated asking around, and realized I didn't even really know what questions I needed to ask.
I also found this, which struck me as helpful but not quite what I wanted to read about: https://www.linkedin.com/today/post/article/20121226171356-6...
Most of the time the Wordpress template is so wonky that its better to just rewrite it (especially if you're migrating away from Wordpress in the first place).
In the past, we used subsets of data to test our migration tool. Always. ALWAYS, there was something bizarre in the the excluded subset that caused the migration to break.
When you have a PHP project where there's a mix of procedural function (the code is at the top and the html at the bottom, so it's easy to split right ? Hum and those includes, what do they do exactly ?), ... classes (We used classes to put our functions in so we do OOP, right ? Err... no) and parts of a heavily customized Wordpress. Plus two-three half started in-code reorg/rewrite that went nowhere because the devs jumped the ship.
You're told you have to I18N the whole spaghetti meatball soup. You try a few weeks and then see you're going nowhere and start to realize that you're doomed to a rewrite. Oh, and you're a team of two. Sometimes you really have to. And migrating data .... such a pain in the ass ! Test encoding I'm looking at you ! At of hand int code (would not call that enum) I'm looking at you too.
But. It's not always the case. And right now I'm starting the reorganize and refactoring a project that was already in a descent shape. And it's much more interesting, less straining and business owner have more visibility on the work done and are happier.
So yes, whenever you can, incremental rewrite is the way to go. But never say never. There are contexts where doing a rewrite is just plain necessary because the code is in such a bad shape you just can't add any feature anymore.
While successfully in production, the web servers were crashing multiple times a day (still Apache-modphp). So, I started to work on the PHP upgrade from 4-5 to get onto php-fpm + Nginx to stabilize the site. Lo and behold, the legacy db abstraction code (poor man's ORM) was not playing well with later PHP releases so I started down the path to patch what I could.
I successfully upgraded PHP, fixed a few of their outstanding issues, but I was weeks into this project and had not completed any of the new feature sets on my plate yet. Looking into the horizon I imagined months of wrestling with legacy code and the deeper I went (there was a SOAP layer I hadn't even looked at yet, they wanted REST/JSON) the more worried I became.
We evaluated a bunch of options and opted for a re-write. We're still deep into it but there is light at the end of the tunnel and I'm pretty sure it will be worth it in the end.
The analogy used is of a ship at sea. Replacing the ship in a single step, while it is underway, is impossible. Replacing too many planks at once is foolhardy. (I thought the analogy was Oakeshott's but I can't find the original source).
So instead you replace the planks one at a time. Yes, this is tremendously slow, wasteful and frustrating. But when the ship is a system on which you utterly rely, there is no other safe path.
Hayek expanded on the problems of pure rationalism more generally; it's a trap we as a profession tend to fall into. We substitute our preferences (for a tidy, orderly system) for an appraisal of actual reality. Real systems have accumulated complexity that cannot be wished away. Worse: the complexity is often essential and not merely accidental.
What the author describes sounds like a complete nightmare.
On the projects I've worked on, I would like to think that a rewrite will not be needed.
If it does, it will be because we've now understood the problem space so well that we can start from scratch and avoid all the 'pivots' that happened in the original code base.
If so, do you rewrite those tests? Or do you keep them, no matter how ugly they are, because new code still has to pass all the old tests?
If you're rewriting a project then you're probably going to have to bin the unit tests. If you're rewriting it's because you want to make some drastic structural changes and unit tests are just too tightly coupled. Integation tests on the other hand can be an absolute godsend, you can swap out the entire stack and they'll still provide just as much value as before.
Get your data source in order, then refactor in small achievable projects from there. Better for motivation and better for end users.
edit: unflagged after modifications
I think the content is really good, so title changed -- don't think the author will mind (he had two alternate titles for the original post).
Thanks for taking the time to update it.
Ground-up rewrites are occasionally necessary, always difficult, and often hideously political. There's no binary right answer on this question. Sometimes it's the only answer, but that's a terrible place to be.
My shortest job tenure ever (108 days in the winter of 2011-12) was in a company where I was brought in to fix the cultural and social collapse that had followed a badly-thought-out "rearchitecture" led by a 25-year-old CTO's protege who (a) was on his first real job, having burned out of college and landed in retail, and (b) didn't know what the fuck he was doing. I'm still deciding whether I should share that story, in full, with the world.
The reason I'm hesitant is that good engineers (who don't deserve it) might suffer more than the management (which does). I learned the hard way that whistleblowing often hurts the more vulnerable good guys first.
It's a startup of about 100 people. I could demolish this company's reputation with a few keystrokes, but the executives (who deserve it) would probably be just fine. It's the engineers who'd suffer-- having a disgraced company on their resumes-- and I hate that that's the case, but such is life.
What happened is that there was a "War of Three VP/Eng's". There was an unofficial VP/Eng (call him Dan) who did a really good job. However, with four product pivots in three years (the CEO was an idiot, but kept being able to raise capital because of family connections) there were legacy and code-quality issues. Top management threw Dan (even though he was a true 2.0+ engineer) under the bus, and the CTO pulled in a 25-year-old protege (call him Tony) who was fucking terrible. He knew basic programming, but this was his first white-collar job; story was that he'd burned out of college and was in retail before the CTO "rediscovered" him. Tony's quite intelligent and he can talk with the best of them, but he's technically mediocre and a genuine psychopath.
Dan had engineer support; Tony had management support but engineers disliked him because he was an obvious case of an incompetent protege. Tony's "rearchitecture" fucked the company badly. I was brought in as a 3rd VP/Eng (and actually promised the title, at the 6-month mark) to resolve it.
There were good engineers on both the old and new teams, and resolving the technical disagreements was easy. But in March 2012 it was clear that my real reason for being hired into the company was to give management a credible excuse to toss away half of the old team. I either had to sign a bunch of documentation that'd be used to justify Pincus-type moves against 11 of my colleagues, or give up my contention for the VP/Eng role since (apparently) that's just want managers have to do. I chose the latter-- really didn't want to be managerial in such a toxic company-- but Tony convinced management that if I was throwing in the towel for VP/Eng, I should be shown out entirely.
This wasn't only out of ethical altruism. I did the right thing for good reasons, but I also had no choice. If I signed the papers and gotten a bunch of undeserving engineers fired (or at least put through humiliating PIPs that would probably involve equity clawbacks) then I'd lose credibility and Tony would win. Tony (angling for CTO once the existing one burned out, which he did around the time I left) teamed up with management to set it up so that I would be the one getting dirty in the Pincus move. He could be the good guy CTO and I'd be the bad guy VP/Eng, the rubber glove used for evil work and thrown away. It'd be career suicide because no engineer would work with me after that.
It wasn't the rewrite that killed that company. The rearchitecture was a symptom of psychopathic management and a really horrible CTO-protege-cum-unofficial-VP-Eng. Still, it makes me extremely skeptical every time I interview with a company and hear, "We're throwing out all the old code" because I know how politically fucked-up that often is.