The True Cost of Rewrites
8thlight.com
8thlight.com
So i just want to tell you that i've done successful rewrites!
The biggest was rewriting the e-commerce site of a well-known toy company. They had a homegrown site written in .NET, and we rewrote it on top of a commercial e-commerce framework in Java. It took a couple of years, the rewrite team was 2-5 times the size of the team that built the old site, and the client ended up with all the features of the old site, plus internationalisation, plus thorough automated tests and deployment. They seemed to be really happy with it.
The smallest was a data-aggregation server my team uses. It was written in Node, and kept crashing under load. I rewrote it in Java. It was just me, and it took a week or two. It hasn't crashed since, and we've been able to add loads more features.
Success factors in both: putting effort into thoroughly understanding the old thing; resisting any temptation to add new features unless you are exceptionally confident in how much more work it will take; thoroughly testing what you build as you go.
So what makes this rewrite successfull? I am sure in that time you could have easily added the I18N, tests and deployment and refactored the most difficult parts of the code so that it would be more expandable and easier to maintain.
That it was successfully completed and the tech provides a better basis for future features? Sometimes timeframe and team size are almost irrelevant to the success of a rewrite project.
Business needs still matter. Time and money is still a thing.
Are you arguing that success is in the eye of the stakeholder? Sounded to me like it was a success.
> Are you arguing that success is in the eye of the stakeholder? Sounded to me like it was a success.
I think a lot of these stories make sense if you replace software re-write with a home construction project gone overbudget, one where the overage comes out of your pocket.
I'd say that something that gets done eventually at a great cost is not successful if the case made for doing the project was an overly rosy story of how easy/inexpensive it would be to do. It really hurts when you pay for it, and that is the perspective to take.
As for success being in the eye of the stakeholder, sort of -- in my analogy, the construction company might consider it a success all the way to the bank. That project owner may consider it a success to save face. The project sponsor would likely not.
If it was a quick cash-grab then obviously a long and expensive rewrite is deemed unsuccessful. If you go for the long haul and want to bring real value to your customers and the rewrite helped you do that, then it was a success.
Nothing is ever cheap in IT in my opinion. I found that being very upfront with the people with the money landed me more projects than trying to lie about it and reveal the extra expenses one by one over a course of a year.
"No, that's gonna be very far from 50k euros. Be prepared to pay 300k and it's not gonna be a year. Optimisically, two, realistically, two and half to three."
It works surprisingly well with experienced businessmen. They applaud the honesty and we move directly to the next point -- what are the tradeoffs and compromises.
Sounds like success.
There is a reason why stakeholder would like the change despite large cost - if the old system was unstable and source of stress and fires for the stakeholders, they will go for more expensive option.
Kind of like after you had a car that needed to be fixed every few weeks, you will go for more expensive higher quality one next time.
Thanks to Charles Pick (phpnode on here) and pretty tight plan we got the job done in record time. Yes, we could have done even better. But we did not have the budget and the time so we worked within those constraints and got the job done. Would have been super to work with more people, more budget and several months to do it. But as it was it was already a huge improvement over what was there before.
We've instead focused replacing parts of the program over time. If we touch some old code that works but needs to be extended for a new feature, we usually rewrite it and bring it up to date as part of implementing that feature. Other times it might be too sensitive and we just modify a small piece of the old code and let it chug along like it has done for over 15 years.
Instead of spending time rewriting the whole shebang we can deliver features, alternatively not having a second team reimplementing everything avoids a huge financial burden.
With our approach we will likely have some code that's 10-15 years old now and which still be untouched for another 5-10 years.
One of our struggling competitors has, as far as I know, a decent replacement in the form of a rewritten product. But trying to recover the cost they're trying to pitch it as an upgrade. So a license to the new version is much more expensive.
Apparently a lot of clients either do not see the added value over their current version, or, given the higher price, they're considering alternatives such as us.
Maybe I've been lucky but they were all raging successes. At the end everyone is relieved and happy and much back slapping ensues.
I like Spolsky's articles but I hate the dogmatic culture a writer like him creates. His article on rewrites is always a thing whenever I've been part of a discussion on rewrites. It's an anecdote, nothing more.
But without creating that dogmatic culture, would his writing have become popular enough to be disseminated to a wide audience?
Perhaps we should think of it as a corollary to the Gell-Mann amnesia effect.[1]
Personally, I think you have the cause and effect switched around. Would his opinions be considered dogmatic if he weren't so popular?
If you have to maintain untested code I recommend you get yourself a copy of "working effectively with legacy code" as it is mostly a list of recipes on how to add tests to a codebase. I also recommend it to anyone starting to write something new so they can learn what is useful to test.
Forget about unit testing and start working on end-to-end tests. Selenium, Sikuli, Codeception, Wiremock, siege are the kind of tools you want more than whateverUnit. Test your applications at the UI level. Test its performances. Your client does not care about your design pattern usage: they want something which respect their specifications.
I would call that refactoring, whereas a rewrite means starting from scratch. I agree that refactoring is the superior choice if at all possible.
End to end tests are great for providing confidence that the code does what it says on the box. Though end to end tests are slower than unit tests, and it can be tricky to track down why an end to end test failed.
A "Test Pyramid" seems a good idea to me. (Unit tests can be quick, but don't cover much of the system. E2E tests cover a lot of the system, but aren't quick. "Test Pyramid" suggests it's better to have more unit tests relative to E2E tests). "Only unit tests" or "only end to end tests" don't seem like practical things to aim for.
Sometimes, but it's not always that simple. Useful software tends to interact with external systems, which might not be amenable to that sort of automated end-to-end testing. Also, the objectives of a rewrite might include enabling new integrations and/or user interfaces, which deliberately don't work as drop-in replacements for what was there before so wouldn't expect to pass the same end-to-end test suite. Automated testing is useful in the right context, but IME it's rarely the whole story and there are often data migration exercises and new integration tests to be done as well.
I don't think he said it's simple. As someone who tried to add unit tests to legacy code, I can assure you it is anything but simple. However, it is a sound approach.
>Useful software tends to interact with external systems, which might not be amenable to that sort of automated end-to-end testing.
That's what mocks are for. Writing effective mocks is an art, though. Again - not trivial.
I shall respectfully disagree with you here. I mean, yes, obviously that is literally what mocks are for, but I have never found mocks to be a particularly effective or valuable tool for testing. They can take a disproportionate amount of time to write and maintain if whatever real external system they stand in for is complicated. They are inherently fragile if that system is subject to change. Most importantly, even if those tests pass, you don't actually know whether your real system is going to work, and IMHO the greatest benefits of automated testing are found where you can systematically and repeatably exercise exactly the behaviour and interactions you might see in production.
But when you want to re-write the last thing you want is to fossilize your design. You explicitly don't want the same system you started with otherwise it wouldn't be a rewrite, It would be a refactoring.
So, the goal isn’t to keep the monolith as is. Instead, it’s to break up its functionality into small parts. Then, the parts may be joined so that the old interfaces work as they did before. But, they may also be used to build entirely new interfaces.
I usually rename the test class of the old implementation to OldXTest and try to keep the interface of the new code similar enough to enable a reasonably quick transition into the new code with proper unit testing.
I’ve done this several times with success. There was a period of about 10 years where I specialized in rewriting legacy code. The first case, I rewrote 200k lines of C++ (COM/OLE) code in Java. Then, I joined a startup that had to scrap its entire codebase because of scaling issues. Then, I worked for two other companies that had acquired other software companies (with crappy code), and I led the rewrite and integration efforts. These are the notable ones.
Believe me, rewrites can be successful. They must be done carefully. But, there are lots of ways to manage the process and mitigate risks. I’ve done it at least 10 times, and I’ve never had a failure or substantial cost or schedule overrun.
My experience has seen it work well as either 1) bottom-up: pick a small enough section of the codebase such that you can fully understand it, then rewrite just that piece, repeat until complete to get an exact copy of your existing application or 2) top-down: you write a new application from scratch completely, starting with understanding what your business's goals are and how to best serve your clients
Isn't this an absolute requirement for a rewrite anyway? (Your "2" case?)
A good rewrite happens so subtly that end-users and operators will never really realize that a rewrite is underway.
There are only very rare cases where a rewrite in big-bang fashion is indicated and even then the bulk of those will be incompetence on the part of the tech crew because they see no way to turn the job into an incremental one (or do not want to see a way).
Calling incremental improvement a form of rewrite gives it a bad name :)
I’d argue that rewrites that are a second system that (hopefully) eventually get deployed once the features exceed current system is not incremental at all, but is instead Big Bang.
Even if the features of second system are built incrementally, if the business aren’t using it until it is finished then it is Big Bang.
If the plan is to change all the parts one by one until done, it's a rewrite. If the parts are changed one by one for independent reasons, it's just code evolution.
When you have products and systems that have been around for a decade or more, major portions of them become outdated and need to be rewritten (or largely rewritten, the two are synonymous to me). Best practices in crypto change, operating systems and hardware improve/change, etc., and if you want to keep generating value, you had better change along with it. You can see where organizations don't do this: the developers force weird constraints on IT like needing to use old, obsolete versions of operating systems because the software won't run on newer versions.
Along these lines, one of the problems that I think exists with the software industry today is an inability among developers to recognize that software sticks around a lot longer than one might originally envision. And this phenomenon only gets worse (better, for the end user) as the value provided by the software increases. It's a bit of a Faustian bargain: everyone wants their software to be used and provide value, but often don't realize the "soft commitments" being made in the background that can tie you (or the business) to the code for years (or decades).
In both cases:
- it was promised to take 6-9 months, and ended up taking 3-4 years
- it never really finished, the old software had to remain in production (along with the new "rewrite")
- in some form or the other, while the rewrite was happening, the company lost its customer focus and/or its ability to innovate
- good people left
Instead of rewrites, make incremental changes. The eng. team should never be off doing its own thing.
Also, a good lesson from Facebook: instead of rewriting their PHP codebase, they extended the language by creating the "PHP++" language Hack (along with the HHVM runtime), and incrementally changed their codebase to take advantage of Hack:
The only point I'd like to add is that incremental rewrites often have long-lasting value in terms of a culture that emphasizes continuously paying down tech debt.
I agree with most of your comment, but I don’t think this part generalizes. This isn’t feasible for most companies unless you’re operating at Facebook scale.
Afaik, the team that did Hack/HHVM at Facebook is ~5 people. I don't think you need scale for this (the rewrite of the code itself is not a scale thing, the codebase is usually linear-ish in the number of engineers).
My point is: instead of doing a rewrite, be inventive and avoid it, and this is a great example.
Also, not every company could hire the 5 people needed to do their own version of a hack/hhvm project. But at some point it makes sense to find those people instead of rewriting in an unrelated language
Facebook has the scale to invest in creating tooling, IDE extensions, core libraries, documentation, training, test frameworks, bindings, etc for a new language that they create. They also have the organizational scale and career development to make it worthwhile for an engineer to learn their proprietary language.
This doesn’t apply to most other companies.
[1] Well, if the codebase has very comprehensive tests, that's likely better than the old code. But few projects have this.
I have even seen them done to bypass a management/development team that was in disfavor with the company leadership. In effect, put together a new platform with new leadership.
I guess “rewrite” is bad, but “reinvent” with experience from the former system shouldn’t be shied away from. Incremental improvements might only take you so far before competition overtakes you.
On the other hand couple of years back I started purposefully taking up projects in bad shape and fixing them by cutting discussions about rewrites and instead focusing on putting work where it matters -- diagnosing the problems and finding solutions to them.
I found that most of the teams I met stopped or indeed never really had any practice of constantly diagnosing and improving their process and application. This is usually the cause why the application is in bad shape but more frequently than not the team will blame the organization, predecessors or constraints like old technology they are working with. They channel their frustration by focusing on the idea of the rewrite which seems like a relief from the frustration of current codebase. Unfortunately, the rewrite is performed using the same process that failed the previous version with predictable results.
Think of a person that has messy house. This person does not have practice to keep things clean and in order. The person decides he/she will fix the issue by building another house and burning the old one with all belongings. The result is predictable.
The solution is to learn to keep things clean and in order instead of burning the old house and re-building it at great cost and effort.
Can you elaborate a bit more? Using your analogy I imagine this would be someone tidying up every evening before they go to bed. But can you maybe describe what it would look like/how the process might work for a team maintaining an application?
We have plenty of ideas on how to build new software: agile, TDD, DDD. It seems that there is no proven process for doing rewrites.
"Think of a person that has messy house. This person does not have practice to keep things clean and in order. The person decides he/she will fix the issue by building another house and burning the old one with all belongings. The result is predictable."
I believe your unstated epilogue is supposed to be something like "Rather than build a new house, it is best to learn how to take care of the one that you already have." Great. So let's suppose that happens. Now the team has a new skill. They've learned how to be organize their code and fix the problems in the code. So should they incrementally improve the code, or should they do a complete re-write? Your analogy does not stretch that far into the future.
I agree with this, and I've had the same experience:
"I found that most of the teams I met stopped or indeed never really had any practice of constantly diagnosing and improving their process and application. This is usually the cause why the application is in bad shape but more frequently than not the team will blame the organization, predecessors or constraints like old technology they are working with."
My sense is that what these teams need is new leadership. It doesn't really matter if their code is written in Fortran or C or Javascript or Go, what they need is good quality new leadership. Once they get that, their situation will improve. However, the new leadership needs to have the freedom to take the team in the direction of those skills and experiences that the new leadership has gathered over their lifetime. If the new leadership has a lot of experience in Ruby On Rails, then a complete re-write to Ruby On Rails can be justified, because if the new leadership is good, then they will produce something good in Ruby On Rails, and it will be better than what the old organization had before.
And indeed, I think many real life re-writes happen because of exactly this set of circumstances.
An interesting point: code is (probably) not as ugly as it seems -- it just looks that way because you didn't write it. What looks like cruft is actually the accumulation of years of bug fixes and edge case handling.
Rewrites might be the wrong term. "Refactoring" is a better. If you already have functioning code you don't need to rewrite it, but fix the mess.
If you have undocumented legacy code that's a mess, and it doesn't work with new quite different requirements, it might be faster to just rewrite most parts and cherry pick code that seems to do edge catching stuff and opaque interfacing with other black box systems.
I aim at spending half my time refactoring and documenting, so I'll in reality end up with at-least some time. Random Company average I have been at is probably around 0.1%.
Both have their place but refactoring happens regularly whereas rewriting is a more drastic option that should be used only when absolutely necessary. Having a robust test suite is very important with rewrites to prevent regressions.
Undoubtfuly it is all mature and correct advise. However, I feel it is all one-sided and I ought to provide some counter points:
1. The codebase in question may be well beyond the line where any sane person would touch it, seriously.
2. Individuals experienced with the codebase, its structure, implementation, technology stack might be not available (think cobol).
3. Refactoring or incremental rewrite is a process that has to be planned and managed s.t. current product/codebase structure. Oftencase it is this very structure, which demands the full rewrite -- because of it being too convoluted to allow refactoring to take place.
4. You might want to do a rewrite to refresh the tech. stack -- language, design, frameworks -- it is ok to do so.
5. You do not necessarily need to come with the same feature set. Both you and your customers might want to cut down unnecessary cruft. Technical debt usually starts to show its signs with losing flexibility w.r.t user request to change features.
6. Occasionaly you might come up with a separate, different product, with a different name and brand -- which might turn out to be even better!
7. You learn a lot in the process.
Being part of team incremental rewrite, myself, I'd say that you're overestimating the cost of the incremental rewrite and underestimating the cost of the rewrite.
For your first 2 points especially, I'd argue if you can't even begin to incrementally rewrite a system, you're in no position to begin to plan a full rewrite.
I'd say that if code is being used in production it's not yet "dead", and it should be a much higher priority for team members to work with it.
Can you not add additional tests? Refactor out even a single feature at a time into more modern tech?
> 2. Individuals experienced with the codebase, its structure, implementation, technology stack might be not available (think cobol).
There are people that reverse engineer binaries, find a place they can use a buffer overflow to insert a jump and then “program” by jumping around to different parts of the existing compiled code.
Your legacy cobol spaghetti code isn’t impossible to understand—-get better or better motivated engineers.
> 3. Refactoring or incremental rewrite is a process that has to be planned and managed s.t. current product/codebase structure. Oftencase it is this very structure, which demands the full rewrite -- because of it being too convoluted to allow refactoring to take place.
It does take extensive, often tedious effort, but there’s no such thing as too convoluted. A giant ball of tangled string can be pulled apart one knot at a time.
> 4. You might want to do a rewrite to refresh the tech. stack -- language, design, frameworks -- it is ok to do so.
> 7. You learn a lot in the process.
It probably is very true that a rewrite is often better for the resumes of employees. But employment is a fiduciary relationship. By cashing that paycheck you’ve agreed to put the company’s interests before your own. And rewrites kill companies.
When someone new comes in and within six months of starting starts pushing for a rewrite he should be fired and a hard look should be taken at what went wrong with the hiring process.
Sure. If you ever find this [1] a rewrite does look like the only sane decision.
[1] https://ayende.com/blog/4612/it-really-happened-legacy-progr...
The simplest way I can show the error? If making a new piece of software is so expensive that it's far, far too expensive to do, how come folks enter the market everyday with new software which keeps replacing all of that stuff you think is irreplaceable? And they not only replace your stuff with better stuff, they do it at a fraction of the cost you take just keeping the lights on in your shop.
Math is not going to save you now. Looking at how fast features can be deployed is sticking your head in your ass. It's counting things for the purpose of counting them. It doesn't work like that. You may enjoy counting, sgraphing, and tracking points-per-sprint or feature-speed-per-team, but nobody buys or uses software based on feature count or team speed. It's not that these things aren't important. It's that you're confusing managing things with value creation.
I wrote an essay earlier this week about code budgets which I think can help a little by at least helping force conversations around the real issues involved. http://tiny-giant-books.com/1.html?EntryId=rec39SaDeZCZjauRo
But the larger issue here is that organizations don't know how to create and harness value. They're really good at hiring, managing, and a few other things. But those other things don't come directly into play here. It'd be great if they did, but they don't.
You don't know how to create value. Step 1 is admitting that. Without that admission, no amount of charting or graphing is going to help. And yes, you can't rebuild your software. Probably lucky you've got it up enough to provide value right now. I'd hang on to it.
Incremental changes by domain, decomposition, and careful management of the work is not only effective, but also necessary.
After scoping out the work, my recommendation? Build a small app to handle receiving. You do receiving at all of your locations, it's being done by several different separate systems, and it's an opportunity to write a small, cross-platform app that can be used by anybody with zero training right away.
It was shot down! Why? Because large projects aren't done that way around here
That has nothing to do with anything, yet it prevented getting started immediately.
As a hired-gun, I moved on to bigger and better things. The org dropped 100M+ on just the kind of rewrite this author is talking about before giving up in failure. (Actually they changed the goalposts so that they won, then had a big party. But there was very little done compared to the money they spent)
It's the wrong mental model. It's painful to watch, like a kid with a big hammer trying to make a large circular block fit inside a small square hole. It's not going to be good even if somehow you make it happen. It's going to be ugly as crap. You end up destroying the thing you're trying to help.
That is evidence that the author has a point. He also said, at the very end, that this is a two-part article, and part 2 will deal with how to address these problems. I would not be surprised if he advocates something like the approach you suggested in this case.
Yes. They lost me at that point. Why do people feel inventing some cargo cult mathematical formula enhances their point?
Because they don't (yet) have the calcified, dysfunctional processes that led to the need for a rewrite in the first place.
On the other, you've got small, scrappy teams whose survival depends on quality, because they only ever have a month or two of payroll in the bank.
Or put another way, the place where you can save a ton of effort in a rewrite and increase your chances of success is in throwing stuff away, but you usually need more than just the engineering team to achieve that.
The workloads were categorized by type and that script decided on the order by base-ing a timestamp and using the resulting integer as an if for each category. This mostly worked,too. Sometimes, if previous workloads took too long, the following category would not be triggered for several cycles... But most of the time, everything was done in a somewhat timeley manner.
That is a rewrite that needed to be done, and yet would never be doable with your strategy.
There is no new business feature to add, but it still needs to be fixed because as each workload increased, the implemented feature got less stable.
For example, rewriting a crufty ordering system that drops 50% of the edge-cases that have accumulated over 30 years, but making it work on mobile devices.
If you could estimate “keeping feature X costs $200,000 per year in added development time” it would be an easy decision. Through this process of cost estimating you may find that while annoying it only costs $10,000. Any ill will generated from removing it would cost more than that so it should be left alone.
So the full discussion goes something like this: "Can we remove support for Roman Numerals? We are pretty sure nobody uses these anymore, and we estimate that this would cost 20 man/days in this release but result in a saving of 800 man days over the fiscal year..."
"Not sure, it could be useful again, use these 20 days to add support for Aztec calendar to the Insurance reports - you have already postponed this twice!"
(I.e. they get the impression that somehow there are 20 "extra" days available and those have to be diverted to implements something that may have become already obsolete).
The conversation shouldn’t be about it costing 20 days to remove something. It should be about the net savings from doing so.
“Removing rarely used feature XYZ will open up 780 man-days of schedule time this year for new features by improving our efficiency”
As such, every different part of the business may request changes/enhancements - our team provides these to lots of different "business departments" across the whole company.
Again, imagine a hotel chain with hotels all over the world, each national branch may require specific changes due to local laws and regulations, or because they need to start a new incentive campaign or participate in a joint venture with a flight company or whatever.
There is one application, and N different (competing) "customers" each one considering only their own specific plans and priorities. (In case of conflicts, the pecking order will be used to solve who gets more attention: biggest hotels, or hotels in regions that bring more revenue have more "clout").
Now, when I say "we have to postpone your request for X in order to recoup 780 more days later" the guy in front of me will immediately conclude that he will not necessarily get a bigger share of these 780 days - he will have to fight for his piece just as strenuously as before, and in any case this will happen maybe in three months, and he needs his stuff yesterday - so he better insist to have his own specific request included in the release, no matter if it costs 21 days, 20 days, 5 days: he wants this to be done because the rest of his business needs it for a specific date, and everything else is just a way for IT to postpone his request once again.
In other words: everybody wants their own specific request implemented as soon as possible, and anything else has absolutely zero interest for them. Especially if it is some promise of future "gains" from guys who are constantly late.
In my experience, this is not so uncommon when you work on an app that has been developed internally (and so there is no unified marketing department which represents a single stakeholder) - note also that precisely because we have to work for a myriad different "internal customers" we tend to accumulate "technical debt" at a faster rate: there is only one codebase, and has to accommodate all these pesky requirements from all over the world...
“If we could open up 780 days on the schedule - what new features would you like?” At this point you want the client dreaming of all the extras they never thought they’d get to have.
Then at the end you mention the 20 day delay. At this point they will feel the “loss” of the new features they just imagined. It’s an extremely powerful sales technique that works almost everywhere.
Given your earlier comments about the extra 20 days they seem particularly susceptible to this technique.
I agree on the principles nonetheless. You want clients to write a formal request for new features, then development has a backlog of requests and can prioritize them.
Clients might not like the prioritization but that's life, limited bandwidth, it's simple to show there is too much to do and not enough resources.
Everyone else has to justify their programs with a cost/benefit analysis. I don’t think software should be immune.
Applications developed internally don't necessarily have a roadmap (or might have it and abandon it 2 weeks into the new year because). Problem is, nobody will consider you a hero for sticking to the (now obsolete) plan.
I suppose the way to deal with it is a feature switch, turn it off and if nobody notices after a year then remove the code. That or add some instrumentation to identify what is being really being used and what isn't.
Anecdotally, I’ve had precisely the opposite experience at times. Though I can tend toward what you described as a fault, I’ve also experienced abject need to rewrite internally and externally produced software to fulfill business requirements.
Some businesses decide that they want a new thing and find some developers to work on it. Neither seem to care about existing systems, or integrations. It will be one more legacy soon, if it ever gets anywhere.
If the old software is convoluted, it may be because the old business processes are convoluted.
A software rewrite should be done in the context of a business process redesign. Consider how information flows through the organization. Do you really need to do all the things you're doing now? Should different people or departments get different information? Should the product you're selling or the service you're providing change? Should parts of the process be outsourced or insourced?
If an IT department is considering a big rewrite, but doesn't have the authority to look at the business as a whole, then the project is being managed at too low a level. Rewrites make the most sense in the context of a change in the underlying business.
Guess what? It turns out that all those edge-cases were there for a reason. You needed them. And by the time you've added them back, or equivalents, two things have happened. First, you've burned an order of magnitude more time than you budgeted for the rewrite. And second, you have a codebase every bit as gnarly as the one you started with.
First rule of rewrites: don't do it. Second rule of rewrites (for experts only): don't do it yet.
Maybe you used to have much less RAM in the past, but also much less data, or not as many users etc.
I really hate this cute quote about rewriting because it's elitist (you have to be an expert to do it right) and generally sells the false idea that rewrites aren't a valid part of iterative product building.
Junior as in * have been exposed to few real-world development projects, or, * have little understanding for the interplay between development and business. (The junior developer may however be very good at programming in languages X, Y and Z.)
Many developers will be junior under this definition for all of their career.
The exception I have seen from this is consultant sales ppl, who want to use the hype of the new often untested (or for the task non-optimal) technology or paradigm X.
When the original system was written by developers who were not junior, as is often the case with large successes systems that have survived, the result from letting junior developers rewrite the system, will be predictable.
Tip: Be cautious when listen to developers that talk about “technical debt” and similar. Make sure they are not junior developers under the definition above or are just not that good at reading code and understanding real world systems. Similarly, when listening to a sales pitch by a consultancy firm, make sure they have your interest in mind.
So rewriting and knowing the full scope is a luxury that allows a much better design.
The problem is that in the new system there will be new cases that don’t fit the new design.
Rewriting shouldn’t be a decision taken lightly but it also shouldn’t be avoided at all cost.
The alternative isn’t massaging your old system when the technical debt weighs it down too much. The alternative is just maintaining it without improving it, which eventually sees the system overtaken by competitors.
Having the same people who wrote the code work on the rewrite is silly though. How could they be expected to produce something better?
By learning from their mistakes?
“I divide my officers into four classes as follows: The clever, the industrious, the lazy, and the stupid. Each officer always possesses two of these qualities.
Those who are clever and industrious I appoint to the General Staff.
Use can under certain circumstances be made of those who are stupid and lazy.
The man who is clever and lazy qualifies for the highest leadership posts. He has the requisite nerves and the mental clarity for difficult decisions.
But whoever is stupid and industrious must be got rid of, for he is too dangerous.”
I would say the 4th class of people are the ones likely to churn out a horrifying codebase. Such people are also unlikely to be able to recognize and learn from their mistakes.
These are the people that do “negative work”. They cause other people to do more work than if they weren’t there.
Only if you have really, really poor developers on your team.
Every decent developer learns the pros and cons with how they implemented something. They'll bring that with them to the next time they need to implement something similar. And again, and again.
He ran away everyone who was brought in to bring some outside perspectives. The only people left are those who gave up and go with the flow because they can’t find a better job without moving (small city with only three major technology employees) and those who don’t know any better.
One system I worked on was an information system in VB6 which was being used as a base platform to build domain-specific ERP implementations. Deployment in their environment to anything non-browser-based was a nightmare. There was overlap between domains although there was shared data (customer and supplier information, product database, that kind of thing). The obvious thing to do here was to let each domain (i.e. team) build their own solutions in the most effective platform for their domain whilst integrating via defined interfaces. That's not what happened; when I left few teams had anything working and those who did spent the majority of their time on deployment issues.
Another system I worked on was initially built by a very junior team under tight time and budget pressure under leader who believed in never saying "no" to the customer. They created a horrible mess (most of the code was in a couple of files). The system was deployed by two/three customers. Testing was done during deployment; the first version worked well enough to go into production. A year later everything had ground to a halt; there was no formalised testing and the developers couldn't fix one bug without introducing several more. One very smart developer started to "rewrite from the inside" by implementing new features using clean and modern techniques. At that point I took over the team. We doubled down on this "rewrite" whilst adding manual tests (and later automated tests) to cover the most important flows. We only rewrote features when they needed to change and in some cases removed functionality when particularly hairy features no longer worked.
Otherwise I advise against rewrites; if the system was reasonably built and the platform viable then rewriting is just a waste of money.
I would like to know if anyone participated in a similar gradual rewrite and how it went.
It may be very hard for some architectures and also sometimes wrong architecture (prototype became production, some features were dropped along the way, a new feature is dog slow) may require big rewrite, even breaking APIs, then it would be impossible. But if one keeps API one can at least do comparison tests as more features are implemented.
[0] http://roscidus.com/blog/blog/2014/06/06/python-to-ocaml-ret...
> As people move off of the team, features are forgotten and misunderstood; but as long as those features continue to work, there will be customers continuing to depend on them.
You have this in the non-rewrite scenario, too, it's just part of the technical debt, so it's less obvious to see. And it will bite you just as hard if you need to update the feature.
So because rewrites are good at uncovering unknown unknowns, and it might be better to say that the technical debt was underestimated.
Customer wants to migrate from Oracle (because it costs quite a fortune) to free PostgreSQL and he wants to use browser. Also a lot of things changed in those 20 years and many functions are just not used anymore. I think that rewrite is quite justified in this case (will rewrite in Java).
The "this is dumb, we should rewrite it" reaction people tend to have when exposed to very large, complex legacy software systems is totally understandable, but really comes from a fundamental misunderstanding of the forces that drive software development at organizations which have legacy code bases. Yes, the re-write, if it gets "done", is basically guaranteed to be "simpler and cleaner" when it arrives; of course it is! It is being compared to a code base that's 20 years old, not just from an era where the technological compromises were completely different, but also from essentially a different world of what was acceptable at the time. In addition to that, all software that lives long enough tends to become a mess. Let's see how the rewrite is in 20 years.
Oh, but wait, this time will be different. This time we won't make quick fixes and little hacks, we'll be diligent about requirements, we'll refactor and clean up when change is needed ...
Everyone who has not should read "The Big Ball of Mud" ( http://www.laputan.org/mud/ ) which is about the clearest description I have seen of the reality of legacy systems.
- Port from Java to C#. The Java teams no longer maintained it beyond critical fixes. Many original experts were gone. My team had more vested interest and expertise. We wanted to be in charge of this system, and we work in C#.
- Switch from horizontal processing to vertical processing. The original system would query one record, query one associated record, query yet another associated record, etc. to form the complete picture of one entity. The new system reads and writes batches of homogeneous data. (Kind of a leap from OOP to DOD in terms of database accesses.) This optimization was largely necessary as the old system could not keep up with load spikes.
- Along with the previous change, we completely separated the data processing into two giant phases: one read-only step and one write-only step. This allows the read-only step to target replication databases. This reduces the stress on the primary database, which is a huge win.
- Detach the system so that it can be run, tested, and released in isolation. The original system was part of a much larger whole, so any fixes or enhancements had to wait for weekly or biweekly monolithic releases.
Everything you expect from a rewrite happened but not to a devastating degree. It went overbudget, missed key bug fixes, etc. but we worked through them and came out the other side with all the wins we hoped for. Being able to deploy multiple times within a single day is amazing. Our iteration time on new bugs or features is very small. The overall architecture is simpler and easier to navigate. I have confidence that a member of my team could enter the code base cold and fix a bug within a day.
So I dunno. AMA?
What was the estimated cost of developers being literate in more than one language? How significant was this in the decision?
It is also possible that some of the unknown scope becomes obsolete, but if that is happening, then you are already in a long-drawn-out process.
The best way to go through a rewrite or upgrade is to get the business involved at the start and throughout the process. If you fail to do this when something is missed, overlooked or done differently and you will be at fault. If you include them it becomes a not feature not necessary.
I've been on a project where we had a working system, but it had some severe technical platform & product value limitations, and we knew those limitations were costing us real $$$, both in support burden and market share vs legacy incumbents and competitors.
Plus, we had a "ticking time bomb" because it was a large-scale data system, but the prototype (which became prod) was not designed up-front to handle horizontal sharding, and we were at the limits of vertical scaling, and were projected to hit the "max limit" of that vertical scaling within 12 months, given current growth rate.
Thus, we began a rewrite, with full knowledge of how dangerous it was -- we even circulated Brooks's essay on "the second system effect" and had several team discussions about it during the specification stage of the rewrite.
In the end, the project was a success, and powered 3+ years of scaled growth (the current live data storage of the system is 100x what it was when the rewrite began, and our "hard limit" was around 2x). The rewrite also helped us make the system more scalable, competitive, and mature, not by throwing away edge cases, but by choosing an architecture that didn't cut corners on areas that, we discovered during user feedback from the v1, were non-negotiable core areas of value for our customer use cases. We had relaxed many requirements in the "prototype stage" of the v1, merely in the interest of getting something working out the door in front of customers.
The last piece of de-risking we did is to run both systems in parallel with our users for several months. This allowed us to e.g. let 10% of our users into the new system at first, ensure we weren't breaking any of their use cases, then let 20% in, 50% in, and so on. We could also do user interviews throughout. Since the new system involved not just a better data backend, but also faster response times, a modernized UI, and many new features, lots of people wanted in. We even had a waitlist, at one point.
Then, we cut the stragglers over, and cut the old system loose -- which felt great, BTW! Running two production systems in parallel isn't easy, but was absolutely the right thing to do.
With hindsight being 20/20, we feel firmly that our first system was a "prototype that went to prod", and that we followed Brooks's advice to "plan to throw one away, because you will anyway". And that we executed a "successful rewrite". But it certainly wasn't easy.
I'm really proud of that project, but I also feel it was a bit of a harrowing experience, especially near the end, when we were concerned some "showstopping bugs" were going to keep the progress bar at 99% for a couple extra months. But we made it through.
Perhaps the reason my outcome is better is because the need to rewrite wasn't driven by a framework or architecture du jour, but by real business requirements and real scaling requirements. Even then, I think sometimes those requirements can be overstated and the ability for an existing architecture to cope can be understated. I feel confident we made the right call, but I think it takes real expertise -- and healthy dose of skepticism -- to take on the full rewrite risk with eyes wide open.
(p.s. now, 3 years later, the same team is being forced to rewrite a significant portion of the backend, not for any business requirement or scale reason, but because of bitrot of a stable open source database engine version which needs to be upgraded to avoid EOL, and wherein the new version introduces backwards-incompatible breaking changes to the API and schemas. At least in this case, it's "only" a backend migration, and not a total rewrite. But, I'll tell you that it sucks to realize this is just required maintenance, thus a pure development cost with little customer benefit, rather than a project to introduce a step-level change to the product and business. C'est la vie!)
The faster they work, the more over-engineered the results? :-)
After a clifhanger like that, it better be good. :)
Half the comments on this thread have already alluded to it: A careful, piece-by-piece replacement, delivering the new system incrementally (running in parallel with the old system if need be.) Or even a gradual refactoring of the original code base if the base technology wasn't the problem. Read the success stories in this thread - you'll see they tend to follow this pattern.
You can think of software and how it fits with its users purposes as having multiple layers of features, interactions and details.
The top few features usually each have a few sub features that you have to tailor correctly to your users work flows, each of these sub feature often have sub details, interactions with other features, corner cases, data formats or cross compatibility considerations.
It's a fundamentally mathematical issue. Adding a layer of detail is an exponential proposition and the exponential function is explosive in nature. If your software only has 3 major features and each have 3 sub features, that is 9 things to consider. If each sub feature has 3 corner cases you need to get right, that is 27 things. And if each of these 27 things have 3 details that gets you to 81 considerations. The next level is 243. Each of the 243 points tend to take a similar ammount of time to plan out and build weather it's at the top or the bottom of the pyramid.
As a piece of software evolves over the years, the intricate details at the bottom get sculpted out. The software's fit to its purposes can become very fine.
The thing is, people tend to think of software in terms of its top 2-3 layers of details, its 27 most important features. The finer complexity is often not as visible and it's just difficult for humans to keep so many items in mind.
This is true of any complex technologies. People tend to think of cars as machines that have properties concerning speed and direction and suspensions and breaking but rarely think of the complex chemistry of the fuel aeration and combustion process, the complex internal forces and velocities of parts in the transmission, the carefully tweaked metallurgic alloy work that has gone in each of the thousands of metal parts, the carefully chosen properties of the plastics and the rubbers, the thermal properties of everything, the hundreds of electronic systems, the analog circuits, the thousands of specification lines of many dozens of communication protocols for controlling all these parts, the entertainment system, acoustic properties etc, etc, etc. All these things took decades to evolve and refine.
It's easy when planning a rewrite, to plan out the few dozen items visible at the top of the iceberg and underestimate the amount of important details hiding bellow.
Sometimes, some of the details might not be important, but often times they are the essence of your business case and at the core of the favourable economics of your product. It's the years of knowledge accumulated in the subtle details that provide you with a moat and makes it difficult for your product to be copied.
If you rebuild with mostly the top few dozens of features in mind, and only vague ideas about everything else, you are likely to be creating a commodity solution, one that your competitors will have a much easier time to copy than your older finely tailored solution.
It is often viable to take the meat of a system and place it in a new framework with some alterations.
And for the performance issue, just identify what causes it by timing parts of your code. I'm guessing home grown database wrapper.
That just means that there are cases that justify the cost, and they happen to me more extreme than the reasons we generally like to rewrite things.
Eric Normand addresses one facet of this by saying "you can't refactor Aristotle to Newton": https://lispcast.com/building-composable-abstractions/