LibreOffice: cleaning and re-factoring a giant code-base [video]
fosdem.org
fosdem.org
Looks like a project I worked on. It's eerie how similar it is (and no, it was not an office app, it was something different).
At least the comments weren't in German, but there were a couple of ones in Dutch. (Not that the comments were of any help)
No unit tests: CHECK. Twenty-five year of technical debt, CHECK. And now we're supposed to make it work in a different platform
So thank you, but I don't plan on touching a monolithical piece of old C++ ever again. Not even with gloves.
Start by adding integration tests, then (try to) identify the core of the project and isolate that in its own sub-project in the repo. Add unit tests for the functions in the core and then split the remainder into parts that are peripheral and that could be removed/disabled. (say a spellchecker). Next up re-factor the core until you're happy with the state of affairs, but keep the interface the same. If there was no interface to the rest of the code then you'll need to define one and implement that as well as changing the peripheral code to use that interface. Then look at each peripheral piece and decide if it is worth saving, needs rebuilding or refactor it as well.
The basic trick is to reduce scope until a subsection of the work becomes tractable.
With a really large project this is a multi-year effort with a team of seasoned programmers, it's hard and it means that while you do it there will be no honoring feature requests. You need management buy-in, patience and perseverance but that goes for any large effort.
Building fresh and green stuff is obviously much more fun to do but taking an old codebase and making it nice is (to me, at least) just as rewarding.
In my case there were added difficulties:
- The project ran on embedded hw (hence it does not have a user interface, only interface is with hw devices)
- There was a lack of problem domain understanding, so it's not - like in LibreOffice case, for example, 'this is a list of paragraphs', but 'we don't know what's this piece of data' because if refers to a specific domain of knowledge
Not to mention unit tests (both writing and running) are more painful in C++ than in Python for example
This causes all kinds of misery.
I'd see this being C++ over Python and embedded as advantages rather than disadvantages! Each to his own I guess :)
I could never imagine someone on their own accord, going by themselves into a 30-year-old codebase and refactoring it, just to make it "look nice". On the other hand people constantly build new stuff (too much of it really). I think it would be fine to admit that working on new code is more desirable than working on old code. Sometimes the latter needs to be done, but I think most programmers are happy to avoid maintenance work when they can.
Both are fun, assuming you stay in long enough the first will at some point turn into the second, no matter how well your intentions were at the beginning ('this one will be different, this time I'll get it right').
If you're young I can see how you might have that impression. But I know quite a few programmers personally that have been working on the same codebase for more than a decade, and even a couple that have been working on the same codebase for 3 decades.
The web is still young enough that most people that came into programming through building websites have no idea how long most code bases are alive.
There is a neat little proverb here: programs are like children, you can start one in an evening but you'll end up supporting it for the rest of your life.
However I don't agree with glorifying old codebases or putting them as "equal" to working on new stuff. If you're happy doing that, then fine, no one's going to stop you from doing it, but I do think that in general, given the choice, almost anyone would pick building something anew over working on a legacy codebase.
ok
> However I don't agree with glorifying old codebases or putting them as "equal" to working on new stuff.
Ok, so we will disagree on that then. I think that both jobs, new stuff and maintaining old stuff are equally rewarding and are essential skills. If all you can do is make new shiny MVPs you'll never run a business.
> If you're happy doing that, then fine, no one's going to stop you from doing it, but I do think that in general, given the choice, almost anyone would pick building something anew over working on a legacy codebase.
You won't be given the choice, unless you keep running away from your creations.
"Legacy codebase" is another word for "important codebase".
When you're coding, you solve the most important problems first. Thus, in some sense, the older the code base, the more important the problem it solves.
We have a lot more tools than programmers did back then, but tools are incapable of long term thinking.
OK, or maybe not. But it can happen, and it does, all the time, in really good shops with very experienced coders using all the right techniques and best intentions. I know. Dear God, I know.
Of course anyone would rather bootstrap new code and feel the thrill of new object models, shiny, hip new libraries breezing through their editors. But until you've experienced both sides of the equation, the long-term consequences of technical debt growing through a codebase like some kind of...virus...you'll not know how important the second-order effects of refactoring, TDD, and truly clean code really are.
As far as making the code "look nice" -- you've missed the point. The goal is to make it functional, readable and extensible again, and doing that to a 1M+ line codebase (which I've had to do) can be enormously rewarding, the business value gained more than tangible.
One of my "hobbies" that I unfortunately have too little time to delve into, is to find old code that I want to see live on, especially old Amiga code, and make it live on. Some I "only" make compile for AROS (an AmigaOS re-implementation). Some I try to port to e.g. Linux.
Part of it is nostalgia. Part of it the learning experience - code written for a system with 512KB or RAM and a floppy is often structured very differently, and sometimes the approaches are still interesting. Part of it that it is just relaxing to clean up something where all the hard implementation decisions have been done.
With more and more experience of that kind and by working on stinking code-bases, I've come to the conclusion that while in the past I could have thought that trashing the code and starting from scratch would help, now I would probably approach most problems by pushing new code in the form that I want and transition the rest as changes are required.
I had projects that I did myself with great design care, but after 5-6 years due to shifting requirements also started to look like you could have done a better job by starting from scratch again. Reality is, in retrospect all code is suboptimal.
I don't get why designers are always doing stuff like redesigning website X for free, without being asked to, but miss great opportunities to be in the spotlight by participating in projects like GNome, KDE or the likes. There are many examples of good OSS projects that could use some UX/UI love, one that comes to mind right now is Moodle, for example.
Is it designers that don't like to participate in OSS? Are programmers making it harder for them to join these type of projects or is simply the fact that they don't bring to much excitement or challenge to the design community?
But you know, it might be that they work primarily in photoshop, which is runs on Windows / Mac OS??
Simple.
Because if they did assist with FOSS projects much of the time they would be dealing with developers who (a) think they know what users want better (b) have no respect for designers and (c) believe obscure feature X is more important than UI.
Remember this is the same group of people e.g. on Slashdot/Reddit who think the iPad is an overpriced toy and think people are sheep for buying it.
For my personal use case for LibreOffice specifically, I don't give a damn how ugly it is, and I'm willing to have some patience if it's a bit slow, but crashing or hanging on a 160-page document is unacceptable. So adding a designer to the project wouldn't add any value to me.
But FFS why did you have to bring apple in to this. Absolutely nothing to do with this, quit it.
That's because it looks like everything designers are touching (Gnome3, Unity and partly Firefox) starts to turn to crap.
Yes, I completely respect Designers and I understand that they can do wonders for every FLOSS project and I know that this sentence is completely biased...but there are many examples out there were a designer (or multiple) are trying to stuff their vision down our throats without any afterthoughts.
No, actually he made a very good point by bringing Apple into it. Apple, see, is the very place were design and UX reigns supreme over other considerations (of course they also have faults, every team has).
So, OSS developers bashing Apple and its UIs, are not the very best community to nurture biding UI designers.
Within a single release cycle, these designers managed to destroy what was once among the most vibrant of all open source projects. The existing GNOME users have been alienated by the absolutely horrible UI. Among their ranks were many good developers and other contributors, too, whose absence now prevents the project from recovering. A large portion of these users moved on to KDE or Xfce, rather than MATE, causing further harm to the GNOME project and community.
There are other cases, too. Windows 8 and the UI changes starting with Firefox 4 are other recent, and major, examples of the work of designers going terribly wrong. The end result is software that is nearly unusable, or much less usable than previous versions.
The skepticism you describe is very well earned and deserved.
If you could give people a completely fresh, unprejudiced choice between Gnome 2 & 3, Windows 7 & 8, or Firefox 3.5 & 18, I bet many of them would choose the newer version. I don't think (most of) these UI decisions are objectively bad, and over the long run they might even pay off. But when the user upgrades and suddenly can't see how to make it Frobnicate, they get angry.
I also think that the attractiveness of one of these projects can vary quite heavily depending on whether you're a UI/UX designer or a programmer, and there's always the opportunity cost of giving up working on something new and fancy.
Because redesigning famous X website does bring them attention in their circles AND its something that they can do by themselves in little time.
Whereas participating in projects like Gnome/KDE is a messy, long term, no pay, proposition where they'll also have to put up with the opinions of non-designer contributors and be at their coding mercy.
Plus, they like to work on the platforms they use, and almost every professional designer uses either OS X or Windows.
For one, a Windows designer has access to Office if he wants an office suite. Why would he use LibreOffice? The price wouldn't entice him much, Office costs less than what he makes in a day's project.
Then again, I doubt designers ever hit the limits of Office on OS X.
Like level one of a libreoffice effort would be to organize all the buttons on all the dialogs and make sure they are consistent, or something like that. There are probably a couple hundred standard ones, it's not sexy. It is some work though.
A level 10 would be doing something like reinventing the ribbon from office, which might be more of a gtk+ task than a libreoffice task, and the libreoffice guys might not even think that's in their domain or their problem. I think that could be a justifiable position too.
You have to actually collaborate a lot with the team just to find all of the work for both of these. Ux guys don't seem to be dedicated enough to those projects. And fwiw, the projects themselves might not be open to spending energy "reskinning" especially when libreoffice has so much other technical debt.
I think 'reinventing the ribbon from office' is exactly what I don't want as a daily user of Libreoffice/OpenOffice. Surely we can have a competition to find a new paradigm for an office UI? Ubuntu searchable menus (HUD) coupled with execution of the command from the search results would be something I'd be interested in.
You mean like the OSX help system? Yes, feature discoverability is the main issue I have with Office ribbons. For this reason alone it's IMO not well thought through. iWorks still has the best / most consistant UI, if only its features weren't so lacking. I'd much rather have LibreOffice copy these concepts, such as the info window, table insert hotkeys and spreadsheets with multiple tables per sheet.
Is that something beyond what Ubuntu already has? Just hit enter or click on the result and it executes.
Great: pop your screen mockups somewhere first though. LibreOffice/OpenOffice has a LOT of users, and many of us have the interface and shortcuts wired into our neuron structure. Basically, gradualism might pay benefits...
The worst I've seen is where users have been encouraged to embed UI elements (such as HTML or nonprinting ASCII) in raw data or settings to trigger certain functionality. That'll drive a UI designer to drink.
https://wiki.documentfoundation.org/Development/WidgetLayout
With the previous code, the dialogs couldn't flex or expand, so they had to be sized for the longest translated text. Which, as you can imagine, was not only ugly but poor for usability.
It is now much easier for them to redesign dialogs for UX/UI reasons.
An oldie but a goldie: http://www.joelonsoftware.com/articles/fog0000000069.html
EDIT While I was reading the thread, it appears @jodrellblank had also posted it in the discussion elsewhere in the thread.
I think you need to have a really good excuse not to completely rebuild your product from scratch every now and again. Because if you don't, someone else will.
And I suppose we should also burn down factories and rebuild them every now and then.
The world of software is moving a lot faster than the physical world, so things can become outdated pretty quickly. Compare a 15 year old browser to a 15 year old factory.
Sure, it's nice to have the latest and greatest. But I don't think it gives you magical powers. You still need to do all the other stuff.
That, and the well-known problems with rewriting (the track record is notso hotso) mean that really, the onus for making a case to rewrite is on the rewriters.
If the market starts shifting to alloy wheels, and the reduced volume of sales for your steel wheels starts to affect the profitability of your factory, then maybe it makes sense to start looking at investing in updating your process to make steel wheels more cheaply at lower volume or to dump steel wheels entirely and start manufacturing alloys.
The Air Force Heavy Press Program made the rounds on the internet last year when Alcoa refurbished its 50,000 ton forging press in Cleveland [1]. It's been around since 1955, and it's almost unique in its capabilities within the USA. The idea that it should be replaced simply because it's old is, frankly, insane.
That's obviously a pretty extreme example, but the fundamental requirements for a metal lathe or end mill haven't changed in years. Yes, there are multi-axis CNC machines that can do vastly more operations before somebody has to change the setup on them than on a basic lathe or mill. If you aren't making something that requires a lot of steps, the basic lathe or mill might still viable.
Just because the rate of change in requirements over time is huge in software doesn't mean that it's true everywhere. If an old factory can continue to meet the requirements, it's probably hard to justify changing it just to be up to date.
EDIT: Forgot the link
[1] http://www.theatlantic.com/magazine/archive/2012/03/iron-gia... [1]
Every line in your legacy code is there for a reason. Sometimes the reasons are bad, but often they're there because of some very valuable lesson you learned from a customer, a mistake you made that your competitor is also likely to eventually make. Throwing away your legacy code means that it's likely you'll make the same mistake again.
I know this probably means that I don't unit test right. It also obviously means I don't code right, because I have bugs. The point is, these are facts of (real) life, and to be successful you have to manage the reality. Often this means that the best course of action is to stick with the legacy code.
What legacy code does is inhibit your ability to make the fundamental changes you need to implement in order to stay competitive. And I'm not saying throw you code out every other week. And I'm not saying every attempt to rewrite your product should replace what you already have. But the for all the costs associated with re-writing your product, they pale into insignificance when compared to being put out of business.
I'm not saying you should never rewrite a system from scratch. But I do think it's extremely costly to do so, and that it's extremely difficult to predict when you'll be driven out of business if you don't. So I don't think it's usually justifiable.
I'd be very interested if someone could supply an example of a company that was put out of business because they failed to re-write their product.
But perhaps we can agree, then: rewriting code is a business decision, driven by the business needs of the company. Premature rewriting is just another form of premature optimization, and can get you in trouble by putting your resources in the wrong places. But when you have a tangible threat- when you are unable to adapt to accommodate needs that you are convinced are on your horizon, then you should do it.
There's always the possibility that your existing code doesn't suck, and was designed to be flexible for changes.
Ideally, you've done a good job and can rewrite small pieces of your system iteratively as needed, rather than having to chuck the whole thing.
Also, code that might look crap and seem to do weird stuff often handles intricate edge cases. But due to the lack of comments, it's not obvious what's going on.
I've been involved in several component refactorings when as a team we've gone "look how complicated the old code is, we can make it much simpler".
When we've re-written it and then tested it with actual production use cases, it then becomes obvious why the code was so convoluted previously. So time was wasted, but at least there are comments now (and in those cases, the code's slightly better).
Because that's exactly what someone else is going to do. The company that makes the product that competes with yours is unlikely to have picked your product at random and simply decided to compete against it on a whim. There's a good chance they've been using your product and have become so frustrated with it's shortcomings that they've had a crack a building there own version. They've been on your support forums that detail every bug and bad descision, they've read the blog post detailing exactly what is wrong with your features and how they could be done better. They've stared at their screen in a full on rage wondering why the hell your product doesn't do things in a way that makes sense. They have their own experience and they are not afraid to use it.
There could be one, tens or hundreds of people who decide to do this. Many will fail, but if one succeeds then you have a serious problem. The continued success of your product can no longer rely on it's features unless you're prepared to make some fundamental changes. Only now you have a deadline, whereas before you could have worked at your own pace.
You're right when you say your legacy code is your greatest asset. It's proof of what is a good idea and what is a bad idea. But that doesn't mean you should keep it around any longer than you have to.
Ideally, yes! Realistically, there's always some hairy, gnarly wisdom "baked into" the code.
The definitive article is Joel Sposky's article on the perils of writing from scratch. If you only want to read a few paragraphs:
"Back to that two page function. Yes, I know, it's just a simple function to display a window, but it has grown little hairs and stuff on it and nobody knows why. Well, I'll tell you why: those are bug fixes. One of them fixes that bug that Nancy had when she tried to install the thing on a computer that didn't have Internet Explorer. Another one fixes that bug that occurs in low memory conditions. Another one fixes that bug that occurred when the file is on a floppy disk and the user yanks out the disk in the middle. That LoadLibrary call is ugly but it makes the code work on old versions of Windows 95.
Each of these bugs took weeks of real-world usage before they were found. The programmer might have spent a couple of days reproducing the bug in the lab and fixing it. If it's like a lot of bugs, the fix might be one line of code, or it might even be a couple of characters, but a lot of work and time went into those two characters.
When you throw away code and start from scratch, you are throwing away all that knowledge. All those collected bug fixes. Years of programming work."
http://www.joelonsoftware.com/articles/fog0000000069.html
Ideally, yeah, you take that wisdom with you to the new project but realistically (even if you've followed great documentation practices, etc, for all of those hairy kludges) the rewrite-from-scratch is never as easy as one thinks.
Code becomes legacy pretty quickly, in my experience... Architectural decisions that were obvious before you had real customers rarely stand up to real usage.
The company that's "going to put you out of business" also doesn't currently have any working or revenue-generating code, and are unlikely to immediately understand the architectural challenges...
> I think you need to have a really good excuse not to completely rebuild your product from scratch every now and again
It pretty much just doesn't work in the real world; how's that for an excuse? As other people have said, you need to scaffold interfaces, and rebuild small parts. Any company that isn't continually reinvesting in making their codebase better will accumulate technical debt pretty fast; but the throw it away and start from scratch approach rarely works on big systems.
It is just someone else saying the same thing, or it's a real world example, or it's anecdotes-aren't-data.
The company that's going to put you out of business is using an early version of their own legacy code.
It's not always obvious when you should attempt a rewrite from scratch. Having done a few refactorings, I've noticed one desirable trait in good code bases: If you can implement new logic easily in a few lines of code, it's probably a good design. But if you need to get out the surgeon's knife and write tons of new code every time you want to implement new logic, it might be fundamentally flawed enough to warrant a complete rewrite. A good design is adaptable and will allow you to be resilient against newcomers.
And furthermore, while the developers are rewriting the code, who is going to maintain the old code? The demands for new features and bug fixes aren't going to stop just because you've decided to rewrite the code.
"Things You Should Never Do" http://www.joelonsoftware.com/articles/fog0000000069.html
Sure, we know have Firefox, but that's little consolation for the corporate entity that was Netscape.
Even assuming their newly-minted code is better (which others have disputed), what they don't have is loyal customers and brand recognition. If you're worth competing with, you do.
>> I think you need to have a really good excuse not to completely rebuild your product from scratch every now and again.
There could be many reasons. 1) That's expensive. 2) May introduce new bugs. 3) Interface changes may alienate existing customers. 4) Switching technologies may require new developers.
I wouldn't say "plan to rebuild from scratch periodically." I'd say "design your software to be flexible, watch your competitors, and iterate to ensure they don't leave you in the dust." If it becomes clear that you can't compete on features without a rewrite, then consider it.
What I was surprised to see is the word 'resurrect'. As if they wanted to imply LibreOffice was dead.
Loaded it into Office and it worked fine.