Keeping track of technical debt in source code
philippe.bourgau.net
philippe.bourgau.net
>>>The great thing about TODO comments is that, as a very old programming trick, they are already supported out of the box by most tools IntelliJ, SonarQube, Rails, CodeClimate and I guess many others. Only one day after I refactored to TODO comments, a team mate fixed one that had appeared in his IDE’s TODO tab
Instead
1. Quickly experiment with refactoring but stop if it feels wrong. We often underestimate how quickly we can improve the codebase. It's often better to try something quickly than waste time planning to do it.
2. Defer if it's a big refactor or you're on a very tight deadline. Write a reminder in your physical notebook so you don't prematurely pollute your board. If it feels high priority after that and you still haven't had time to do it, then make a chore ticket for it.
Placing TODO and ASAPs in code helps ensure you don't leave any loose ends, and if you call another developer onto your feature branch allows them to get a feel for how you're getting on and how complete the code is.
Using a physical notebook distances the workload from the code and hides the fact something is incomplete from others.
A TODO graveyard can be prevented when you review the code - a developer who thinks a feature branch littered with TODO annotations is complete code is either dreaming up future use cases or failing to complete the work assigned.
> Using a physical notebook distances the workload from the code and hides the fact something is incomplete from others.
Exactly. In fact, I use both (except the notebook is an org-mode file, not a dead-tree book). TODOs / FIXMEs in code are for things very context-dependent and connected to code, for which it's important to see them in surrounding context. In more abstract cases, I just log a task in my notebook, to be done later.
Examples of both - code:
//TODO resourcify
(i.e. move to external string resources database later)vs. notebook:
** TODO Refactor the refactor of the SomeClass class.
--Also, often TODO / FIXME comments serve for me as anchors in some larger refactoring processes. I tend to label them specifically for easy grepping (e.g. with unique Unicode chars like ). The point is, refactoring process may take few days, during which I also may need to do some other tasks, so whenever there's a piece of code that needs to be redone but not right now, I tag it for later; I can then use grep/IDE search to check if I didn't miss anything.
//XXX this is bad
//XXXX this is more bad
//XXXXX this is EXTREMELY bad
Allows for easy grepability and a rough, at a glance evaluation of the seriousness of the hack/debtIf you want other developers to see these TODOs in mainline, then they shouldn't exist and instead be in your teams project tracker.
Would you put every task as a TODO in code? No of course not. Its not easy to manage by non-coders, low visibility to anyone not in the file in question, no priority or organizational quality, no way to prune or plan these TODOs, etc etc. If you would wouldn't put every task in a TODO, why put any task?
If you invest heavily in TODOs you just end up with things that should tickets and aren't or vague desires that would never end up as tickets because they're not well thought out anyway.
Often, by the time you officialize a TODO into the ticket pipeline, you could've just fixed it yourself. But the reason you dropped a TODO was because you were already busy with something else more important.
TODOs just sit at a local, unofficial scope in the code. I'm not sure what kind of things you put into TODOs, but if you were to generate a ticket from each TODO in the code I work on, you'd have generated a bunch of low value ticket garbage.
The problem with a rule like "no TODOs allowed in the code" is that TODOs are issues that aren't usually worthy of a ticket, or at least not the time to turn it into one at that exact moment. You'd just be blind to more issues. Not a reason to pat yourself on the back for 'clean code'.
Take the Tensorflow codebase (and Google and Microsoft code, and linux in general). There are currently ~1,028 TODO's in Tensorflow, most of which refer to a particular engineer in the form TODO(<name>):<comment>. This is pragmatism and good judgement in action.
Software engineering is all kinds of shades of grey, competing considerations, trade-offs and balances. It's what makes it interesting.
In other places, I've seen just blanket `// TODO: comment` is pretty bad. There's no accountability and also you don't know when it's obsolete.
(I've got a thing in vim that runs git blame on the current line, and pops up the commit message in the quickfix window).
I find the usefulness/annoying ratio to be very high.
Also, not all tech debt is solved by refactoring.
TODO comments aren't suitable for tracking all technical debt, but they're good for targeted instances.
If you want to fix/change something eventually at some indeterminate time in the future then comment TODOs are a good place for that sort of issue to live.
Whether fixed or not, it's valuable to future readers to know there's a gap.
TODOs are for optional/longer-term things as you described, like inefficient implementations because the current assumption is that it will ever only operate on 20 elements.
HACKs are more along the lines of "hey if you're debugging this bit of code, your issue is probably here"
Some technical debt spans multiple lines, or across files, modules, etc.
Issue trackers are everywhere nowadays. You've found something that needs doing? File an issue. Treat it like any other, it's not special.
The nice thing about putting todos in the code is, somebody will see that the next time they're working on that chunk of code. The bad thing is that they're highly localized, and frequently give little space to give a full explanation of the problem.
The combination solution that really does get you the best of both worlds is to create a ticket, and then reference the ticket # with TODOs when there are bits of the code that are influenced by that technical debt. That way if you're seeing a todo in the code, you can easily get to a central ticket with more information. But also, when you're working on one of those tech debt tickets, you can search for the ticket # in the code base to verify that you really have found all known (by which I mean, somebody cared enough to make a note of it) places where the tech debt is manifesting itself in the code. Also, tech debt is less likely to become a dead letter. Programmers read the code incidentally to their work. The only people who habitually read the ticket queue incidentally to their work have PMO certifications.
I used to think the same; however the reality is if nobody cares about a given issue it's not an issue worth working on. Same with TODO or other notes in comments.
When that thing finally does get fixed or upgraded, it sure is nice to be able to just look up the master ticket for "Hacks to work around X" and easily find all the little things you can clean up right away.
But, of course, if your tooling and team procedures force you to choose between "fix immediately" and "DFC", I suppose that would result in a lot of decisions to not even worry about technical debt. Over time, you could probably even lull yourself into thinking it was never technical debt in the first place.
I agree. That is not what I said.
I try to write a test when I open a ticket, demonstrating the desired behavior, and mark it such that it's checked for failure (so we notice if we've fixed it in passing) unless an environment variable is set to the particular ticket number. Sometimes I also include a test demonstrating current behavior - this helps surface the tickets if someone is working on something related.
It's a cop-out, and unfortunately a hallmark of a class of developer most of us miss. It took working with roughly the same group of people for four+ years, twice, before I even noticed them. It's probably two groups with similar symptoms which I think contributed to my confusion.
First group are the fakers. They are people who want to appear high minded and will agree with all of your plans for reducing tech debt, except the first time there is any schedule tightness they'll throw it all away and claim that they wanted to do more, But The Deadline. You can spot these people because the next time there is schedule slack due to requirements not being ready, they will complain about there not being anything to do. Meanwhile their legitimately conscientious peers are already 12 hours into some refactor that the boss didn't tell them not to work on.
The second are Flow junkies. There are people on your team that do heavy lifting. Big rework, new architecture, they can do amazing things, occasionally startlingly fast (doing judo on the code to make it do things you didn't think it could).
Problem is, they are so addicted to The Flow that they have to wrestle dragons even if they don't exist. They can take the code far in directions nobody else understands, which must be because they are smart and you feel a little imposter syndrome.
But what really happens is that they make the code in their own mental image because they can, and that stacks the deck so that they are >3x more productive in that code than everybody else. This creates the 10x myth, because they are actually about 3x more productive in arbitrary code, but 10x more productive in code they've arranged so only they can really follow it. In the middle of the Flow stopping to clean up pesky things like argument consistency might break the Flow,
Their tell is that they often "don't have time" to fix annoying inconsistencies or bugs in their code, expecting people to avoid them, and yet they seem to have plenty of time to reinvent the wheel, writing frameworks or event processing systems when we already have tons of those to chose from. Because those Big Problems have the sort of time and space in them to both require and allow for a Flow State to happen.
I know this type because I'm an ex Flow junkie. I spent too much time studying ergonomics and realized that just because I could follow the code easily didn't mean it was great code. Though I could achieve Flow while doing truly annoying bookkeeping refractors, over a long period of time I found it harder to achieve the Flow state. Maybe I got old or maybe I was less motivated or maybe open offices killed it, who knows,
So I subscribed to the philosophy of, "a rising tide lifts all boats", and am just fine being a 3x developer who makes the code and process more productive for others.
git grep -EI "TODO|FIXME"
Then if I want numbers I can do:
git grep -EI "TODO|FIXME" |wc --line
Credit to https://www.commandlinefu.com/commands/view/12842/get-a-list...
In all seriousness, the great thing about a standard like this is that it is tool agnostic.
Then, there needs to be awareness and a sense of urgency. Technical debt accrues interest.
In my case, I use deprecation tags:
- @deprecated on javadoc for Java (http://download.java.net/java/jdk9/docs/api/java/lang/Deprec...), JSDoc for JavaScript (http://usejsdoc.org/tags-deprecated.html), Yard for Ruby (http://www.rubydoc.info/gems/yard/file/docs/Tags.md#deprecat...)
- Obsolete attribute on .NET (https://msdn.microsoft.com/en-us/library/22kk2b44(v=vs.90).a...)
- deprecated pragmas and macros in C++
- DeprecationWarning on Python (https://docs.python.org/3/library/warnings.html#warnings.war...)
By doing this:
- You get a deprecation warning.
- Some editors may highlight it.
This signals people to stop using that code.
This creates a sense of urgency to switch to or create a refactored alternative that is cleaner.
An alternative interpretation is that technical debt depreciates. That is, in a Darwinian way, clearly the product survived WITHOUT implementing the changes someone thought it needed. So is it really technical debt, or just baggage?
Sometimes, debts are written down.
Whether it is "baggage" or "debt" depends on whether it is accumulating interest (that is, it makes other changes such as adding features harder and harder to do).
But if your service is large enough, a maintenance window costs you money, reputation, customer satisfaction and makes you breach contracts and service level agreements... you may want to find ways to minimize maintenance downtime.
In that case you will want to make your code straightforward enough so mean time to detection and mean time to repair is low by emphasizing maintainability, namely, clean code with accompanying tests.
...and this is one reason that many business die, the assumption that past performance is indicative of future results.
The reality is that you are making a bet with unlimited downside (e.g. a naked call in stock parlance) when you don't address technical debt.
"Sometimes, debts are written down."
One word: Enron.
At my $dayjob I do Java, for which the IDE also lets me conveniently view the TODO/FIXME entries. I also have Sonar, which reminds me of them through low-importance warnings.
In Fish Shell I have:
function grp
grep -rni $argv
end
and I simply do "grp todo", "grp fixme", etc. Example output (it's actually also colored): src/main.lisp:63: ;; TODO FIXME introduce at some point a mapping for page list -> generation function, or something.
src/main.lisp:89:;;; FIXME deprecated
src/main.lisp:94:;;; FIXME unused and probably throwaway
src/site-components.lisp:9: ;; FIXME use some page structure map to map from (name language) to base filepath
src/site-components.lisp:31: (let ((colorize:*css-background-class* "no-paren-fx")) ;FIXME make configurable maybe?
That's good enough for me :). Also I have appropriate highlights set up in editors & IDEs I use :).In my experience there are lot of productive programmers who, overall, add value (stuff gets done, it's not a bug-infested disaster) but simply don't care. Might not even have heard the expression technical debt. Might care more about short-term reputation wins from Getting Stuff Done. They aren't going to leave TODOs.
That way it's
1.) Clear what the context on the todo is 2.) It's clearly in our backlog and can be prioritized accordingly.
We have some scripts that run at build time to verify all todos have the appropriate issue as well as manual code reviews on all changes and I honestly don't remember the last time it was done incorrectly.
// TODO(<Issue Number>): <Short message>The problem is that the TODOs get flagged as failures in the context of a pull request and require a manual override. The overall attitude in the tooling is TODOs as a code smell or problem vs. a solution.
Maybe this is not a big deal (since you can interpret the information however you choose) but it creates cultural resistance to embracing TODOs. We've recently started dabbling with TODOs again on my team but I still haven't overcome a bit of guilt I feel, like maybe "I'm doing it wrong" -- even though I think TODOs are a perfect fit for highlighting small tidbits of technical debt as the author describes here.
This is really good feedback, which we are addressing. We're going to change things up so that by default TODO issues are emitted as "Info" severity instead of "Minor", and we are going to change our PR integration so it does not fail PRs on "Info" issues.
As an aside, on Tuesday we launched a Grep engine, which is much more powerful than FIXME: https://codeclimate.com/changelog/58ecfa297705a149790008b2
It allows full customization of the emitted issues.
If your project uses them to indicate e.g. things that can not be merged into master, then obviously that's one thing, but when they're simply comments, the hate here seems misplaced. Other than (maybe) redis, I don't think I've ever seen code that did not have "potential things to do" that would have been nice to see commented.
All those issues have already been solved with a simple ticket system (github issues, JIRA tickets etc'). Why reinvent the wheel? What is the benefit here?
The #TODO comments live close to the code. If it gets updated (whether that's an actual fix for the tech debt while working in the area, or removing code altogether due to refactoring, there's no guarantee that the ticket will get updated - and then you spend time trying to figure out which actual tech debt tickets are up to date or not.
It also means that they're easily discoverable - you can stumble on one while working on something, without having to hunt through Jira for what to fix. You're also likely to understand how to fix it without having to come up to speed, since you're already working in that code.
If you only allows TODO after a discussion (e.g. TODO: We decided to forego implementing X at this time, but we may need it later, as per project lead Y).
It does not prevent prioritization or ownership, and there's nothing to prevent a corresponding entry in some management system (and indeed, you can put the ID from that system in the TODO).
That said, often there's things you want to signpost specifically for future developers (including yourself) which may not rise to the level of needing a ticket assigned or external tracking (or that may be complicated due to bureaucracy or or you measure goals being met). There are plenty of reasons to have an additional channel to signal information directly to a dev. Just don't make that instead off the official channels of communication without good cause.
"// TODO(dmi): update after moving to Java8" doesn't necessarily mean that it's _me_ that has to fix it, but that someone can come to me and say "Tell me more about this". This took me a while to learn/remember, but it's very helpful. The hardest part for me to keep in mind was that someone else's name on a TODO shouldn't mean I ignore it.
For example, IntelliJ, just "shift+shift+a.n.n+enter", bam! VSCode has a nice plugin that will do it, same with sublime.
PS. In VS - Go to View: "Task List", it gives an overview of all TODO's.
We examined the impact of SATD on software quality for forty open-source projects. To measure this, we took into account three criteria commonly associated with quality: (i) on the file level, the relationship between defects and SATD; (ii) on the change level, the potential of SATD to introduce future defects and (iii) the complexity SATD changes impose on the system. The results of our study indicate that: (i) SATD and defects exhibit no correlation at the file level, (ii) SATD changes make the system more susceptible to future defects than non-SATD changes do and (iii) SATD changes are more difficult to perform on the system. [2] (an abbreviated version [3])
There are more patterns to identify technical debt than "TODO, FIXME or XXX". You can find the rest of the patterns here [1].
It is very interesting to see developers confessing their workarounds and hacks through source code comments, as it can alert later contributors to the technical debt induced.
References:
[1] http://das.encs.concordia.ca/uploads/2016/01/Potdar_ICSME201...
[2] https://www.scribd.com/document/345197805/Sultan-Wehaibi-mas...
[3] http://das.encs.concordia.ca/uploads/2016/01/Wehaibi_SANER20...
Is this because it allows the devs more control over priority...like putting it in the work queue might result in outsiders deciding priority?
That's why sometimes the code is a better spot to leave a hint for the next poor sap -- or yourself -- who stumbles across the thing again.
Engineering owns these and gets to prioritise them as necessary. It's usually done with an eye to the current backlog of stories and bugs that product has prioritised.
Sometimes chores get done opportunistically. You have a time gap, you grab a short chore and do it.
Or you might spin one out of a story as a kind of marker of known tech debt.
Sometimes engineering sees a heavy piece of tech debt that is a millstone and just goes ahead with it. Usually this ties up a track of work that becomes unavailable for product work, sometimes for weeks.
Technical debt tickets were created in the issue tracker. The ticket offers plenty of space to detail the technical reasons for implementing potentially suboptimal code in the first place (a justification for the introduction of technical debt), a general outline of the approach to resolving the debt and to list out business justifications for resolving the debt.
A TODO comment referencing the ticket was added to the code. This reduced comment bloat and allowed a developer to easily find the relevant source from a ticket and vice versa.
Technical debt issues were added to each sprint/development cycle by product managers with developer input during planning sessions.
This process also had the benefit of letting developers fill in small amounts of possible development downtime. If you know you have to leave the office in 30 mins (a dentist appointment for example that you can't miss), you may be uninclined to start work on a large feature. Picking up almost any small tech debt ticket turns that time into something more productive than nothing.
Todos (and fixme, xxx) are very useful as review or editing anchors.
For example, if I'm adding a substantial new feature and it will affect the code in a few different places, I'll often go through first and leave easily searchable markers at the exact locations where I expect to make changes. Then I can do a quick comparison between where I'd expect to need changes given the design and where I've found to make them to ensure they match up before I start changing anything, and then I can move quickly through all of the required changes without losing focus.
Something similar could be used with todo comments in general, e.g. in your favorite language something along the lines of:
import todo
todo.add("Fix this later")
Which would keep track of those and you can somehow automate reporting, various restrictions and what not.As much of a rule-follower as I am... I think that if I always had to code for the eventual lawsuit, I'd quit coding.
I really like reading todos written by other devs to see where a feature was left incomplete. It's so much easier to see it right in the code rather than lost in a ticketing system.
Holy crap that's fragile!
I'm still looking for the value add of an IDE. I'm not negative toward them (I know many many people use them and get a lot of value out of them) but I personally have never found one that was easier than simply writing and compiling code. Could be my use cases (the kind of code I write/projects I do).
Part of the problem I think is taking your hands of the keyboard to manipulate the mouse.
Works well for me as I'm working through new code or refactoring since it makes the TODOs stick out so much against the dark theme I use.
Eg: // #TODO refactor
• HACK
• TODO
• UNDONE
• UnresolvedMergeConflictA single line is not enough to explain work needed to be done or getting said work done, but it can be enough to explain existing work, sometimes.
//TODO: JTF, Use new error handling strategy here
This way it is clear who put it there and makes it MUCH easier to keep track of your debt and not your buddies.2. Unless the dev team is very small, most TODOs belong in your issue tracking system. That's where they can be enumerated, described, discussed, prioritized, assigned, etc.
It helps me understand the original intent of the developer, but I don't trust that it's actually the right thing to do, and mostly shouldn't be there.
// TODO(author): fix this
that way you know who was TODO'ing
It's a bit case-by-case, but I think I'd rather encourage contributors to admit to the lacks in their code and accept them - but get them to commit to follow ups that will fix the problems.