"Write an integration test for xyz" is not really what OKRs are designed for. You can do that (and lots of people do), but it tends to be frowned on. The reason for this is that it's not clear how it rolls up into the larger company OKRs. Where does "wrote integration test for xyz" fit into the company's "increase monthly active users by x%" KR?
I think the general trickiness is that OKRs are inherently backward-looking. A lot of technical improvements have to do with risk mitigation. The risk of nasty bugs (testing), the risk of future functionality being slow to develop because of poor architecture (refactoring), etc. But it isn't clear (to me) how to write OKRs for mitigating future risks.
You should always be making progress on OKRs, but it’s not the only thing you do. Engineering has its own overhead, not every hour is billable.
Let's consider a KR of 'increase active users by X%'.
* Interviewing/hiring isn't going to move the needle on getting X% more active users by itself. Having more engineering time available for new (critical) feature work might, however and you have to do interviewing to get new hires.
* Training in and of itself isn't going to move the needle either. But having a more knowledgable engineering team probably will, by being able to move faster or better. You have to do training to improve knowledge.
* Maintenance work might not move the needle on active users FORWARD (or it might, if you have a lot of bugs that prevent users from actually using your stuff) but it might prevent it from going BACKWARD.
This gets it backwards. This is a task. Why are you performing this task? Especially without being tied to specific development? If you can't justify doing it with your/teams/companies stated OKRs, you actually shouldn't be doing it. Even if you feel you should because 'its the right thing'. Even looking at it this way gets it backward.
You should be looking at your OKRs and figuring out what you should do to get there. But, you might have to break it down a little further.
Indeed, a KR of 'increase monthly active users by x%' is not directly actionable by an engineer. But an engineering department can come up with its own OKRs that fit in that direction.
For instance, to achieve that goal it might be necessary to develop new features. However, if 70% of the engineering time is spent in bug-fixing then they're not going to be able to do that. (Do you know where your team spends its time? That might be worthwhile to figure out).
So, an objective of 'Increase development time for features' (or some such) might be considered. If you find you're dealing with a lot of bugs, one of the key results might be to reduce bugreports by 50%. To do that you might argue that you should add some integration tests so that you can change code and catch issues before they get to production so that in the future you'll be spending less time on bugfixing / rework, so that you can work on new features that will entice new users to full fill the company objective.
"Don't have a catastrophic data breach"
There isn't a middle ground where there is a linear correlation between a serious compound failure and fixing bugs or writing tests.
You could come up with a methodology for a "security score" and track that. However it is both a transparent workaround for having specific bugs / tasks in OKRs and it is likely that your invented metric will be bogus and will lead you to do irrational things.
If a sizable chunk of work can't be covered by OKRs then why are you using them or how do you use them with the understanding that they cover partially and inconsistently work and achievements.
Put differently: you have a way of working that includes applying measures for security. For instance, you might do threat modeling at an early stage of your project/iteration. Those sessions may result in concrete work to perform. You might also perform a security test at the end of an iteration that confirms things have been covered. That adds to your workload and affects your ability to achieve goals in a certain way. You apply this way of working to move toward a goal.
So, over time, you will find out that applying these principles have a certain outcome. It may be positive: "we're not seeing security issues and development pace is OK" and it might be negative ("security tests are coming back bad", "we're hacked" or "development pace is snail-like"). When it's negative, that means you have to do things that will slow down achieving your goals.
That kind of negative feedback should lead to a poor(er) performance score on the KR when reviewing them. It's this kind of feedback that should factor back into your decisions moving forward to achieve the goal. In this example, if you have negative findings, you might choose to educate engineers on security matters, streamline your engineering principles, increase the team size and/or even replace people.
My point was that you should have a certain set of sane engineering principles (security being one area they should cover). They should be sufficient to todays standards. These principles are not/should not be business goals: they are tools in achieving goals in a responsible and reproducible way.
I am also saying that if you get feedback that these principles are keeping you from you should include them your evaluation in determining next steps to move forward without dictating a specific manner of how you should deal with them; that's up to the specific situation at hand.
There is no planning process that can effectively deal with long tail existential risks like that. It's silly to fault OKRs for not solving it when nothing else has.
Best case scenario is something along the lines of:
Objective: reduce annual carry cost of IT Data breach insurance KR: Do project X to insurance co.'s satisfaction to lower our premiums KR: prevent regressions in existing compliance by running regular audits
While this outsources the scoring problem to an insurer, at least they have multiple customers to amortize over and extract some data from.
You might improve on some measure of "velocity", but on a small dev team that's so noisy as to be nearly meaningless. I guess you could track employee frustration as a primary target.
Currently where I work we have a fixed allocation % of dev/test time for technical improvements. In practice it works really well if you approach it strategically and use it to achieve things over the long term. Things you can just chip away at that steadily improve the health of the code, architecture and product. I find if it's used on a purely operational basis where you just ad-hoc improve stuff you find then it's not super effective, but it's pretty solid when the team is onboard with leveraging it to achieve better and better DX over the long term. Otherwise "refactoring" barely ever gets done and products slowly suffer bit-rot.
I think OKRs are a perfect fit for such a scenario. We don't use them, but I kinda use them internally as part of my mental model.
That more or less mandates building new products, features and services.
Refactoring, tests etc don't fit anywhere in the OKR model.
Particularly for startups, as they grow, one engineering team objective is to remain efficient as the team is growing. The way to achieve that? By refactoring code and introducing automation (like CI, CD, etc).
This gets reflected in the cost of developer time per feature / project. Which can be represented as a KR in the. OKR framework.
Unfortunately it is mostly spent responding to infrastructure / dependency churn, just treading water rather than getting ahead, but it helps. No one is going to criticize you for refactoring what needs refactoring, as long as you are also making progress on your OKRs.