> available 99.5% of the time with decent alerting for operators to kick it
An operator should never have to "kick" a service. It should repair itself, except for the occasional hardware replacement if one is working with bare metal. And for anything that's being sold as a product, as opposed to an internal tool, I think 99.9% availability should be the minimum.
But I don't know enough about Kubernetes to say whether it's overkill at the scale of just a few servers.