Tortoise: Shell-Shockingly-Good Kubernetes Autoscaling
github.com
github.com
Also, Mercari has embraced a microservices architecture, currently managing over 1000 Deployments, each with its dedicated development team.
To effectively drive FinOps across such a sprawling landscape, it's clear that the platform team cannot individually optimize all services. As a result, they provide a plethora of tools and guidelines to simplify the process of the Kubernetes optimization for service owners.
But, even with them, manually optimizing various parameters across different resources, such as resource requests/limits, HPA parameters, and Golang runtime environment variables, presents a substantial challenge.
Furthermore, this optimization demands engineering efforts from each team constantly - adjustments are necessary whenever there’s a change impacting a resource usage, which can occur frequently: Changes in implementation can alter resource consumption patterns, fluctuations in traffic volume are common, etc.
Therefore, to keep our Kubernetes clusters optimized, it would necessitate mandating all teams to perpetually engage in complex manual optimization processes indefinitely, or until Mercari goes out of business.
To address these challenges, the platform team has embarked on developing Tortoise, an automated solution designed to meet all Kubernetes resource optimization needs.
This approach shifts the optimization responsibility from service owners to the platform team (Tortoises), allowing for comprehensive tuning by the platform team to ensure all Tortoises in the cluster adapts to each workload. On the other hand, service owners are required to configure only a minimal number of parameters to initiate autoscaling with Tortoise, significantly simplifying their involvement.
It happens shockingly often that applications only support working with a single replica and even worse when those applications cannot run concurrently with replicas of themselves which prevent smooth rolling updates.
IME if applications are fault tolerant of restarts, or support concurrent replicas then scaling up and down to meet demand is absolutely fine.
Also this reads like a cry for help:
> Therefore, to keep our Kubernetes clusters optimized, it would necessitate mandating all teams to perpetually engage in complex manual optimization processes indefinitely, or until Mercari goes out of business.
I had the same initial reaction, though. "Kubernetes? Shellshock? Oh noes!"
Additionally, many startups develop apps for scenarios where the database scaling can be solved. Typically they have many customers and can shard the data, or they can tolerate eventual consistency and hence can just throw caching layers at the problem.
Transactional enterprise apps can't do this, they're single instance and caching could cause all sorts of weird data loss or corruption. Hence, the database becomes the most common bottleneck.
For more advanced use cases there is keda - https://keda.sh/
There should be a k8s metric of 'cluster utilization'/'packing metric' the autoscaler from k8s should reschedule pods but rescheduling only exists for node drain, Priority and resource issues but not for packing a cluster.
You can only do this with Deschedule, which is not part of the k8s distri itself...
Thats it about my generic k8s rant (i do love k8s nonetheless).
Whats missing is an annotation on the specs to clarify how aggressive an autoscaler can handle certain pods (like do not move this stateful workload ever, reschedule this type of pod at night or look at this metric to see when its less active).
From an algorithm standpoint, its probably the same issue as java has with old and young generation: there are plenty of stable pods, which just need to run always. The base load, easy to pack, easy to keep aligned, seldom repackaking needed. Than you have the young generation: unclear how long the workload is required. Might mean that you have a big node running nearly empthy for an additional x hours just for the pod to finish its task or for the load to scale down again. If you know more about the type of workload, you could decide to spin up a big node or a lot more smaller nodes from a number of node pools. If you know that this job runs for 20 hours but is gone after, put it on a small node.
You could also do a strategy (depending on the cluster size) to only have one big node between 1-99% and make sure that all other nodes are always packed.
The project itself is tbh. shitty in describing how it works. I was neither able to read in the README or in the linked blogpost HOW they are doing it. The key information is missing. What algorithm does it use to utilize it?
I do love to have more autoscalers though.
It has the same issues as the project from this hn thread: Its not part of k8s and it also doesn't tell you their strategy at a glance.
I'm using managed kubernetes, but from a relatively small provider and I doubt they did any custom coding for their autoscaler.
I agree that it could do more. Sometimes I have to manually juggle pods around to remove one node, if scheduler did bad job.
It has to know how to schedule new nodes which is provider dependend.
But thats my critisism: I believe that the autoscaler should be a k8s core component and not a 'select whatever opensource project you like based on no evidence OR if you are lucky based on experience someone else made over the years with all the choices out there'.
But as you write it yourself: You still need to do something manually and i bet you don't have that much control over telling the autoscaler to act differently depending on the pods.
The decision or packing algorithm could be architectured to be more extensable perhaps or finetuned but the core should be there and i do miss specs which would tell any autoscaler how to act/interact with different pods.
I would also love to see an autoscaler simulator.
Its a hard problem at the heart of k8s which gets fully ignored.
https://keda.sh/ - project website
https://github.com/kedacore/ - code
This allows our scheduler to place/pack the workload according to our preference.