Every time I look at k8s networking seriously it gives me great pause on whether I should continue to run such a complex system. IPv6+EtcD would solve this matter really well.
Every time I look at k8s networking seriously it gives me great pause on whether I should continue to run such a complex system. IPv6+EtcD would solve this matter really well.
Networking in k8s is the poorest designed aspect of Kubernetes. So much so, I'm honestly surprised it didn't kill off K8s early on. You still can't get proper service layer load balancing out of the box w/ iptables.
I'm not sure the history - it's the simplest thing possible, I guess, but doesn't perform very well at the most basic of tasks. To try and solve this, we have the current mess of service meshes, CNI, etc. It's going back to the middleware days of yore - which I think anyone with operational experience should know we really need to avoid. Having multiple layers of network proxies and masquerading between service calls, on an internal network, is just ridiculous and difficult to debug or operate at scale.
Broadly speaking kubernetes saves time by providing an api for existing Linux technology. For example, iptables, ipvs, storage, self-healing, virtual ips, etc via yaml.
Kubernetes was created to give anyone "Google like, production grade" infrastructure without the vendor lock-in of the cloud. And kubernetes on physical machines is where it shines most.
EDIT: it is very easy to complain that kubernetes is complex, but try to do virtual ips with keepalived, corosync, pacemaker to understand how much time kubernetes saves by providing the same capability in a generic way which is easy to use. I have the feeling that people take kubernetes for granted without knowing how much manual work one has to go through in order to build a similar experience using the standard building blocks available.
EDIT 2: straight out of kubernetes home page: "Kubernetes builds upon 15 years of experience running production workloads at Google" that is a lot of saved time just by using kubernetes.
https://twitter.com/copyconstruct/status/1253968037579030532
https://twitter.com/copyconstruct/status/1254273432847446016
If you happen to have the problems Kubernetes solves, it's great. But even Google doesn't have exactly those problems.
(And there is a huge difference between setting up Kubernetes and hosting Kubernetes. If you know how to sit in a chair, you can be in a plane's pilot seat just fine, but piloting is an entirely different skill.)
Still, kubernetes makes a wonderful job abstracting all the hard work of using the underlying tools.
Look how simple is to create a ipvs round-robin by just defining a "kind: Service" object, try to do the same using the bare tools in a way that is synchronized across N computers.
Discovery via DNS is pretty standard too.
Mind to elaborate on what you would like to see in those areas?
Discovery over DNS suffers from the TTL issue right? You’re basically dos’ing yourself with DNS updates?
Also, I would disregard the "moving toward" in that post. Borg has always had mutable container limits, for 10+ years (Large-scale cluster management at Google with Borg, §5.5).
That's not even remotely accurate unless you run Kubernetes on AWS, Azure, GCP.
> straight out of kubernetes home page: "Kubernetes builds upon 15 years of experience running production workloads at Google".
Sounds like you need to read that sentence carefully. They never say they use Kubernetes.
It seems to me you're a Kubernetes fanboy.
I’m being snarky but it is a long time observation of mine. When I see a network config I am more often than not shocked by its needless complexity.
K8 networking is far more complex than typical networking and with less visibility
Communication between pods running across different computers gets encapsulated in another IP packet. This is pretty standard across the industry.
We haven't noticed unusual CPU throttling, though we do have some workloads that turned out to be burstier than expected and had to adjust their CPU limits to match.
Note that when it comes to subtle Linux thread scheduling behavior, your experience will depend on which runtime you use, and if using runc then which version of the Linux kernel your workers run. We weren't affected by the CFS bug introduced in Linux v4.18 because we never ran Kubernetes workloads on a machine with the affected kernel, and if a similar bug occurs in the future it might not affect workloads running within gVisor or Firecracker.
Additionally, Stripe has historically cared more about security than efficiency. This lead to an architecture where services run on dedicated VMs, which naturally strands capacity and reduces the impact of bugs that appear at high utilization and/or high core count.
Often see quota's get exhausted through short bursts that don't show up in metrics that then causes CFS throttling to occur even though it looks like the pod is no where near its limit. Also struggled with application startup requiring far more CPU than at runtime leading to ridiculously slow startup times if you had a low limit.
So far our solution has been to just remove CPU limits, but hoping things will get better.
Removing the limits really improved our latency tail, and so far hasn't resulted in CPU saturation at the node level but your mileage may vary
https://github.com/kubernetes/kubernetes/issues/67577
https://github.com/torvalds/linux/commit/512ac999d2755d2b710...
If I recall correctly you need 4.18+ to get the fix.
https://github.com/aws/amazon-vpc-cni-k8s/blob/master/docs/c...
In the end there is no ideal scenario; boils down to what works best for the use case (or what is the 'least worst' solution). Sometimes it gets you down, but those imperfections can turn a churn job into an interesting one.
Can you please expand on this and if possible point to some references that'd help in case I want to go down this route myself? Thx.
First, background reading:
https://en.wikipedia.org/wiki/6to4
https://en.wikipedia.org/wiki/IPv6_rapid_deployment
The basic idea is you create a SIT tunnel device and assign it an IPv6 /64 composed of two parts:
1. A network prefix between 32 and 56 bits long. This prefix is the same for all machines in the network.
2. A subnet derived from the machine's IPv4 address, minus the netmask.
For example, if your IPv4 addresses are allocated from 192.168.1.0/24 and the machine has 192.168.1.155, then the network prefix should be 56 bits long (64 - (32 - 24)) and the machine's prefix is `xxxx:xxxx:xxxx:xx9B::/64`.
The Linux kernel knows how to wrap the IPv6 with IPv4 so it can route within your local network to any other machine with a similarly configured tunnel device. If you want to send packets to 192.168.1.200 then they get addressed to `xxxx:xxxx:xxxx:xxC8::1` or whatever, they'll transit the IPv4 network like normal, and on arrival the receiving machine's kernel will strip off the IPv4 wrapper and route the IPv6 locally.
How's this useful? Well, if each machine has a /64 prefix then each pod can be allocated an IPv6 within that prefix without coordinating with other machines. Let's say the pod gets `xxxx:xxxx:xxxx:xxC8::aaaa:bbbb:cccc`. Anything with a correctly configured tunnel and that pod IP can route it traffic, no proxy or iptables needed.
I’ve been meaning for a while now to experiment with this same idea in Erlang. I.e., hack up the Erlang runtime to use an IPv6 address as its PID type, such that each Erlang node running on a machine gets its own /64 subnet to hand out; and each Erlang actor-process on that node gets an IP allocated from its node’s /64 range.
This could just be a way of letting Erlang nodes talk to each-other through tunnels. Or it could be a way of having Erlang “VMs” exposed directly to the Internet as their own little machines.
https://www.cni.dev/plugins/main/macvlan/
Basically just do normal ipv4 via your dhcp server rather than an overlay.
-- Edit
For arguments sake, I just set this up:
root@nas:/opt/cni/bin# ./dhcp daemon
cat /etc/cni/net.d/01-macvlan.conf { "name": "mynet", "type": "macvlan", "master": "eno1", "ipam": { "type": "dhcp", "routes": [{ "dst": "192.168.1.0/24"}] } }
PODIP = 192.168.1.181:8096
Works in my browser; so routes correctly.
Got its dhcp from my pihole.
But in production, I'd rather my ability to launch a new pod not be dependent on a DHCP server being reachable and functional. In that case, this particular trick is rather neat, since assignment of IP addresses is fully static/local (without having to agree upfront what range of IPs each node can use for bringing pods online), while retaining the benefit of everything being directly routable. You can now also run a ridiculous amount of pods on a single node.
Though to counter your point, you don't actually need to use an external DHCP server in my example either, you can just define the block you're giving the server via the macvlan/ipvlan plugin, and I presume, again, it works with both IPV4 or IPV6.
So I guess my wider point is, k8s probably doesn't need to replaced to have the networking work how you like.
Like others on this thread I completely agree that k8s networking is over-the-top complex. Still, it is hard to understand how this addresses relatively common use cases like providing intelligent load balancing for clients outside the cluster or secure multitenancy within it.
Solving this requires a step back to think through simplifying networking itself. I thought amazon did the world a favor by removing L2 networking from their model. We need a better set of primitives that abstracts out as much of the low level implementation as possible. The current situation feels like data management before the advent of the relational model.