Running Istio In Production
engineering.hellofresh.com
engineering.hellofresh.com
"N Microservices for $BUSINESS_DOMAIN ? Crazy!"
Can't we just taka it on good faith that the engineers who do these blog posts have at least a modicum of competence and therefore their solution has _some_ merit (that is worth discussing) rather than just "Micro-services bad" dismissals?
<sideshow bob>Yes I realise this can also be taken as a shallow dismissal of the shallow dismissals</sideshow bob>
On a technical resource, no. In a technical article, technical points need to be explained. Decisions need to be checked.
Unless your argument is "it cannot be bad because it must have been checked by many people", we are totally in the position to provide any sort of constructive feedback.
My comments were aimed at the shallow dismissals at the level to which I gave an example (and examples of which can be found in the comments on this post).
If a response to these articles is along the lines of "OMG they are using microservices, BAD!" then that is not constructive or useful unless you are trying to invoke Cunningham's Law (Or whatever the one was about getting help on a linux mailing list by saying something cannot be done).
_However_ if you want to have a constructive discussion about if isio is even the right choice here then have at it if you think there is enough info to go into that. (Personally I don't think there is enough info in the blog post to have that discussion... since that is not the point of the blog).
Again, my comment was mainly aimed at the shallow 1-sentence dismissal comments that crop up on these posts.
Lets solve this problem by introducing another level of indirection and not solving the root cause(?).
At this point i really believe that software architects who don't code, don't belong to this industry. If the implementers and operators are suffering, there should be a feedback channel.
It’s a great way to get a uniform metrics and troubleshooting experience.
Most important though, especially if you have strategically compiled binaries as microservices, a service mesh lets you roll out improved routing logic easily because it’s all abstracted away, not encoded in client libraries in X number of languages. Same for tracing. Wanna add a field to all traces being generated? Go recompile and redeploy a 100 services... or change one service mesh config.
Other than that, envoy can usually withstand much more traffic than the service it overlays, so you can use it to provide DoS protection in depth, by limiting on service proxies everywhere. Saved us a couple outage escalations already.
services-to-service peer-to-peer communication creates problems on its own. and i dont even know what are the benefits. it does not guarantee you anything per-se. all the redundancy and bandwith improvements need to be... coded... like with any other approach.
In the end Helm just introduced more problems than it solved. Rather than applying configs haphazardly and relying on 3rd party services, it was ultimately much simpler to just download the configs for whatever service was needed (nginx ingress controller for me) and committing them to source control.
My biggest take away from the k8s community is that lots of people write terrible documentation and other people write blog posts and SO answers without actually understanding how kubernetes works under the hood.
There is still an ocean of depth in regards to k8s I don’t know yet, but I feel a lot stronger about getting the intermediate basics. I’m at a point where I’m being productive again.
Learning how Kubernetes works is much easier if you first get a firm grasp of the basics and then start bolting stuff on like istio knative and all the other cool stickers “modern architects” wet dream about.
Edit: punctuation
It also enables lots of functionality when it comes to CD. It enables things like canary releases, testing, rollback etc etc in a simpler way by keeping it in the Kubernetes space (and not relying on slow external LBs).
You probably won't know why you need Istio till you need it.
Looking forward to upcoming parts!
Even if you only sipped the microservices kool-aid, I can easily see dozens of services.
Granted, it seems reasonable that a handful of monoliths could get the job done. Without having worked at HelloFresh, I'm inclined to think there's more to the story that we don't know. Maybe there's a good reason to have as many services as they do.
Next time, I want to ask them how they keep track of all those services!
edit: replaced an incorrectly-used idiom
Computers are awesome at automating things, that goes for dev tooling as well.
If you’ve touched the ITSM space you’re used to managing and maintaining many thousand of assets. A few hundred microservices is nothing, really.
My team use what you could call a simplified CMDB (configuration management database) which is cross referenced against the service discovery.
The cmdb keep info about every service, such as persistent data-sources, vm’s etc, but most important - relationships: domain, team, services and resources.
A microservice is basically a ”ci” (configuration item) with a managed lifecycle.
Log shipping is what we do from thousands of servers already (you should at least!), adding a shipper for a few 100 containers on a set of hosts is no big deal.
Fluent(d/bit) -> some kind of elastic? There are a few resonable patterns available that works and scales pretty well.
Failures and issues with the actual code - well I might have been lucky... DDD with somewhat senior devs where no spaghetti action takes place. The tooling we keep usually seem to pinpoint issues fairly well.
We’re on the scale of roughly 40 devs and my team of 3 support them with tooling that handles service lifecycle and operational stuff.
It let’s us be pretty fluent with what and how teams build and iterate stuff. I guess it requires a certain scale and experience though.
Or maybe we're in a tech bubble and if you want engineers, you have to acquiesce to their demands of working with the latest shiny tools while they create mountains of technical debt, because if you don't, they'll just go to another startup that allows that behavior, or they'll go twiddle their thumbs at a FAANG while banking $300k+ total comp for their 5-years of experience.
I'm not trying to say one could replicate this business with a few scripts, but it sure as hell doesn't need a service mesh that looks like a mutated SARS-CoV3.
All of those systems need each other's information. The shop needs to know if a product is available, and what price it is. Customer Service needs to know how it was sent and the tracking code etc,...
We've built hundreds of microservices to manage this. This is definitely not something that is easy in a monolith. (We came from a monolithic architecture, we're really happy we're now using microservices).
We're now also rolling out Istio on GKE. So who knows, they might be on to something at Hello Fresh.
But, at some point, companies that want to keep talented tech people need to let them go build what they want to build. Maybe those things are over-architected for what the company needs right now, but it's tough to say if that's a bigger risk than losing talented tech people.
Obviously part of the answer is "Because I need a tech unicorn valuation". But maybe your marketing to investors shouldn't be what you're actually basing your hiring decisions on.
It's amazingly simple to configure, the docs are pretty OK, and the benefits seems huge. Mutual TLS within minutes? BAM, done. Don't worry about cert rotation, Citatel does that for you.
Block all egress, and whitelist what you need? Like, that's a killer feature! Also, the inability to do 90-10 canary releases with plain K8s baffles me. With Istio? Simple...
I don't know, I'm sure I'll find the pains of Istio in the coming months, but in my dev cluster, it looks amazing.
This quietly disincentivized them to lean into any OSS stacks that will take away from this future revenue stream.