New setup for 2020
changelog.com
changelog.com
This is (part) of what keeps me in the stone age. You are provisioning load balancers and DNS - but just one step removed through k8s
And my prior is that we need to understand that, be aaare of it, have a model of what is going on to help develop and debug.
And so it feels a bit like "magic abstraction". And then to peek through the abstraction you suddenly not only need to know about DNS and which machine is running bind, but also how kubernetes internally stores it's DNS config and how it spits that out and what version they changed it with.
In other words you have to become expect in two things to debug it.
And maybe it's worth it - but I struggle to see why it's not simpler to keep my install scripts going.
(OK I guess I am writing my own answer - but surely the point is what is the simplest level of thin install scripts needed to deploy containers?)
Does this imply there is a cloud abstract layer that should come (assuming all providers can put aside commercial interests etc)
And is k8s the simplest possible abstraction? And if not - what is?
In the unhappy path, sure, you need someone who knows how to debug networking issues, and in some cases it’s going to be harder to debug because of the layers of indirection. But the total amount of toil is significantly reduced.
A bad abstraction doesn’t carry its weight in complexity. A good abstraction allows you to ignore the lower levels most of the time without missing something important; I’d put k8s firmly in the latter category.
If it's a sufficiently robust abstraction, you don't, you just learn the abstraction. Kubernetes has reached that point for many folks.
I no longer have a detailed mental model of how my compiler or LLVM works, I just trust that it does. When was the last time you needed to (or were capable of) debugging a bug in your compiler? A couple of human generations of work went into making that happen.
Note that it turns out compiling code well, or making a reliable orchestration system, is an enormously complex problem. At some point, the complexity outstrips the ability of even generalists in the field to keep up, yet the systems keep getting more reliable.
So in these types of cases, you can either do it yourself poorly (you're an amateur), do it yourself well (congrats, you've become an expert), or delegate.
This isn't really limited to computing. I delegate maintenance on my car to a mechanic, while I'm pretty sure a generation ago, everybody (in the US) changed their own oil and understood how the carb worked. Times change.
But yes, mostly I am not an army.
Their problem is being easy to repair.
Most of the regular people can take their cars to a shop that will have the really basic in stock (ie. oils and filters) and can order most of the rest from a distributor to have delivered same or next day for most cars. You will only take longer to get specialized parts or parts for unusual cars. Also most people can work around not having a car for a few days even if it is a hassle.
While army base does have specialized mechanics to properly fix the trucks, if one break in the middle of a mission after engagement the team in that truck need to be able to patch it up and get it going on the field.
So the army need trucks that people that are primarily soldiers and not mechanics can patch up "easily" in the field without tons of specialized training, without access to special or weird tools and minimal to no access to parts.
In many cases, it's a good tradeoff, because you can now use standard tooling on everything.
Just like it's cheaper to ship an entire (physical) shipping container that's half-full than to ship the same stuff loosely. Or why companies will send you two separate letters on the same day with a small note that this is more efficient for them than collating them.
I assume that k8s also makes it much easier to move to a different cloud provider if you're unhappy with one (or the new one offers better pricing). Instead of rewriting your bespoke scripts that only you understand, anyone familiar with the technology will know which modules to swap to make it work with the new provider.
For someone that is known as the King of Bash (self-proclaimed) - https://speakerdeck.com/gerhardlazu/how-to-write-good-bash-c... - and after a decade of Puppet, Chef, Ansible and oh wow that sweet bash https://github.com/gerhard/deliver - even if all my workstations and work servers (yup, all running k3s) are provisioned with Make (bash++), I still think that K8S is the better approach to running production infrastructure. The advantage to using simple and well-defined components (e.g. external-dns, ingress-nginx, prometheus-operator etc.) that adhere to a universal API, and are maintained by many smart people all around the world, is a better proposition than scripting in my opinion.
At the end of the day, I'm in it for the shared mindset, great conversations and a genuine desire to do better, which I have not seen before K8S & the wider CNCF. I will go on a limb here and assume that I love scripting just as much as you do, but go beyond this aspect and you will discover that it's more to it than "thin install scripts that deploy containers" (which are not just glorified jails or unikernels).
I love the idea of using Kubernetes, it sounds amazing initially, but then every single article I read about it turns into some epic blog post that leaves me worried that the whole house of cards could easily come crashing down.
Maybe in the future the abstraction will become rock solid, easy to install and manage and ‘just work’, but it doesn’t feel like that to me now. There’s too many ‘we had to come up with / use hack X to integrate it with software Y’.
Until then, if you are on a small team with a small budget, I reckon keeping it simple is the better approach. Standard OSs with some bash scripts for provisioning, build and deploy. Even if it’s more manual work, and takes a bit longer, having an understanding of the platform you are building on is crucial.
BTW if you are looking for a description of a way todo this sort of thing ‘the boring way’:
Robust NodeJS Deployment Architecture
https://blog.markjgsmith.com/2020/11/13/robust-nodejs-deploy...
>>> >>> with the advent of the first programming languages you no longer had to think in terms of registers, loading operands from memory, storing the results back, spilling registers to the stack.
>>> You _are_ handling registers and memory spills, but just one step removed through the use of C.
The analogy may not be perfect, but I think it makes obvious some of the things also mentioned in sibling comments: it's all about habit, maturity and thus trust.
If you trust your tools are working correctly, if you know how to deal with their well known quirks, you'll just rebase on top a layer and hopefully boost your productivity and tackle more complex problems more easily.
Maturity is important because if today you're more likely to blame yourself than the compiler when your stuff doesn't works, it's just because you're lucky to work with mature and popular toolchains. (And I'm not only talking about the past when compilers where new and unproven, it still happens today on some niche embedded toolchains).
So, yes, it's indeed rational to be wary of new unproven abstraction layers as they could bring more pain than help.
It's hard to judge when that line is crossed though.
I personally like to know how stuff works under the hood anyways. I find it useful in practice and it gives me confidence in using the higher layers, when they make sense, or stay with the lower level layers, when they make sense.
Occasionally I still write some assembly. But most of the time, for most of the stuff, it just makes more sense to use a higher level programming language.
I see k8s in a similar way. We have operating systems, programming languages, etc; all sorts of abstractions that help us separate concerns and have specialists dealing with the nitty gritty details of some stuff, so that everybody else can be specialized in something else (just like in real life)
But this hard-line-to-underlying reality is unlikely to exist looking at how this months k8s will configure last years AWS route53.
It just works 80% of the time is a disaster, 98% of the time might be bearable. is it above 99%?
When compilers created buggy code every other day, when memory allocators were unreliable because memory fragmentation would make it likely for new allocations to fail in the lifetime of a normal program execution, etc, it would be, as you said "a disaster if it works <95%" of the times.
Did k8s reach 99%? The jury is still out. Probably not yet, but in principle I don't see anything wrong pursuing that path just because it's "another layer". We use abstraction layers all the time; they allowed progress (Yes often we went too far)
Some things I noticed:
* The internal network is not private. But people don't realise it. You share a /16 with other Linodes. So many open databases, file shares and other services in there.
* Block storage performance is really poor, around 100 iops. Same as a SATA disk from 10 years ago.
* No proper snapshot / image functionality.
* Linode Kubeternetes Engine was based on Debian Oldstable when it launched.
* Excessive CPU steal, even on dedicated cores. 25% CPU steal is considered normal. Over 50% happens a lot.
* Problems with their hosts. I can only guess what the reason is but 4 to 8 hours of unannounced downtime of a VM happend to me 6 times in the past 2 years.
Yes, support is friendly. But my international phone bill is huge because the fastest way to get them to do something is to call.
We are moving everything away. Most of our servers are with another provider already. And we haven't had any similar issues there. I've never called them!
And I forgot to mention the connectivity issues at Linode. When the whole London datacenter was unreachable for 2 hours we lost some customers.
But if I'm calling you, you have most likely already failed. If I'm calling for information, your documentation has failed to make that information accessible (it either wasn't documented, or not easy enough to find). If I'm calling to resolve an issue (technical or billing), it would have been a lot better if it didn't happen in the first place.
When failures happen, it's always a series of unfortunate incidents. When we've hit issues with Linode, we reached out and worked on what we can improve in our changelog.com setup, and discussed about the improvements that we can expect on the Linode end. Our common interest is a more resilient system, which requires a healthy collaboration, and Linode has been a great technology partner for us. Expect to see these write-ups on changelog.com as soon as these improvements have shipped, and we have hard data to support the claims ; )
I'm sorry to hear that things have not been as smooth for you on Linode. I hope that you will find an infra provider that you will be able to rely on and work with as we do. Not all collaborations will work out, and that's OK. It's also OK to be annoyed, fed up with the way things are and look for something different, something more suitable for you. My only ask is that you share your migration story with the changelog.com community. That is something that I would want to hear about.
Despite all that, management decided to stay with Linode for the following reasons:
* Change is hard, and "better the devil you know" mentality.
* The instance pricing looks cheap compared to AWS. e.g. c5.xlarge ($124) vs 8GB-Balanced ($40) that Linode charges. In reality it isn't so cheap because it's poor oversold technology.
* AWS/GCP have exorbitant bandwidth pricing. Linode bandwidth is very generous, as its pooled across all servers in the account.
* Having someone to pick up the phone 24/7 when there's a problem is a big plus in theory. However, it's much better not to need to call in the first place because things just work.
* Migrating providers can be an expensive and time-consuming endeavour.
* Technical debt, interdependencies, manually configured snowflake servers and infrastructure, no documentation, etc. makes changes risky.
* Not enough DevOps on the team, and too many fires to put out, and shiny features to ship means cloud provider migration is low on the priority list.
A bunch of bare metal hosts run on Scaleway / Online, and different VMs & managed services run in Digital Ocean, Linode, AWS & GCP. I sometimes spin the odd bare metal instance on Equinix Metal (former Packet).
A diverse fleet means that there's always something new to learn and try out. A single large host would make me anxious, as no internet provider or power grid is 100% reliable and available. Also, software upgrades sometimes fail, and things get messed up all the time, which is when I find it most efficient to just start from scratch. A single host makes that less convenient.
Every approach has its pros and cons, which is why my main workstation is a 20 Xeon W with 64GB RAM & 1TB NVME : ). Yes, there is a backup workstation which doubles up as a mobile one meaning that it can work without power or hard internet for almost a day. Options are good ; )
Is it a configuration issue on their side or do the LKE volumes are really limited to 6MB/s on linode?
How can you be happy with this for production??
We have mostly sequential reads & writes (mp3 files) that peak at 50MB/s, then rely on CDN caching (Fastly makes us happy in this respect).
CDN caching is something that we are currently improving, which will make things quicker and more reliable.
The focus is on reality vs the ideal, and the path that we are taking to improving not just changelog.com, but also our sponsors' products. No managed K8S or IaaS is perfect, but we enjoy the Linode partnership & collaboration ;)
Then again, depends on what you're doing.
> It’s worth noting that we don’t really need what we have around Kubernetes. This is for fun, to some degree. One, we love Linode, they’re a great partner… Two, we love you, Gerhard, and all the work you’ve done here… We don’t really need this setup. One, it’s about learning ourselves, but then also sharing that. Obviously, Changelog.com is open source, so if you’re curious how this is implemented, you can look in our codebase. But beyond that, I think it’s important to remind our audience that we don’t really need this; it’s fun to have, and actually a worthwhile investment for us, because this does cost us money (Gerhard does not work for free), and it’s part of this desire to learn for ourselves, and then also to share it with everyone else… So that’s fun. It’s fun to do.
As for persistent volumes, might be better to just offload Postgres to a managed DB service and downsize the K8S instances, or use something like CockroachDB which is natively distributed and can make use of local volumes instead.
Other static files such as css, js, txt make sense to remain bundled with the app image, which is stateless and a prime candidate for horizontal scaling. Also, CDN caching makes small static files that change infrequently a non issue, regardless of their origin.
The managed Postgres service from Linode's 2021 roadmap is definitely something that we are looking forward to, but the simplest thing might be to provision Postgres with local volumes instead. We are already using a replicated Postgres via the Crunchy PostgreSQL Operator, so I'm looking forward to trying this approach out first.
CockroachDB is on my list of cool tech to experiment with, but that will use an innovation token, and we only have a few left for 2021, so I don't want to spend them all at once.
If you're using an operator then local volumes is a good middleground if it automates the replication already. CockroachDB also has a kubernetes operator although it's only for GKE currently. There are also other options like YugabyteDB which is another cloud-native postgres-compatible DB.
I've just added YugabyteDB to my explore list.
Yes, we could have mitigated that entirely with CDN stale caching, but it was good to see what happens today, and then iterate towards better Fastly integration.
Concourse worked well for us, we didn't have any issues that were being enough to remember. You may be interested in this screenshot that captured the changelog.com pipeline from 2017: https://pipeline.gerhard.io/images/small-oss.png
I missed the simple Concourse pipeline view at first, but CircleCI improved by leaps and bounds in 2020, and the new Circle pipeline view equivalent is even better (compared to Concourse, clicking on jobs always works): https://app.circleci.com/pipelines/github/thechangelog/chang...
The Circle feature which I didn't expect to like as much as I do today, is the dashboard view (list of all pipeline/workflow runs). This is something that Concourse is still missing: https://app.circleci.com/pipelines/github/thechangelog
My favourite Circle 2020 feature is the Insights: https://app.circleci.com/insights/github/thechangelog/change.... Yup, we were one of the first ones to ask for it in 2019.
In 2021, I expect us to spend one migration credit on GitHub Actions, as a Circle replacement. Argo comes as a second close, but that requires an innovation credit which is more precious to us. Because we are already using GitHub Actions for some automation, it would make sense to consolidate, and also leverage the GitHub Container Registry, as a migration from Docker Hub. Watch https://github.com/thechangelog/changelog.com to see what happens : )
I couldn't use the Circle CI links, they're auth gated.
The code in this repo tells the truth about what it is, and even shows how it works: https://github.com/thechangelog/changelog.com
vendor-managed DBs limit plugins + versions, excited to see this space advance
There’s a very active news feed with submissions, commenting, newsletter subscriptions and management, a blog, episode requests, live streams, etc.
Check out the source to see what all the app does:
Without some additional insight that I don't have, it does seem that this is an enormously over-engineered "solution" for a website -- I've NFI why it requires 99.99% uptime!
Perhaps they look it as a goal or challenge, an opportunity to showcase their knowledge and skills to potential customers, or, hell, maybe they just enjoy that kind of thing? If that's the case, I completely understand and can even relate (my home network is a textbook example of an "over-engineered solution": close to a dozen "enterprise-class" servers in the basement, ~35 various subnets, VMware Enterprise Plus clusters, BGP for anycast, and so on).
AFAICT, though, this is just some developers running a blog and podcasts aimed at other developers? I mean, we're not exactly talking about a "mission critical" web site that's going to result in death and destruction the next time it goes down or Linode shits itself, right?
Or am I missing something?
--
EDIT: I've read through the rest of the comments now ...
> This is for fun, to some degree ... We don’t really need this setup. One, it’s about learning ourselves, but then also sharing that ... It’s fun to do.
... and I completely understand!
1. We're great engineers because we can set and maintain such an impressive set-up.
2. We're terrible engineers because all of this could probably done one server for dynamic content + S3 for dynamic content. Or not even S3, maybe just some Cloudflare or Akamai caching.
Of course, like the posters above, I could be missing something due to my outsider/consumer view of changelog.
The devil is in the details, there is more to it than dynamic & static content, we are using Fastly, otherwise we couldn't serve all the traffic that we do.
The best part is that it's all public - https://github.com/thechangelog/changelog.com - and we welcome contributions, especially those that simplify our setup without compromising on resiliency and availability. I'm looking forward to yours ; )