I suspect that VendorOps and complex tools like kubernetes are favored by complexity merchants which have arisen in the past decade. It looks great on resume and gives tech leaders a false sense of achievement.
Meanwhile Stackoverflow, which is arguably more important than most startups, is chugging along on dedicated machines.
1: https://stackoverflow.blog/2016/02/17/stack-overflow-the-arc...
It seems like the trend in this space is to jump directly to the highest layer of abstraction. Skipping the fundamentals and point to $buzzword tools, libraries and products.
You see glimpses of this in different forms. One is social media threads which ask stuff like “How much of X do I need to know to learn Y?” Where X is a fundamental, possibly very stable technology or field and Y is some tool du jour or straight up a product or service.
CORBA belongs there. And perhaps even the semantic web stuff. Definitely XML!
XML is great - if you don't believe it, then you might be interested in the pains of parsing JSON: http://seriot.ch/projects/parsing_json.html And don't get me started on YAML...
The incremental cost of being on a cloud is totally worth it to me to have managed databases, automatic snapshots, hosted load balancers, plug and play block storage etc.
I want to worry about whether our product is good, not about middle of the night hardware failures, spinning up a new server and having it not work right because there was no incentive to use configuration management with a single box, having a single runaway task cause OOM or full disk and break everything instead of just that task’s VM, fear of restarting a machine that has been up 1000 days etc.
For me, the allure of the cloud is not some FAANG scalability fad when its not required, but automatic OS patching, automatic load balancing, NAT services, centralized, analyzable logs, database services that I don't have to manage, out of the box rolling upgrades to services, managing crypto-keys and secrets, and of course, object storage. (These are all the things we use in our startup)
I'd go on a limb and say dedicated servers may be very viable/cost effective when we reach a certain scale, instead of the other way round. For the moment, the cloud expenses are worth it for my startup.
If you are doing your environment config as code it ultimately shouldn't matter if your target is a dedicated server, a vm or some cloud specific setup configured via an api.
Doesn't really matter which one, the point is that you have one that can build you a new box or bring a misbehaving box back to a know good state by simply running a command.
Sure, the actual pile of Go(o) that drives it can be improved (and it is indeed improving).
That said, the hard truth is that your approach is the correct one, almost all businesses (startups!) overbuild instead of focusing on providing value, identifying the actual niche, etc. (Yes, there's also the demand side of this, as the VC money inflated startups overbuild they want to depend on 3rd parties that can take the load, scale with them, SLAs and whatnot must be advertised.)
Kubernetes is fantastic! Whenever I see a potentially competing business start using Kubernetes, I am immediately relieved, as I know I don't have to worry about them anymore. They are about to disappear in a pile of unsolvable technical problems, trying to fix issues that can't be traced down in a tech stack of unbelievable complexity and working around limitations imposed by a system designed for businesses with traffic two-three orders of magnitude larger. Also, their COGS will be way higher than mine, which means they will need to price their product higher, unless they are just burning through VC money (in which case they are a flash in the pan).
Good stuff all around.
Ansible is imperative, it can work toward a static state, and that's it. If it runs into a problem it throws up its SSH connection, and cries tons of python errors, and gives up forever. Even with its inventory it's far from declarative.
k8s is a bunch of control loops that try to make progress toward their declared state. Failure is not a problem, it'll retry. It'll constantly check its own state, it has a nice API to report it, and thanks to a lot of reified concepts it's harder to have strange clashes between deployed components. (Whereas integrating multiple playbooks with Ansible is ... not trivial.)
And even if k8s puts on too many legacy-ness, there are upcoming slimmer manifestations of the core ideas. (Eg. https://github.com/aurae-runtime/aurae )
And now you’re suggesting that people use a “simpler” non “standard” implementation?
So of course you can implement a minimal complexity solution or you can use something "off the shelf".
k8s is complexity and some of it is definitely unneeded for your situation, but it's also rather general, flexible, well supported, etc.
> you’re suggesting that people use a “simpler” non “standard” implementation?
What I suggested is that if k8s the project/software gets too big to fail, then folks can switch to an alternative. Luckily the k8s concepts and API is open source, so they can be implemented in other projects, and I gave such an example to illustrate that picking k8s is not an Oracle-like vendor lock-in.
I always hear about all the great stuff you get practically for free from cloud providers. Is that stuff actually that easy to set up and use? Any time I tried to set up my LAMP stack on a cloud service it was such a confusing and frightening process that I ended up giving up. I'm wondering if I just need to push a little harder and I'll get to Cloud Heaven
Being able to say "I want a clustered MySQL here with this kind of specs" is much better (time-wise) than doing it on your own. The updates are also nicely wrapped up in the system, so I'm just saying "apply the latest update" rather than manually ensuring failovers/restarts happen in the right order.
So you pay to simplify things, but the gain from that simplification only kicks in with larger projects. If you have something in a single server that you can restart while your users get an error page, it may not be worth it.
The cloud is not easy but damn trying to get cooling and power efficiency of an small server room anywhere near the efficiency levels most big data-center publish is next to impossible as is multi vendor internet connectivity.
With the cloud all of that kind of goes away as it's managed by whatever data-center operator that cloud is running on but what people forget is that that is also true for old fashioned colocation services which is often offering a better cost/value then cloud.
And while it's definitely harder to manage stuff like AWS or Azure because it bleeds a lot of abstractions small scall vpc providers hide from you or that you dont really get with a single home server, it's not hard on the scale of having to run a couple of racks worth of vmware servers with SAN based storage.
With Cloud stuff you have more configuration to do because it is about configuring virtual servers etc. Instead of carrying the PC in a box to your room you must "configure it" to make it available.
DO databases have backups you can configure to your liking, store them on DO Spaces (like S3). DB user management is easy. There's also cache servers for Redis.
You can add a load balancer and connect it to your various web servers.
I think it took me about 30 min to setup 2x web servers, a DB server, cache server, load balancer, a storage server and connect them all as needed using a few simple forms. Can't really beat that.
If you have any more info or opinions then please do share.
Lots of articles around the internet about hardening a Linux server, the ones I've tried take a bit more than 30 min to follow the steps, a lot longer if I'm trying to actually learn and understand what each thing is doing, why it's important, what the underlying vulnerability is, and how I might need to customize some settings for my particular use case.
I'm sure you can find example setup scripts online (configure autoupdates, firewall, applications, etc.), should be a matter of running 'curl $URL' and then 'chmod +x $FILE' and 'bash $FILE'. I didn't need configuration management (I do use my provider's backup service which is important I guess).
Something like this: https://raw.githubusercontent.com/potts99/Linux-Post-Install...
(seen in https://www.reddit.com/r/selfhosted/comments/f18xi2/ubuntu_d... )
Obviously the same can be said for long running VMs, and this can be solved by having a disciplined team, but I think it's generally more likely in an environment with a single long running dedicated machine.
Hetzner has all of this except managed databases.
https://www.hetzner.com/managed-server
The webhosting packages also include 1..unlimited DBs (MySQL and PostgreSQL)
does it have automatic backup and fail over?
> With booked daily backup or the backup included in the type of server, all data is backed up daily and retained for a maximum of 14 days. Recovery of backups (Restore) is possible via the konsoleH administration interface.
But i get the impression that the databases on managed servers are intended for use by apps running on that server, so there isn't really a concept of failover.
A single drive on a single server failing should never cause a production outage.
A lot of configuration issues can be tracked down to self-contained deployments. Does that font file or JRE really need to be installed on the whole server or can you bundle it into your deployment.
Our deployments use on-premise and EC2 targets. The deployment script isn't different, only the IP for the host is.
Now, I will say if I can use S3 for something I 100% will. There is not an on-premise alternative for it with the same feature set.
The point of DevOps is "Cattle, not pets." Put a bullet in your server once a week to find your failure points.
I'd love if you jumped into our Discord/Slack and brought up some of the issues you were seeing so we can at least make the experience better for others using Dokku. Feel free to hit me up there (my nick is `savant`).
Let me preface and say that I'm an application dev with only a working knowledge of Docker. I'm not super skilled at infra and the application I struggled with has peculiar deployment parameters: It's a Python app that at build-time parses several gigs of static HTML crawls to populate a Postgres database that's static after being built. A Flask web app then serves against that database. The HTML parsing evolves fast and so the populated DB data should be bundled as part of the application image(s).
IIRC, I struggled with structuring Dockerfiles when the DB wasn't persistent but instead just another transient part of the app, but it seemed surmountable. The bigger issue seemed to be how to avoid pulling gigs of rarely changed data from S3 for each build when ideally it'd be cached, especially in a way that behaved sanely across DigitalOcean and my local environment. I presume the right Docker image layer caching would address the issue, but I pretty rapidly reached the end of my knowledge and patience.
Dokku's DX does seem great for people doing normal things. :)
Plus who doesn't want to play with the newest, coolest toy on another's dime?
There's definitely an argument to move (some) stuff off cloud later in the journey when flexibility (or dealing with flux/pivoting) becomes less of a primary driver and scale/cost start dominating.
Sure, you can get something working on Hetzner but be prepared to answer a lot more questions.
Agreed on your last point of enterprise asking for this which again is just sad that these business "requirements" dictate how to architect and host your software when another way might be the much better one
Disaster recovery is a very real problem. Which is why I test it regularly. In a scorched-earth scenario, I can be up and running on a secondary (staging) system or a terraformed cloud system in less than an hour. For my business, that is enough.
We went through a growth boom, and like all of them before, it meant there were lots of inexperienced people being handed lots of money and urgent expectations. It’s a recipe for cargo culting and exploitative marketing.
But growth is slowing and money is getting more expensive, so we’ll slow down and start to re-learn the old lessons with exciting new variations. (Like Here: managing bare metal scaling and with containers and orchestration)
And the whole cycle will repeat in the next boom. That’s our industry for now.
A solid dedicated server is 99% of the time far more useful than a crippled VPS on shared hardware but it obviously comes at an increased cost if you dont need all the resources they provide.
Yes - but stupid yearnings to "do what all the cool kids are doing now" are at least as strong in those who would normally be referred to as "managers", vs. "employees".
outcome A) huge microservices/cloud spend goes wrong, well at least you were following best practices, these things happen, what can you do.
outcome B) you went with a Hetzner server and something went wrong, well you are a fool and should have gone with microservices, enjoy looking for a new job.
Thus encouraging managers to choose microservices/cloud. It might not be the right decision for the company, but it's the right decision for the manager, and it's the manager making the decision.
(As does being a consultant wanting an extension and writing software that works, as I found out the hard way.)
There's a disconnect between founders and everyone else.
Founders believe they're going to >10x every year.
Reality is that they're 90% likely to fail.
90% of the time - you're fine failing on whatever.
Some % of the 10% of the time you succeed you're still fine without the cloud - at least for several years of success, plenty of time to switch if ever necessary.
I only have one ansible setup, and it can work both for virtualized servers and physical ones. No difference. The only difference is that virtualized servers need to be set up with terraform first, and physical ones need to be ordered first and their IPs entered into a configuration file (inventory).
Of course, I am also careful to avoid becoming dependent on many other cloud services. For example, I use VpnCloud (https://github.com/dswd/vpncloud) for communication between the servers. As a side benefit, this also gives me the flexibility to switch to any infrastructure provider at any time.
My main point was that while virtualized offerings do have their uses, there is a (huge) gap between a $10/month hobby VPS and a company with exploding-growth B2C business. Most new businesses actually fall into that gap: you do not expect hockey-stick exponential growth in a profitable B2B SaaS. That's where you should question the usual default choice of "use AWS". I care about my COGS and my margins, so I look at this choice very carefully.
You're not a software company, fundamentally you make and sell Sprockets.
The opinions here would be hire a big eng/IT staff to "buy and maintain servers" (PaaS is bad) and then likely "write a bunch of code yourself" (SaaS is bad) or whatever is currently popular here (last thread here it was "Pushing PHP over SFTP without Git, and you're silly if you need more" lol)
But I believe businesses should do One Thing Well and avoid trying to compete (doing it manually) for things outside of their core competency. In this case, I would definitely think Sprocket Masters should not attempt to manage their own hardware and should rely on a provider to handle scaling, security, uptime, compliance, and all the little details. I also think their software should be bog-standard with as little in-house as possible. They're not a software shop and should be writing as little code as possible.
Realistically Widget Masters could run these sites with a rather small staff unless they decided to do it all manually, in which case they'd probably need a lot larger staff.
However, what I also see – and what I think the previous poster was talking about – are businesses where tech is at the core of the business, and there it often makes less sense, and instead of saving time it seems to cost time. There's a reason there are AWS experts: it's not trivial. "Real" servers also aren't trivial, but also not necessarily harder than cloud services.
But those agencies don't want to have to staff for maintaining physical hardware either...
But Sprocket Masters still has to have an expensive Cloud Consultant on retainer simply to respond to emergencies.
If you're going to have someone on staff to deal with the cloud issues, you may as well rent a server instead.
I personally have run Dedicated Servers for our business in the earlier days but as we expanded and scaled, it became a lot easier to go with the Cloud providers to provision various services quickly even though the costs went up. Not to mention it is a lot easier to tell a prospective customer that you use "cloud like AWS" than "Oh we rent these machines in a data center run by some company" (which can actually be better but customers mostly wont get that). Audit, Compliance and others.
For many even lot less than that. I run a small side project[1] that went viral a few times (30K views in 24 hours or so) and it is running on a single core CPU web server and a managed Postgres likewise on a single CPU core. It hasn't even been close to full utilization.
Could I switch some of them to lambda functions? Or switch to ECS? Or switch to some other cloud service du jour? Maybe. But the amount of time I spent writing this comment is already about six month's worth of savings for such a switch. If it's any more difficult than "push button, receive 100% reliably-translated service", it's hard to justify.
Some of this is also because the cloud does provide some other services now that enable this sort of thing. I don't need to run a Kafka cluster for basic messaging, they all come with a message bus. I use that. I use the hosted DB options. I use S3 or equivalent, etc. I find what's left over is almost hardly worth the hassle of trying to jam myself into some other paradigm when I'm paying single-digit dollars a month to just run on an EC2 instance.
It is absolutely the case that not everyone or every project can do this. I'm not stuck to this as an end-goal. When it doesn't work I immediately take appropriate scaling action. I'm not suggesting that you go rearchitect anything based on this. I'm just saying, it's not an option to be despised. It has a lot of flexibility in it, and the billing is quite consistent (or, to put it another way, the fact that if I suddenly have 50x the traffic, my system starts choking and sputtering noticeably rather than simply charging me hundreds of dollars more is a feature to me rather than a bug), and you are generally not stretching yourself to contort into some paradigm that is convenient for some cloud service but may not be convenient for you.
Have you ever had to manage one of those environments?
The thing is, if you want to get some basic more-than-one-person scalability and proper devops then you have to overprovision by a significant factor (possibly voiding your savings).
You're inevitably going to end up with a bespoke solution, which means new joiners will have a harder time getting mastery of the system and significant people leaving the company will bring their intimate knowledge of your infrastructure with them. You're back to pets instead of cattles. Some servers are special, after a while automation means "a lot of glue shell scripts here and there" and an OS upgrade means either half infra is KO for a while or you don't do OS upgrades at all.
And in the fortunate case you need to scale up... You might find unpleasant surprises.
And don't ever get me started on the networking side. Unless you're renting your whole rack and placing your own networking hardware, you get what you get. Which could be very poor in either functionalities or performances... Assuming you're not doing anything fancy.
If you want 100.0000% uptime, sure. But you don't usually. The companies that want that kind of uptime normally has teams dedicated to it anyway.
And scaling works well on bare-metal too if you scale vertically - have you any idea the amount of power and throughput you can get from a single server?
It's concerning to keep hearing about "scaling" when the speaker means "horizontal scaling".
If you requirements are "scaling", then vertical scaling will take you far.
If your requirements are "horizontal scaling on demand", then, sure, cloud providers will help there. But, TBH, few places need that sort of scaling.
I'm not saying 100% uptime on bare metal is cheap, I'm saying 100% uptime is frequently not needed.
Because the industry is full of people who are chasing trends and keywords and to which the most important thing is to add those keywords to the CVs.
IME aiming for scalability is exceedingly wrong for most services/applications/whatever. And usually you pay so much overhead for the "scalable" solutions, that you then need to scale to make up for it.
I really doubt this. The one theme you see on almost every AWS proponent is some high amount of delusion about what guarantees AWS actually provides you.
Yeah, sure. Nobody gets fired for buying from IBM.
Well, if large companies had any competence in decision-making, they would be unbeatable and it would be hopeless to work on anything else at all. So, yeah, that's a public good.
Source: almost thirty years of ops.
So far I haven’t noticed that if I spend more for the company that I also get paid more.
getting the equivalent reliability with irons is a lot more expensive than renting "two dedicated servers" - now you might be fine with one server and a backup solution, and that's fair. but a sysad to create all that, even on a short contract for the initial setup and no maintenance, is going to go well beyond the cloud price difference, especially if there's a database in the mix and you care about that data.
Today, cloud is similar - the time to market is quicker as there are less moving parts. When the economy tanks and growth slows, the beancounters come in and do their thing.
It happens every time.
This only makes AWS richer at the expense of companies and cloud teams
It is trivial to provision and unprovision an EC2 instance automatically, within seconds if your deployment needs to scale up or scale down. That's what makes it fundamentally different from a bare metal server.
Now, I'm not denying that it might still be more cost effective when compared to AWS to provision a few more dedicated servers than you'll need, but when you have really unpredictable workloads it's not easy to keep up.
if you are spinning up and shutting down VMs to meet demand curve - something is seriously wrong with your architecture.
Did it ever occur to you, how come stackoverflow uses ~8 dedicated servers to serve entire world and doesn't need to spin up and shutdown VMs to meet global customer demand?
--
When planning compute infrastructure, it is important to go back to basics and not fall for cloud vendor's propaganda
You said it yourself. Not all applications need to serve entire world, therefore demand will be lower when people go to sleep.
Even with global applications there are regulations that require you to host applications and data in specific regions. Imagine an application that is used by customers in Europe, Australia, and the US. They need to be served from regional data centres, and one cluster will be mostly sleeping when the other is running (because of timezones). With dedicated servers you would waste 60-70% of your resources, while using something like EC2/Fargate you can scale down to almost 0 when your region is sleeping.
There is a method to the madness, here it is called "job-security-driven development".
Because it's a threat to their jobs.