Almost every infrastructure decision I endorse or regret
cep.dev
cep.dev
Every so often I price out RDS to replace our colocated SQL Server cluster and it's so unrealistically expensive that I just have to laugh. It's absurdly far beyond what I'd be willing to pay. The markup is enough to pay for the colocation rack, the AWS Direct Connects, the servers, the SAN, the SQL Server licenses, the maintenance contracts, and a full-time in-house DBA.
https://calculator.aws/#/estimate?id=48b0bab00fe90c5e6de68d0...
Total 12 months cost: 547,441.85 USD
Once you get past the point where the markup can pay for one or more full-time employees, I think you should consider doing that instead of blindly paying more and more to scale RDS up. You're REALLY paying for it with RDS. At least re-evaluate the choices you made as a fledgling startup once you reach the scale where you're paying AWS "full time engineer" amounts of money.
Assuming you have any. You might not, because of AWS.
Hard disagree. An r6i.12xl Multi-AZ with 7500 IOPS / 500 GiB io1 books at $10K/month on its own. Add a read replica, even Single-AZ at a smaller size, and you’re half that again. And this is without the infra required to run a load balancer / connection pooler.
I don’t know what your definition of “large” is, but the described would be adequate at best at the ~100K QPS level.
RDS is expensive as hell, because they know most people don’t want to take the time to read docs and understand how to implement a solid backup strategy. That, and they’ve somehow convinced everyone that you don’t have to tune RDS.
If you don't have a reserved instance, then you're giving up potentially a 50% discount on on-demand pricing.
An r6i.12xl is a huge instance.
There are other equivalents in the range of instances available (and you can change them as required, with downtime).
For MySQL and Postgres, RDS stripes across four volumes once you hit 400 GiB. Doesn't matter the type.
The latency variation on gp3 is abysmal [0], and the average [1] isn't great either. It's probably fine if you have low demands, or if your working set fits into memory and you can risk the performance hit when you get an uncached query.
12K IOPS sounds nice until you add latency into it. If you have 2 msec latency, then (ignoring various other overheads, and kernel or EBS command merging) the maximum a single thread can accomplish in one second is (1000 msec / 1 sec / 2 msec) = 500 I/O. Depending on your needs that may be fine, of course.
> If you don't have a reserved instance, then you're giving up potentially a 50% discount on on-demand pricing.
True, of course. Large customers also don't pay retail.
> An r6i.12xl is a huge instance.
I mean, it goes well past that to .32xl, so I wouldn't say it's huge. I work with DBs with 1 TiB of RAM, and I'm positive there are people here who think those are toys. The original comment I replied to said, "large SaaS," and a .12xl, as I said, would be roughly adequate for ~100K QPS, assuming no absurdly bad queries.
[0]: https://www.percona.com/blog/performance-of-various-ebs-stor...
[1]: https://silashansen.medium.com/looking-into-the-new-ebs-gp3-...
After initial setup, managing equivalent of $5k/m RDS is not full time job. If you add to this, that wages differ a lot around the world, $5k can take you very, very far in terms of paying someone.
The query times are incredible.
OTOH, as noted, EBS does not perform as well as native NVMe and is hilariously expensive if you try. And quite a few use cases are just fine on plain old NVMe.
Sure, maybe if you are some sort of SaaS with a need for a small single DB, that also needs to be resilient, backed up, rock solid bulletproof.. it makes sense? But how many cases are there of this? If its so fundamental to your product and needs such uptime & redundancy, what are the odds its also reasonably small?
Most software startups these days? The blog post is about work done at a startup after all. By the time your db is big enough to cost an unreasonable amount on RDS, you’re likely a big enough team to have options. If you’re a small startup, saving a couple hundred bucks a month by self managing your database is rarely a good choice. There’re more valuable things to work on.
By the time your db is big enough to cost an unreasonable amount on RDS, you've likely got so much momentum that getting off is nearly impossible as you bleed cash.
You can buy a used server and find colocation space and still be pennies on the dollar for even the smallest database. If you're doing more than prototyping, you're probably wasting money.
Optionality and flexibility are extremely valuable, and that is why cloud compute continues to be popular, especially for rapidly/burstily growing businesses like startups.
If you want a different opportunity cost, get people with different experience. If RDS is objectively expensive, objectively slow, but subjectively easy, change the subject.
I don't think that's accurate. I've self-managed databases, and I still think that RDS is compelling for small engineering teams.
There's a lot to get right when managing a database, and it's easy to screw something up. Perhaps none of the individual parts are super-complicated, but the cost of failure is high. Outsourcing that cost to AWS is pretty compelling.
At a certain team size, you'll end up with a section of the team that's dedicated to these sorts of careful processes. But the first place these issues come up is with the database, and if you can put off that bit of organizational scaling until later, then that's a great path to choose.
Its about tradeoffs, and some tradeoffs are often more applicable than others - getting a ping at 7am on a Sunday because your ec2 instance filled it's drive up with logs and your log rotation script failed because it didn't have a long enough retey is a problem I'm happy to outsource when I should be focusing on the actual app.
If the cost of a hosted DB is going to sink the company, then of course, I will figure it out and run it myself. But it’s not, for most startups. And therefore that knowledge isn’t providing much leverage.
Starting an AI company with deep expertise in training models - that is an example of knowledge providing huge leverage. DB tech is not in this bucket for most businesses.
I’ve been the guy managing a critical self-hosted database in a small team, and it’s such a distraction from focusing on the actual core product.
To me, the cost of RDS covers tons of risks and time sinks: having to document the db server setup so I’m not the only one on the team who actually knows how to operate it, setting up monitoring, foolproof backups so I don’t need to worry that they’re silently failing because a volume is full and I misconfigured the monitoring, PITR for when someone ships a bad migration, one click HA so the database itself is very unlikely to wake me at 3am, blue/green deploys to make major version upgrades totally painless, never having to think about hardware failures or borked dist-upgrades, and so on.
Each of those is ultimately either undifferentiated work to develop in-house RDS features that could have been better spent on product, or a risk of significant data loss, downtime, or firefighting. RDS looks like a pretty good deal, up to a point.
RDS is a great choice, for prototyping and only for production if you know what you're doing when setting it up.
FWIW, this is common in all cloud deployments, people assume that running something "severless" is a magical silver bullet.
Maybe it takes you a month the first time around and a week the 10th time around. First product suffers, the other products not so much. Now it just takes a week of your time and does not require you to pay large AWS fees, which means you are not bleeding money
I like to set up scrappy products that do not rack up large monthly fees. This means I can let them run unprofitable for longer and I don't have to seek an investor early, which would light up a large fire under everyone's butts and start influencing timelines because now they have the money and want a return asap.
I'll launch a week later - no biggie usually. I could have come up with the idea a month later, so I'm still 3 weeks early ;)
It doesn't work for all projects, obviously, but I've seen plenty of SaaS start out with a shopping spree, then pay monthly fees and purchase licenses for stuff that they could have set up for free if they put some (usually not a lot) effort into it. When times get rough, the shorter runway becomes a hard fact of life. Maybe they wouldn't have needed a VC and could have bootstrapped and also survived for longer.
While I generally agree as far as initial setup time goes, I favor RDS because I can forget about it, whereas the hand rolled version demands ongoing maintenance, and incurs a nonzero chance of simple mistakes that, if made, could result in a 100% dataloss unrecoverable scenario.
I’m also mostly talking about typical, funded startups here, as opposed to indie/solo devs. If you’re flying solo launching a tiny proof of concept that may only ever have a few users, by all means run it yourself if you’d like, but if you’ve raised money to grow faster and are paying employees to iterate rapidly searching for PMF…just pay for RDS and make sure as much time as possible is spent on product features that provide actual business value. It starts at like $15/month. The cost of simply not being laser-focused on product is far greater.
Databases are not particularly difficult to migrate between machines. Of all the cloud services to migrate, they might actually be the easiest, since the databases don't have different API's that need to be rewritten for, and database replication is a well-established thing.
Getting off is quite the opposite of nearly impossible.
Also, Aurora gives you the block level cluster that you can't deploy on your own - it's way easier to work with than the usual replication.
I mean, I'm looking after a 4 instance Aurora cluster which is great feature wise, is slightly overprovisioned for special events, and is more likely to shrink than grow 2x in the next decade. If we start experiencing any issues, there's lots of optimisations that can be still gained from better caching and that work will be cheaper than the instance size upgrade.
There’s still a defined cost to swapping your DB code over to a different backend. At the point where it becomes uneconomical, you’re also at a scale you can afford rewriting a module.
That’s why we have things like “hexagonal architecture”, which focus on isolating the storage protocol from the code. There’s an art to designing such that your prototype can scale with only minor rework — but that’s why we have senior engineers.
No one runs their own electricity supply (well until recently with renewables/storage), they buy it as a service, up to a pretty high scale before it becomes more economic to invest the capex and opex to run your own.
These questions always sound flawed to me. It's like asking won't I regret moving to California and paying high taxes once I start making millions of dollars? Maybe? But that's an amazing problem to have and one that I may be much better equipped to solve.
If you are small, RDS is much cheaper, and many company killing events, such as not testing your backups are solved. If you are big and you can afford a 60K/yr RDS bill than you can make changes to move on-prem. Or you can open up excel and do the math if your margins are meaningfully affected by moving on-prem.
On some level, AWS/GCP/California relies on you doing this calculation for the things that you can do it on easily (the savings of moving away), while not doing this calculation on things where it's hard to do (new development). That way, you can pretend that your new features are a lot more valuable than the $Xk/year you will save by moving your infra.
Yes, I've done the math. The piece you are missing is, saving money on infra will bring in $0 new dollars. There is a floor to how much money I can save. There is no ceiling to how much money the right feature can bring in. Penny pinching on infra, especially when the amount of money is saved is less than the cost of an engineer is almost always a waste of time while you are growing a company. If you are at the point where you are wasting 1x,2x,3x of an engineers salary of superflous infrastructure - then congratulations you have survived the great filter for 99% of startups.
>That way, you can pretend that your new features are a lot more valuable than the $Xk/year you will save by moving your infra.
Finding product market fit is 1000x harder than moving from RDS to On-prem. If you haven't solved PMF, then no amount of $Xk/year in savings will save you from having to shut down your company.
The thing is, most features, particularly later in the life of a company, don't have an easy-to-measure revenue impact, and I suspect that many features are actually worth $0 of revenue. However, they cost money to implement (both in engineering time and infra), making them very much net negative value propositions. This is why Facebook and Google can cut tons of staff and lose nothing off their revenue number.
Also, there's a bit of a gambling mentality here which is that a feature could be worth effectively infinite revenue (ie it could be the thing that gives you PMF), so it's always worth doing over things with known, bounded impact on your bottom line. However, improving your efficiency gives you more cracks at finding good features before you run out of money.
So moving to/from Aurora/RDS/own EC2/on-prem should be a matter of networking and changing connection strings in the clients.
Your operational requirements and processes (backup/restore, failover, DR etc) will change, but that's because you're making a deliberate decision weighing up those costs vs benefits.
You can use DNS to mitigate the pain of changing those connection strings, decoupling client change management from backend change process, or if you had foresight, not having to change client connection strings at all.
I mean, DNS change can work, but when you're doing that one-in-years change, why risk the extra failure modes.
It would have cost a negligible amount. But the sheer amount of time I wasted before I gave up was honestly quite surprising. Let’s see:
- I wanted one simple extension. I could have compromised on this, but getting it to work on RDS was a nonstarter.
- I wanted RDS to _import the data_. Nope, RDS isn’t “SUPER,” so it rejects a bunch of stuff that mysqldump emits. Hacking around it with sed was not confidence-inspiring.
- The database uses GTIDs and needed to maintain replication to a non-AWS system. RDS nominally supports GTID, but the documented way to enable it at import time strongly suggests that whoever wrote the docs doesn’t actually understand the purpose of GTID, and it wasn’t clear that RDS could do it right. At least Azure’s docs suggested that I could have written code to target some strange APIs to program the thing correctly.
Time wasted: a surprising number of hours. I’d rather give someone a bit of money to manage the thing, but it’s still on a combination of plain cloud servers and bare metal. Oh well.
Sounds like you are walking massive edge
Very small businesses with phone apps or web apps are often using it. There are cheaper options of course, but when there is no "prem" and there are 1-5 employees then it doesn't make much sense to hire for infra. You outsource all digital work to an agency who sets you up a cloud account so you have ownership, but they do all software dev and infra work.
> If its so fundamental to your product and needs such uptime & redundancy, what are the odds its also reasonably small?
Small businesses again, some of my clients could probably run off a Pentium 4 from 2008, but due to nature of the org and agency engagement it often needs to live in the cloud somewhere.
I am constantly beating the drum to reduce costs and use as little infra as needed though, so in a sense I agree, but the engagement is what it is.
Additionally, everyone wants to believe they will need to hyperscale, so even medium scale businesses over-provision and some agencies are happen to do that for them as they profit off the margin.
AWS and the like are rarely a cost effective option, but it is something a lot of agencies like, largely because they are not paying the bills. The clients do not usually care because they are comfortable with a known brand and the costs are a small proportion of the overall costs.
A real small business will be fine just using a VPS provider or a rented server. This solves the problem of not having on premise hardware. They can then run everything on a single server, which is a lot simpler to set up, and a lot simpler to secure. That means the cost of paying someone to run it is a lot lower too as they are needed only occasionally.
They rarely need very resilient systems as they amount of money lost to downtime is relatively small - so even on AWS they are not going to be running in multiple availability zones etc.
People pay for RDS because they want to believe in a fairy tale that it will keep potential problems away and that it worked well for other customers. But those mythical other customers also paid based on such belief. Plus, no one wants to admit that they pay money in such irrational way. It's a bubble
Come time for force major upgrade shoved down our throat? Downtime, surprise, surprise
One of these is not like the others (DBAs are not capex.)
Have you ever considered that if a company can get the same result for the same price ($100K opex for RDS vs same for human DBA), it actually makes much more sense to go the route that takes the human out of the loop?
The human shows up hungover, goes crazy, gropes Stacy from HR, etc.
RDS just hums along without all the liabilities.
We technically aren't supposed to talk about pricing publically, but I'm just going to say that we run a few 8XL and 12Xl RDS instances and we pay ~40% off the sticker price.
If you switch to Aurora engine the pricing is absurdly complex (its basically impossible to determine without a simulation calculator) but AWS is even more aggressive with discounting on Aurora, not to mention there are some legit amazing feature benefits by switching.
I'm still in agreeance that you could do it cheaper yourself at a Data Center. But there are some serious tradeoffs made by doing it that way. One is complexity and it certainly requires several new hiring decisions. Those have their own tangible costs, but there are a huge amount of intangible costs as well like pure inconvenience, more people management, more hiring, split expertise, complexity to network systems, reduce elasticity of decisions, longer commitments, etc.. It's harder to put a price on that.
When you account for the discounts at this scale, I think the cost gap between the two solutions is much smaller and these inconveniences and complexities by rolling it yourself are sometimes worth bridging that smaller gap in cost in order to gain those efficiencies.
Genuinely curious, how do you that?
We pay a couple of million dollars per year and the biggest spend is RDS. The bulk of those are 8xl and 12xl as you mention and we have a lot of these. We do have savings plans, but those are nowhere near 40%.
It looks like a reserved instance is 35% off sticker price? Add probably a discount and you'd be around 40% off.
Just like any other asset.
Besides, running things locally can be refreshingly simple if you are just starting something and you don't need tons of extra stuff, which becomes accidental complexity between you, the problem, and a solution. This old post described that point quite well by comparing Unix to Taco Bell: http://widgetsandshit.com/teddziuba/2010/10/taco-bell-progra.... See HN discussion: https://news.ycombinator.com/item?id=10829512.
I am sure for some use-cases cloud services might be worth it, especially if you are a large organization and you get huge discounts. But I see lots of business types blindly advocating for clouds, without understanding costs and technical tradeoffs. Fortunately, the trend seems to be plateauing. I see an increasing demand for people with HPC, DB administration, and sysadmin skills.
So much this. The "keep know how" has been so greatly avoided over the past 10 years, I hope people with these skills start getting paid more as more companies realize the cost difference.
My money is on openness continuing to grow and more and more pieces of the stack being completely owned by openness (kernels anyone?) but one doesn't know.
I hear tell of a shop that was running on ephemeral instance based compute fleets (EC2 spot instances, iirc), with all their prod data in-memory. Guess what happened to their data when spot instance availability cratered due to an unusual demand spike? No more data, no more shop.
Don't even get me started on the number of privacy breaches because people don't know not to put customer information in public cloud storage buckets.
If I had a small business with very clever people I'd be very afraid of what happens if they're not available for a while.
My nasty little secret is for single server databases I have zero fear of over provisioning disk iops and running it on SQLite or making a single RDBMS server in a container. I've never actually run into an issue with this. It surprises me the number of internal tools I see that depend on large RDS installations that have piddly requirements.
On what disk is the actual data written? How do you do backups, if you do?
For example, how long does it take to rent another rack that you didnt plan for?
And not to mention that the cost of cloud management platforms that you have to deploy to manage these owned assets is not free.
I mean, how come even large consumers of electricity does not buy and own their own infrastructure to generate it?
They sure do? BASF has 3 power plants in Hamburg, Disney operate Reedy Creek Energy with at least 1 power plant and I could list a fair bit more...
>For example, how long does it take to rent another rack that you didnt plan for?
I mean, you can also rent hardware a lot cheaper then on AWS. There certainly are providers where you can rent out a rack for a month within minutes
Most companies don‘t need to scale up full racks in seconds. Heck, even weeks would be ok for most of them to get new hardware delivered. The cloud planted the lie into everyone‘s head that most companies dont have predictable and stable load.
Might be other alternatives than using Docker so if anyone has tips for something simpler or easier to maintain, appreciate a comment.
I would have a hard time doing servers as cheap as hetzner for example including the routing and everything
I think there is an unreasonable fear of "doing the routing and everything". I run vpncloud, my server clusters are managed using ansible, and can be set up from either a list of static IPs or from a terraform-prepared configuration. The same code can be used to set up a cluster on bare-metal hetzner servers or on cloud VMs from DigitalOcean (for example).
I regularly compare this to AWS costs and it's not even close. Don't forget that the performance of those bare-metal machines is way higher than of overbooked VMs.
We are using hetzner cloud.. but we are also scaling up and down a lot right now
If you mean dealing with the physical dedicated servers that can be rented from Hetzner, that's what the person you replied to was talking about being not so difficult.
If you mean everything else at the data centre that makes having a server there worthwhile (networking, power, cooling, etc.) I don't think people were suggesting doing that themselves (unless you're a big enough company to actually be in the data centre business), but were talking about having direct control of physical servers in a data centre managed by someone like Hetzner.
(edit: and oops sorry I just realised I accidentally downvoted your comment instead of up, undone and rectified now)
Aka if I add up power (including backup) + backbone connection rental + server deprication I can not do it for the hetzner price..
That was quite imprecise, sorry about that.
But I think that people (at least jwr, and probably even nyc_data_geek saying "on prem") are talking about cloud (like AWS) vs. renting (or buying) servers that live in a data centre run by a company like Hetzner, which can be considered "on prem" if you're the kind of data centre client who has building access to send your own staff there to manage your servers (while still leaving everything else, possibly even legal ownership and therefore deprecation etc. to the data centre owner).
What you're thinking of - literally taking responsibility for running your own mini data centre - I think is hardly ever considered (at least in my experience), except by companies at the extremes of size. If you're as big as Facebook (not sure where the line is but obviously including some companies not AS big as Meta but still huge) then it makes sense to run your own data centres. If you're a tiny business getting less than thousands of website visits a day and where the website (or whatever is being hosted) isn't so important that a day of downtime every now and then isn't a big deal, then it's not uncommon to host from the company's office itself (just using a spare old PC or second hand cheap 1U server, maybe a cheap UPS, and just connected to the main internet connection that people in the office use, and probably managed by a single employee, or company owner, who happens to be geeky enough to think it's one or both of simple or fun to set up a basic LAMP server, or even a Windows server for its oh-so-lovely GUI).
Why? The only reason I'm using Hetzner and not AWS for several of my own projects (even though I know AWS much better since this is what I use at work) is an enormous price difference in each aspect (compute, storage, traffic).
If all you need are some cloud servers, or a basic load balancer, they are pretty much the same.
If you need a plethora of managed services and don't want to risk getting fired over your choice or specifics of how that service is actually rendered, they are nothing alike and you should go for AWS, or one of the other large alternatives (GCP, Azure etc.).
On the flip side, if you are using AWS or one of those large platforms as a glorified VPS host and you aren't doing this in an enterprise environment, outside of learning scenarios, you are probably doing something wrong and you should look at Hetzner, Contabo, or one of those other providers, though some can still be a bit pricey - DigitalOcean, Vultr, Scaleway etc.
Well, in my case at least, what they have in common is that I can choose to run my business on one or the other. So it's not about intuition, but rather facts in my case: I avoid spending a significant amount of money.
I (of course) do realize that if you design your software around higher-level AWS services, you can't easily switch. I avoided doing that.
I was once part of an acquisition from a much larger corporate entity. The new parent company was in the middle of a huge cloud migration, and as part of our integration into their org, we were required to migrate our services to the cloud.
Our calculations said it would cost 3x as much to run our infra on the cloud.
We pushed back, and were greenlit on creating a hybrid architecture that allowed us to launch machines both on-prem and in the cloud (via a direct link to the cloud datacenter). This gave us the benefit of autoscaling our volatile services, while maintaining our predictable services on the cheap.
After I left, apparently my former team was strong-armed into migrating everything to the cloud.
A few years go by, and guess who reaches out on LinkedIn?
The parent org was curious how we built the hybrid infra, and wanted us to come back to do it again.
I didn't go back.
A customer had an interest in merging the data from an older account into a new one, just to simplify matters. Enterprise data. Going back years. Not even leaving the region.
The AWS rep in the meeting kinda pauses, says: "We'll get back to you on the cost to do that."
The sticker shock was enough that the customer simply inherited the old account, rather than making things tidy.
Have people lost the ability to write export and backup scripts?
You can peer two vpc's and as long as you are transferring within the same (real) AZ, it's free: https://aws.amazon.com/about-aws/whats-new/2021/05/amazon-vp...
Even peered VPC's only pay "normal" prices: https://aws.amazon.com/ec2/pricing/on-demand/#Data_Transfer
"Data transferred "in" to and "out" from Amazon EC2, Amazon RDS, Amazon Redshift, Amazon DynamoDB Accelerator (DAX), and Amazon ElastiCache instances, Elastic Network Interfaces or VPC Peering connections across Availability Zones in the same AWS Region is charged at $0.01/GB in each direction."
It's the amount of data where it makes more sense to put hard drives on a truck and drive across the country rather than send it over a network, where this becomes an issue (actually, probably a bit before then).
> Q: Can I export data from AWS with Snowmobile? > > Snowmobile does not support data export. It is designed to let you quickly, easily, and more securely migrate exabytes of data to AWS. When you need to export data from AWS, you can use AWS Snowball Edge to quickly export up to 100TB per appliance and run multiple export jobs in parallel as necessary. Visit the Snowball Edge FAQs to learn more.
https://aws.amazon.com/snowmobile/faqs/?nc2=h_mo-lang
Why would they make it convenient to leave?
When you have enough data, that cost is quite significant.
I've heard some people here on HN say that it's slow, but I haven't noticed a difference. We're mainly dealing with multi-megabyte image files, so YMMV if you have a different workload.
I guess permissions might be more complex, as in EC2 instance profiles wouldnt grant access, etc.
I've made a career out of inheriting other peoples whacky setups and supporting them (as well as fixing them) and almost always its documentation that has prevented the client getting anywhere.
I personally dont care if the docs are crap because usually the first thing I do is update / actually write the docs to make them usable.
For a lot of techs though crap documentation is a deal breaker.
Crap docs aren't always the fault of the guys implementing though, sometimes there are time constraints that prevent proper docs being written. Quite frequently though its outsourced development agencies that refuse to write it because its "out of scope" and a "billable extra". Which I think is an egregious stance...doxs Should be part and parcel of the project. Mandatory.
If there is only one thing that juniors should learn about writing documentation (be it comments or design documents), it is this: document why something is there. If resources are limited, you can safely skip comments that describe how something works, because that information is also available in code.
(It might help to describe what is available, especially if code is spread out over multiple repositories, libraries, teams, etc.)
(Also, I suppose the comment I'm responding to could've been slightly more forgiving to GP, but that's another story.)
It's also completely against their interest to write docs as it makes their replacement easier.
That's why you need someone competent on the buying side to insist on the docs.
A lot of companies outsource because they don't have this competency themselves. So it's inevitable that this sort of thing happens and companies get locked in and can't replace their contractors, because they don't have any docs.
OK so the docs are in sync for a single point of time when you finish. Plus you get to have the context in your head (bus factor of 1, job security for you, bad for the org.)
How about if we just write clean infra configs/code, stick to well known systems like docker, ansible, k8s, etc.
Then we can make this infra code available to an on prem LLM and ask it questions as needed without it drifting out of sync overtime as your docs surely will.
Wrong documentation is worse than no documentation.
I can always guarantee a stream of consciousness one note that should have most of the important data, and a few docs about the most important parts. It's up to management if they want me to spend time turning that one note into actual robust documentation that is easily read.
Even with documentation on the hybrid setup, they'd need to get a new on-prem environment up and running (find a colo, buy machines, set up the network, blah blah).
If price is the only factor, your business model (or executives' decision-making) is questionable. Buy only the cheapest shit, spend your time building your own office chair rather than talking to a customer, you aren't making a premium product, and that means you're not differentiated.
But you are totally right it can be expensive. I worked with a startup that had some inefficient queries, normally it would matter, but with RDS it cost $3,000 a month for a tiny user base and not that much data (millions of rows at most).
RDS make perfect sense for them
That: and trust is hard earned over a long tail which is harder if you are trying to compete on price.
Small databases or test environment databases you can also leverage kubernetes to host an operator for that tiny DB. When it comes to serious data and it needs a beeline recovery strategy RDS.
Really it should be a mix self hosted for things you aren't afraid to break. Hosted for the things you put at high risk.
DynamoDB as a replacement, pay per request, was essentially free.
I found Dynamo foreign and rather ugly to code for initially, but am happy with the performance and especially price at the end.
Just to pay someone else enough money to provide the same service and make a profit while do it
It's like accounting and finance. Yeah a lot of companies use tax firms, but they all have finance and accounting in-house.
Amazon may have their own shipping fleet, most retailers of smaller scale pay someone else to do it for profit.
They need to replicate everything in multiple availability zones, which is going to be more expensive than replicating data centres.
They still need to test their cloud infrastracuture works.
Seems like a solid cost effective approach for when a company reaches a certain scale.
> Data is the most critical part of your infrastructure. You lose your network: that’s downtime. You lose your data: that’s a company ending event. The markup cost of using RDS (or any managed database) is worth it.
You need well-run, regularly tested, air gapped or otherwise immutable backups of your DB (and other critical biz data). Even if RDS was perfect, it still doesn't protect you from the things that backups protect you from.
After you have backups, the idea of paying enormous amounts for RDS in order to keep your company from ending is more far fetched.
Frankly, this is anti-competitive, and the FTC should look into it, however, Microsoft has been anti-competitive and customer hostile for decades, so if you're still using their products, you must have accepted the abuse already.
I haven't run a Postgres instance with proper backup and restore, but it doesn't seem like rocket science using barman or pgbackrest.
1. Hiring someone full time to work on the database means migrating off RDS
2. Database work is only about spend reduction
I know this is an unpopular opinion but I think google cloud is amazing compared to AWS. I use google cloud run and it works like a dream. I have never found an easier way to get a docker container running in the cloud. The services all have sensible names, there are fewer more important services compared to the mess of AWS services, and the UI is more intuitive. The only downside I have found is the lack of community resulting in fewer tutorials, difficulty finding experienced hires, and fewer third party tools. I recommend trying it. I'd love to get the user base to an even dozen.
The reasoning the author cites is that AWS has more responsive customer service and maybe I am missing out but it would never even occur to me to speak to someone from a cloud provider. They mention having "regular cadence meetings with our AWS account manager" and I am not sure what could be discussed. I must be doing simper stuff.
I suspect the team that manages it was OKR-ed into using AJAX but come from a classic ASP background, so don't understand what all this "single page app" fad is all about and hope it blows over one day
It amazes me a company that makes that much money has such a crappy client.
You'd think a company with 1.5m employees could find half a dozen decent front end developers, but apparently not
How do they do this jedi mind trick?
GCP SDK docs must be mentioned separately as it's a bizarre auto-generated nonsense. Have you seen them? How can you even say that GCP docs are good after that?
wdym? As far as I see, it's either CLI or Terraform. GCP SDK is complete garbage, at least for Python compared to AWS boto3. I have personally made web UI for AWS CLI man pages as a fun project and can index everything myself if needed. Googling works fine. If you are not happy with it then ChatGPT is to the rescue. I honestly do not see any problem at all.
We started using Azure Container Apps (ACA) and it seems simple enough.
Create ACA, point to GitHub repo, it runs.
Push an update to GitHub and it redeploys.
I don't have a ton of Azure or cloud experience but I run an Unraid server locally which has a decent Docker gui.
Getting a docker container running in Azure is so complicated. I gave up after an hour of poking around.
AWS has "fun" features like the ability to just lose track of some resource and still be billed for it. It's in here... somewhere. Not sure which region or account. I'll find it one day.
GCP is made by Google, also known as children that forgot to take their ADHD medication. Any minute now they'll just casually announce that they're cancelling the cloud because they're bored of it.
Azure is the only one I've seen with a sane management interface, where you can actually see everything everywhere all at once. Search, filter, query-across-resources, etc... all work reasonably well.
How can you even compare it to AWS is a mystery to me. There are pages showing all your resources, not sure why you think it's a problem. Could be a problem from long time ago?
[0] https://azure.microsoft.com/en-gb/products/container-apps
https://learn.microsoft.com/en-us/azure/container-instances/... is the one I was going to plug, because coming from a kubernetes background it seems to damn near be the PodSpec and thus both expresses a lot of my needs and also is very familiar https://learn.microsoft.com/en-us/azure/templates/microsoft....
Your link does seem to be a lot more "container, plus all the surrounding stuff" in line with the "apps" part, whereas mine more closely matches my actual experience of what you said: container, go
The "what the fucking hell is wrong with you people?" part is that their naming is just all over the place, and changes constantly, and is almost designed to be misleading in any sane conversation. I quite literally couldn't have guessed whether Container Apps was a prior name of Container Instances, a super set of it, subset, other? And one will observe that while I said Container Instances, and the URL says Container Instances, the ARM is Container Groups. Are they the same? different? old? who fucking knows. It's horrific
I’ve had technicians at both GCP and Azure debug code and spend hours on developing services.
Almost every time Google pulled in a specialist engineer working on a service/product we had issues with it was very very clear the engineer had no desire to be on that call or to help us. In other words they'd get no benefit from helping us and it was taking away from things that would help their career at Google. Sometimes they didn't even show up to the first call and only did to the second after an escalation up the management chain.
They often provide that consulting for free, and we know their biases. There's nothing hidden about the fact that they will push us to use AWS services.
On the other hand, they will also help us optimize those services and save money that is directly measurable.
GCP might have a better API and better "naming" of their services, but the breadth of AWS services, the incorporation of IAM across their services, governance and automation all makes it worth while.
Cloud has come a long way from "it's so easy to spin up a VM/container/lambda".
Our account team don't even do that. We use a lot of AWS anyway and they know it, so they're happy to help with competitor offerings and integrating with our existing stack. Their main push on us has been to not waste money.
AWS wants happy customers to stick around for a long time, not one month of goosed income
If you still decided to move away, and want to take data with you, yeah... there is a cost. Heck there is a cost to delete the data you have with them (like S3 content).
Its a good way to do business.
And never did I miss something in GCP that I could find in AWS. Not sure the breadth is adding much compared to a simpler product suite in GCP.
Pub/Sub - Kinesis
Cloud CDN - CloudFront
Cloud Domains - Route 53
...
Also, I don't think trying to emulate AWS's support and consistent API makes sense as a strategy for other cloud providers. They will never beat AWS at their own game, it is light years ahead. If cloud providers want to survive they need to fill a different niche and try different things.
Idk why you don't see AWS as a utility providing building blocks.
> I don't think trying to emulate AWS's support and consistent API makes sense as a strategy for other cloud providers.
Those are such essential things. It's very hard to imagine prioritising something else and succeeding in anything.
One, and you often times only need one.
Google Cloud Run - Lambda
Sure I get the reference to the underlying algebraic representation of coding but come on, Lambda tells us nothing of what it does.
Products (not brands, products) should be named in a way that means something to the customer afaic.
> Google Cloud Run - Lambda
ECS is the AWS equivalent of Cloud Run. GCP Cloud Functions are the equivalent of AWS Lambda.
ECS / Cloud Run = managed container service that autoscales
Lambda / Cloud Functions = serverless functions as a service
The only weird part I experienced with AWS is their SNS API. Maybe due to legacy reasons, but what a bizarre mess when you try doing it cross-account. This one is odd.
I have been trying GCP for a while and DevX was horrible. The only part that more-or-less works is CLI but the naming there is inconsistent and not as well-done as in AWS. But it's relative and subjective, so I guess someone likes it. I have experienced GCP official guides that broken, untested or utterly braindead hello-world-useless. And also they are numerous and spread so it takes time to find anything decent.
No dark mode is an extra punch. Seriously. Tried to make it myself with an extension but their page is Angular hell of millions embedded divs. No thank you.
And since you mentioned Cloud Run -- it takes 3 seconds to deploy a Lambda version in AWS and a minute or more for GCP Could Function.
I also worked at a smaller company using GCP. GCP refused to do a small quota increase (which AWS just does via a web form) unless I got on a call with my sales representative and listened to a 30 minute upsell pitch.
As being on a number of those calls, its just a bunch of crap where they talk like a scripted bot reading from corporate buzzword bingo card over a slideshow. Their real intention is two fold. To sell you even more AWS complexity/services, and to provide "value" to their person of contact (which is person working in your company).
We're paying north of 500K per year in AWS support (which is a highway robbery), and in return you get a "team" of people supposedly dedicated to you, which sounds good in theory but you get a labirinth of irresponsiblity, stalling and frustration in reality.
So even when you want to reach out to that team you have to first to through L1 support which I'm sure will be replaced by bots soon (and no value will be lost) which is useful in 1 out of 10 cases. Then if you're not satisfied with L1's answer(s), then you try to escalate to your "dedicated" support team, then they schedule a call in three days time, or if that is around Friday, that means Monday etc.
Their goal is to stall so you figure and fix stuff on your own so they shield their own better quality teams. No wonder our top engineers just left all AWS communication and in cases where unavoidable they delegate this to junior people who still think they are getting something in return.
I’ve found a lot of the time the issues we run into are self-inflicted. When we call support for these, they have to reverse-engineer everything which takes time.
However when we can pinpoint the issue to AWS services, it has been really helpful to have them on the horn to confirm & help us come up with a fix/workaround. These issues come up more rarely, but are extremely frustrating. Support is almost mandated in these cases.
It’s worth mentioning that we operate at a scale where the support cost is a non-issue compared to overall engineering costs. There’s a balance, and we have an internal structure that catches most of the first type of issue nowadays.
I am so tired of the support team having all the real metrics, especially in io and throttling, and not surfacing it to us somehow.
And cadence is really an opportunity for them to sell to you, the parent is completely right.
In my experience all questions I've had for AWS were: 1. Their bugs, which won't be fixed in near future anyway. 2. Their transient failures, that will be fixed anyway soon.
So there's zero value in ever contacting AWS support.
I'm not saying there's anything wrong, and I'm oversimplifying a bit, but I still find this amusing.
The argument for RDS seems to be “we can’t automate backups”. What on earth?
Totally on your side with this one - but alas, people associate value with complexity.
True story bro
I'm sure that's possible if you're storing the backup on the same server you're restoring on and everything is on top of the line nvme storage. Otherwise your backup just started to run and will need another few days to finish. And that's only if you're running single master.
You're massively underestimating the challenge to get that kind of automation done in a stable manner - and the maintenance required to keep it working over the years.
With sqlite you only need the scp part.
You can even push your backup file to an S3 bucket... with one command!
Honestly, this argument mystifies me.
Of course you can make it as complicated as you want to, too. I've also worked on replicating anonymized data from a production OLTP database to a data warehouse. That's a lot more work.
https://about.gitlab.com/blog/2017/02/01/gitlab-dot-com-data...
It took them a data loss incident to find this out? This is just one of the many red flags mentioned in the article, IMO this incident isn't about relying on cloud backups vs self managing it
I would rather pay for RDS. Databases are the one thing you don't want to screw up.
For me that's short-sighted.
Hundreds of dead startups because after all that unnecessary spending, they still have unnecessary buggy software that got sold to other startups that, when push comes to shove, will cut spending in those same startups that offer half-baked buggy products.
What you say is definitely what they preach. But I don't agree or see that as a good logic.
Way too many startup founders decide to build shitty products with short-sighted solutions like these, following whatever is trendy (crypto, AI, etc) because investors advise them to. Guess what: the investor doesn't care about creating a good business. He wants a unicorn. So they advise them to make all-or-nothing moves knowing it will most likely kill the startup.
It's definitely "a strategy". But I think it's short-sighted as hell.
These investment rounds are only there to provide money to start. Unless the founders sign away the control of their company to investors, they are still at the helm of the ship. They can choose how they will approach growth.
We can see how the entire scene is gearing towards profitability now that the money dried up, so this growth focus is no longer the only game in town.
You can still take investments and accelerate growth without having to recklessly go all-in. But I've never taken VC money. Maybe that's baked into the contracts?
I can automate backups and I'm extremely happy they with some extra cost in RDS, I don't have to do that.
Also, at some size automating the database backup becomes non-trivial. I mean, I can manage a replica (which needs to be updated at specific times after the writer), then regularly stop replication for a snapshot, which is then encrypted, shipped to storage, then manage the lifecycle of that storage, then setup monitoring for all of that, then... Or I can set one parameter on the Aurora cluster and have all of that happen automatically.
And, when factoring in all costs and considering all things the service takes care of, it seems like a reasonable assumption that in a free market a team that specializes in optimizing this entire operation will sell you a db service at a better net rate than you would be able to achieve on your own.
Which might still turn out to be false, but I don't think it's obvious why.
The point isn’t that you can’t do it, the point is that it’s less work for extremely high standards. It is not easy to configure multi region failover without an entire network team and database team unless you don’t give a shit about it actually working. Oh yea, and wait until you see how much SOC2 costs if you roll your own database.
My contrarian view is that EC2 + ASG is so pleasant to use. It’s just conceptually simple: I launch an image into an ASG, and configure my autoscale policies. There are very few things to worry about. On the other hand, using k8s has always been a big deal. We built a whole team to manage k8s. We introduce dozens of concepts of k8s or spend person-years on “platform engineering” to hide k8s concepts. We publish guidelines and sdks and all kinds of validators so people can use k8s “properly”. And we still write 10s of thousands lines of YAML plus 10s of thousands of code to implement an operator. Sometimes I wonder if k8s is too intrusive.
But in general k8s provides incredibly solid abstractions for building portable, rigorously available services. Nothing quite compares. It's felt very stable over the past few years.
Sure, EC2 is incredibly stable, but I don't always do business on Amazon.
Many of the core concepts of Kubernetes should be taken to build a new alternative without all the footguns. Security should be baked in, not an afterthought when you need ISO/PCI/whatever.
Who exactly needs millions of lines of code?
And in the end it's hard to say if you've actually gained anything except now this different code manages your AWS resources like you were doing with CF or terraform.
How many LOCs in the linux kernel again?
I don't know what you have been doing with Kubernetes, but I run a few web apps out of my own Kubernetes cluster and the full extent of my lines of code are the two dozen or so LoC kustomize scripts I use to run each app.
0 https://github.com/kube-hetzner/terraform-hcloud-kube-hetzne...
It's really not about what I do and do not do with Kubernetes. It's on you to justify your "millions upon millions lines of code" claim because it is so outlandish and detached from reality that it says more about your work than about Kubernetes.
I repeat: I only need a few dozen lines of kustomize scripts to release whole web apps. Simple code. Easy peasy. What mess are you doing to require "millions upon millions" lines of code?
Sometimes I think that managed kubernetes services like EKS are the epitome of "give the customers what they want", even when it makes absolutely no sense at all.
Kubernetes is about stitching together COTS hardware to turn it into a cluster where you can deploy applications. If you do not need to stitch together COTS hardware, you have already far better tools available to get your app running. You don't need to know or care in which node your app is suppose to run and not run, what's your ingress control, if you need to evict nodes, etc. You have container images, you want to run containers out of them, you want them to scale a certain way, etc.
Lifting and shifting an "EC2 + ASG" set-up to Kubernetes is a straightforward process unless your app is doing something very non-standard. It maps to a Deployment in most cases.
The fact that you even implemented an operator (a very advanced use-case in Kubernetes) strongly suggests to me that you're doing way more than just lifting and shifting your existing set-up. Is it a surprise then that you're seeing so much more complexity?
> Moving off JIRA onto linear
I don't get the hype. Linear is fine and all but I constantly find things I either can't or don't know how to do. How do I make different ticket types with different sets of fields? No clue.
> Not using Terraform Cloud No Regrets
I generally recommend Terraform Cloud - it's easy for you to grow your own in house system that works fine for a few years and gradually ends up costing you in the long run if you don't.
> GitHub actions for CI/CD Endorse-ish
Use Gitlab
> Datadog Regret
Strong disagree - it's easily the best monitoring/observability tool on the market by a wide margin.
Cost is the most common complaint and it's almost always from people who don't have it configured correctly (which to be fair Datadog makes it far too easy to misconfigure things and blow up costs).
> Pagerduty Endorse
Pagerduty charges like 10x what Opsgenie does and offers no better functionality.
When I had a contract renewal with Pagerduty I asked the sales rep what features they had that Opsgenie didn't.
He told me they're positioning themselves as the high end brand in the market.
Cool so I'm okay going generic brand for my incident reporting.
Every CFO should use this as a litmus test to understand if their CTO is financially prudent IMO.
So if that's the upgrade path you're going down I'd expect it to be fantastic.
Linear has been such a breath of fresh air, with such a solid desktop app (on Mac OS) that I don’t ever want to go back. Stuff happens instantly, the layout and semantics are an excellent “90% good enough” that I would happily relegate jira to only the most enterprise of enterprise projects.
There are lots of things where Jira falls short, but the pain points on an under-resourced self hosted instance of ten years ago are nothing like the ones you'll find on Jira cloud today.
It takes markdownish input but converts it to rich text as you type - so asterisk-space starts a bullet point list, etc.
I actually can't remember if it has a dedicated markdown mode anymore; the rich text editing supports the usual shortcuts that mean I tend to stick with it.
Given Atlassian bought OpsGenie in 2018, this either somewhere between quite late and unsurprising.
Anything Atlassian does is mostly quite late and its integration story is so pathetic that it's unsurprising.
Try to have a bitbucket pipeline that pushes to confluence. Seems like a basic integration to have, after all, Confluence has an API (well, actually it has 3 different ones) so surely Atlassian would make a basic thing like "publish a wiki page" a thing you get out of the box.
Nope.
I suppose it comes back to the comparative priorities (as evaluated by recurrent revenue) of ticking rfq boxes vs solving actual problems.
In fact, OpsGenie has mostly been on Auto-pilot for a few years now.
To me its whole schedule interface is atrocious for its price, given from an SRE/dev perspective, that's literally its purpose - scheduled escalations.
OpsGenie’s cheapest is $9 per user month but arbitrarily crippled, the plan anybody would want to use is $19 per user month
So instead of a factor of ten it’s ten percent cheaper. And i just kind of expect Atlassian to suck.
Datadog is ridiculously expensive and on several occasions I’ve run into problems where an obvious cause for an incident was hidden by bad behavior of datadog.
We use iOS “Critical Alerts” and similar on Android that breaks through any Do-Not-Disturb settings. https://heiioncall.com/blog/better-alerting-for-heii-on-call... Would you be willing to give that a shot? It wakes me every time :)
(It’s configurable too; we have vibrate-only or silenced modes. Think old-school beeper.)
In the rare case that it doesn’t wake you, we have configurable escalation strategies to alert someone else on your team after a configurable number of minutes.
I usually do not respond immediately to phone notifications, which I can handle async. Phone calls are by definition sync.
Heii On-Call will keep alerting you with these “Critical Alerts” until you’ve manually acknowledged. (Or until it escalates to a teammate and they acknowledge…)
And at least on my phone they sound nothing like normal phone notifications, which I personally always have on vibrate and/or DND anyway.
Give it a try and I think you’ll like it.
With SMS, Phone Calls and Critical Alerts / DnD override.
We're 5 USD/user.
We try to build as close to our users as possible. Happy for any new try outs! :)
(I am co founder)
I loved Datadog 10 years ago when I joined a company that already used it where I never once had to think about pricing. It was at the top of my list when evaluating monitoring tools for my company last year, until I got to the costs. The pricing page itself made my head swim. I just couldn’t get behind subscribing to something with pricing that felt designed to be impossible to reason about, even if the software is best in class.
Their pricing setup is evil. Breaking out by SKUs and having 10+ SKUs is fine, trialing services with “spot” prices before committing to reserved capacity is also fine.
But (for some SKUs, at least) they make it really difficult to be confident that the reserved capacity you’re purchasing will cover your spot use cases. Then, they make you contact a sales rep to lower your reserved capacity.
It all feels designed to get you to pay the “spot” rate for as long as possible, and it’s not a good look.
I understand the pressures on their billing and sales teams that lead to these patterns, but they don’t align with their customers in the long term. I hope they clean up their act, because I agree they’re losing some set of customers over it.
Then they have things that I wanted to try for a long time, but... support doesn't care? Repeated "would you like to use this? / very likely, can we try it out? / (silence)". I love their product, but they are so annoying to deal with at the billing level.
I, quite literally, was griping to my Datadog CSM about this exact thing last week. They'll email me and be, "Oh, you know you're logging volume this month put you into on-demand indexing rates, right?" and my answer is always, "No, because your monitoring platform makes it nearly impossible for me to monitor it correctly."
You can't reference your contracted volume rates when building monitors out and the units for the metrics you need to watch don't match the units you contract with them on the SKU.
Maddening.
Are you referring to the `datadog.estimated_usage.logs.ingested_events` metric? It includes excluded events by default but you can get to your indexed volume by excluding excluded logs. `sum:datadog.estimated_usage.logs.ingested_events{datadog_index:*,datadog_is_excluded:false}.as_count()`
I'll give you a fun example. It's fresh in my mind because i just got reamed out about it this week.
In our last contract with DataDog, they convinced us to try out the CloudSIEM product, we put in a small $600/mo committment to it to try it out. Well, we never really set it up and it sat on autopilot for many months. We fell under our contract rate for it for almost a year.
Then last month we had some crazy stuff happen and we were spamming logs into DataDog for a variety of reasons. I knew I didn't want to pay for these billions of logs to be indexed, so I made an exclusion filter to keep them out of our log indexes so we didn't have a crazy bill for log indexing.
So our rep emailed me last week and said "Hey just a heads up you have $6,500 in on-demand costs for CloudSIEM, I hope that was expected". No, it was NOT expected. Turns out excluding logs from indexing does not exclude them from CloudSIEM. Fun fact, you will not find any documented way to exclude logs from CloudSIEM ingestion. It is technically possible, but only through their API and it isn't documented. Anyway, I didn't do or know this, so now i had $6,500 of on-demand costs plus $400-500 misc on-demand costs that I had to explain to the CTO.
I should mention my annual review/pay raise is also next week (I report to the CTO), so this will now be fresh in their mind for that experience.
Datadog on the other side... their "DD University" is a shame and we as paying customers are overwhelmed and with no real guidance. DD should assign some time for integration for new customers, even if it is proportional to what you pay annually. (I think I pay around 6000 usd annually.
Jira: Its overhyped and overpriced. Most HATE jira. I guess I don't care enough. I've never met a ticket system that I loved. Jira is fine. Its overly complex sure. But once you set it up, you don't need to change it very often. I don't love it, I don't hate it. No one ever got fired for choosing Jira, so it gets chosen. Welcome to the tech industry.
Terraform Cloud: The gains for Terraform Cloud are minimal. We just use Gitlab for running Terraform pipelines and have a super nice custom solution that we enjoy. It wasn't that hard to do either. We maintain state files remotely in S3 with versioning for the rare cases when we need to restore a foobar'd statefile. Honestly I like having Terraform pipelines in the same place as the code and pipelines for other things.
GitHub Actions: Yeah switch to GitLab. I used to like Github Actions until I moved to a company with Gitlab and it is best in class, full stop. I could rave about Gitlab for hours. I will evangelize for Gitlab anywhere I go that is using anything else.
DataDog: As mentioned, DataDog is the best monitoring and observability solution out there. The only reason NOT to use it is the cost. It is absurdly expensive. Yes, truly expensive. I really hate how expensive it is. But luckily I work somewhere that lets us have it and its amazing.
Pagerduty: Agree, switch to OpsGenie. Opsgenie is considerably cheaper and does all the pager stuff of Pager duty. All the stuff that PagerDuty tries to tack on top to justify its cost is stuff you don't need. OpsGenie does all the stuff you need. Its fine. Similar to Jira, its not something anyone wants anyway. No ones going to love it, no one loves being on call. So just save money with OpsGenie. If you're going to fight for the "brand name" of something, fight for DataDog instead, not a cooler pager system.
So on the standard tech hype cycle, that sounds about right.
- It's fast. It's wild that this is a selling point, but it's actually a huge deal. JIRA and so many other tools like it are as slow as molasses. Speed is honestly the biggest feature.
- It looks pretty. If your team is going to spend time there, this will end up affecting productivity.
- It has a decent degree of customization and an API. We've automated tickets moving across columns whenever something gets started, a PR is up for review, when a change is merged, when it's deployed to beta, and when it's deployed to prod. We've even built our own CLI tools for being able to action on Linear without leaving your shell.
- It has a lot of keyboard shortcuts for power users.
- It's well featured. You get teams, triaging, sprints (cycles), backlog, project management, custom views that are shareable, roadmaps, etc...
Datadog's cheapest pricing is $15/host/month. I believe that is based on the largest sustained peak usage you have.
We run spot instances on AWS for machine learning workflows. A lot of them if we're training and none otherwise. Usually we're using zero. Using DataDog at it's lowest price would basically double the cost of those instances.
You're staying within an ecosystem you know and it seems to offer almost all of the necessary functionality
Getting them to use Github/Gitlab is an argument I've never won. Typically it goes the other way and I end up needing to maintain a Monday or Airtable instance in addition to my ticketing system.
For folks running k8 at any sort of scale, I generally recommend aggregating metrics BEFORE sending them to datadog, either on a per deployment or per cluster level. Individual host metrics tend to also matter less once you have a large fleet.
You can use opensource tools like veneur (https://github.com/stripe/veneur) to do this. And if you don't want to set this up yourself, third party services like Nimbus (https://nimbus.dev/) can do this for you automatically (note that this is currently a preview feature). Disclaimer also that I'm the founder of Nimbus (we help companies cut datadog costs by over 60%) and have a dog in this fight.
I'll be dead in the ground before I use TFC. 10 cents per resource per month my ass. We have around 100k~ resources at an early-stage startup I'm at, our AWS bill is $50~/mo and TFC wants to charge me $10k/mo for that? We can hire a senior dev to maintain an in-house tool full time for that much.
Things were more far more manual and much less secure, scalable and reliable in the past, but they were also far far simpler.
With all of that complexity/word salad from TFA, where’s the value delivered? Presumably there’s a product somewhere under all that infrastructure, but damn, what’s left to spend on it after all the infrastructure variable costs?
I get it’s a list of preferences, but still once you’ve got your selection that’s still a ton of crap to pay for and deal with.
Do we ever seek simplicity in software engineering products?
Cloud and SaaS tools are very seductive, but I think they're ultimately a trap. Keep your tools simple and just run them yourselves, it's not that hard.
Doubtfully. Simplicity of work breakdown structure - maybe. Legibility for management layers, possibly. Structural integrity of your CYA armor? 100%.
The half-life of a software project is what now, a few years at most these days? Months, in webdev? Why build something that is robust, durable, efficient, make all the correct engineering choices, where you can instead race ahead with a series of "nobody ever got fired for using ${current hot cloud thing}" choices, not worrying at all about rapidly expanding pile of tech and organizational debt? If you push the repayment time far back enough, your project will likely be dead by then anyway (win), or acquired by a greater fool (BIG WIN) - either way, you're not cleaning up anything.
Nobody wants to stay attached to a project these days anyway.
/s
Maybe.
Either because they only know how to manage AWS instances (it was the hotness and thats what all the blogs and YT videos were about) and are now terrified from losing their jobs if the companies switch stacks. Or because they needed to put the new thing on their CV so they remain employable. Also maybe because they had to get that promotion and bonus for doing hard things and migrating things. Or because they were pressured into by bean counters which were pressured by the geniuses of Wall Street to move capex to opex.
In any case, this isn't by necessity these days. This is because, for a massive amount of engineers, that's the only way they know how to do things and after the gold rush of high pay, there's not many engineers around that are in it to learn or do things better. It's for the paycheck.
It is what it is. The actual reality of engineering the products well doesn't come close to the work being done by the people carrying that fancy superstar engineer title.
You know the old adage "fast, cheap, good: pick two"? With startups, you're forced to pick fast. You're still probably not gonna make it, but if you don't build fast, you definitely won't.
There's an easy bent towards designing everything for scale. It's optimistic. It's feels good. It's safe, defendable, and sound to argue that this complexity, cost, and deep dependency is warranted when your product is surely on the verge of changing the course of humanity.
The reality is your SaaS platform for ethically sourced, vegan dog food is below inconsequential and the few users that you do have (and may positively affect) absolutely do not not need this tower of abstraction to run.
For 99% of businesses it's a wasteful, massive overkill expense. You dont NEED all the shiny tools they offer, they don't add anything to your business but cost. Unless you're a Netflix or an Apple who needs massive global content distribution and processing services theres a good chance you're throwing money away.
The “control plane” was ZooKeeper. Everything had bindings to it, Thrift/Protobuf goes in a znode fine. List of servers for FooService? znode.
The packaging system was a little more complicated than a tarball, but it was spiritually a tarball.
Static link everything. Dependency hell: gone. Docker: redundant.
The deployment pipeline used hypershell to drop the packages and kick the processes over.
There were hundreds of services and dozens of clusters of them, but every single one was a service because it needed a different SKU (read: instance type), or needed to be in Java or C++, or some engineering reason. If it didn’t have a real reason, it goes in the monolith.
This was dramatically less painful than any of the two dozen server type shops I’ve consulted for using kube and shit. It’s not that I can’t use Kubernetes, I know the k9s shortcuts blindfolded. But it’s no fun. And pros built these deployments and did it well, serious Kubernetes people can do everything right and it’s complicated.
After 4 years of hundreds of elite SWEs and PEs (SRE) building a Borg-alike, we’d hit parity with the bash and ZK stuff. And it ultimately got to be a clear win.
But we had an engineering reason to use containers: we were on bare metal, containers can make a lot of sense on bare metal.
In a hyperscaler that has a zillion SKUs on-demand? Kubernetes/Docker/OCI/runc/blah is the friggin Bezos tax. You’re already virtualized!
Some of the new stuff is hot shit, I’m glad I don’t ssh into prod boxes anymore, let alone run a command on 10k at the same time. I’m glad there are good UIs for fleet management in the browser and TUI/CLI, and stuff like TailScale where mortals can do some network stuff without a guaranteed zero day. I’m glad there are layers on top of lock servers for service discovery now. There’s a lot to keep from the last ten years.
But this yo dawg I heard you like virtual containers in your virtual machines so you can virtualize while you virtualize shit is overdue for its CORBA/XML/microservice/many-many-many repos moment.
You want reproducibility. Statically link. Save Docker for a CI/CD SaaS or something.
You want pros handing the datacenter because pets are for petting: pay the EC2 markup.
You can’t take risks with customer data: RDS is a very sane place to splurge.
Half this stuff is awesome, let’s keep it. The other half is job security and AWS profits.
that would have been around the time when containers entered the public/developer consciousness, no?
There is no way one person can thoroughly understand so many complex pieces of technology. I have worked for 10 years more or less at this point, and I would only call myself confident on 5 technical products, maybe 10 if I being generous to myself.
Why do we need entire teams making 1000s of micro decisions to deploy our app?
I’m hungry for a simpler way, and I doubt I’m alone in this.
Does not mean each of these things don’t solve problems. The issue as always about complexity-utility tradeoff. Some of these things have too much complexity for too little utility. I’m not qualified to judge here, but if the suspects have Turing-complete-yaml-templates on their hands, it probably ties them to the crime scene.
The problem was: too much money, too few consequences for burning it.
The existence of the uber-wealthy means that markets can no longer function efficiently. Every market remains irrational longer than anyone who's not uber-wealthy can remain solvent.
Welcome to the new normal.
I recreated the environment in ECS in 1/10th the time and everything just worked.
We have moved off of it though, you can eventually need more features than it provides. Of course that journey always ends up in Kubernetes land, so you eventually will find your way back there.
Logging to Cloudwatch from kubernetes is good for one thing... audit logs. Cloudwatch in general is a shit product compared to even open source alternatives. For logging you really need to look at Fluentd or Kibana or DataDog or something along those lines. Trying to use Cloudwatch for logs is only going to end in sadness and pain.
In ECS, service updates would take 15 min or more (vs basically instant in K8s).
ECS has weird limits on how many containers you can run on one instance [0]. And in the network mode where you can run more containers on a host, then the DNS is a mess (you need to lookup SRV records to find out the port).
Using ECS with CDK/Cloudformation is very painful. They don't support everything (specially regarding Blue/Green deployments), and sometimes they can't apply changes you do to a service. When initially setting up everything, I had to recreate the whole cluster from scratch several times. You can argue that's because I didn't know enough, but if that ever happened to me on prod I'd be screwed.
I haven't used EKS (I switched to Azure), so maybe EKS has their own complex painful points. I'm trying to keep my K8s as vanilla as possible to avoid the cloud lock-in.
[0] https://docs.aws.amazon.com/AmazonECS/latest/bestpracticesgu...
On the other hand, I was able to spin up an entire ECS cluster in a few minutes time with no manual operations and entirely within CloudFormation. ECS costs nothing extra, so creating multiple clusters is very reasonable, though separate clusters would impact packing efficiency. The applications can be fully independent.
> ECS has weird limits on how many containers you can run on one instance
Interesting. With ECS it says for c5.large the task limit is 2 with without trunking, 10 with.
With EKS
$ ./max-pods-calculator.sh --instance-type c5.large --cni-version 1.12.6
29
$ ./max-pods-calculator.sh --instance-type c5.large --cni-version 1.12.6 --cni-prefix-delegation-enabled
110My approach on Azure has been to rely as little as possible in their Infra-as-code, and do everything I can to setup the cluster using K8s native stuff. So, add-ons, RBAC, metrics, all I'd try to handle with Helm. That way if I ever need to change K8s provider, it "should" be easy.
Why not dump your application server and dependencies into rented data center (or EC2 if you must) and setup a coarse DR? Maybe start with a monolith in PHP or Rails.
None of that word salad sounds like startup to me, but then again everyone loves to refer to themselves as a startup (must be a recruiting tool?), so perhaps muh dude is spot on.
Right now I'm engineer No. 1 at a current startup just doing DDD with a Django monolith. I'm still pretty Jr. and I'm wondering if there's a way to scale without needing to get into all of the things the author of this article mentions. Is it possible to get to a $100M valuation without needing all of this extra stuff? I realize it varies from business to business, but if anyone has examples of successes where people just used simple architecture's I'd appreciate it.
If you're looking for successful businesses, indie hackers like levelsio show you how far you can get with very simple architectures. But that's solo dev work - once you have a team and are dealing with larger-scale data, things like infrastructure as code, orchestration, and observability become important. Kubernetes may or may not be essential depending on what you're building; it seems good for AI companies, though.
If you're primarily building a web app, a monolith is fine for quite a while, I think. But a lot of the stuff in the post is still relevant even for monoliths - RDS, Redis, ECR, terraform, pagerduty, monitoring/observability.
That said, anything that's set-and-forget is great to start with. Anything that requires it's own care and feeding can wait unless it's really critical. I think we have a project each quarter to optimize our datadog costs and renegotiate our contract.
Also if you make microservices, you are going to need a ton of tools.
If you are running a marketplace app and collect fees you're going to be able go much further on simpler architectures than if you're trying to generate 10,000 AI images per second.
The things I did to get here are honestly kind of stupid. I started out at a defense contractor after graduating and left in the first six months because all the software devs were jumping ship. Went to a small business defense contractor (yep that's a thing) and learned to build web apps with React and Django. Then the pace of business slowed so after about 18 months I got on the Leetcode grind and got into a FAANG. Realized I hated it, so I quit after about 9 months with no job lined up.
While unemployed I convinced myself I was going to get a job in robotics (I actually got pretty close, I had 3 final level interviews with robotics companies), but the job market went to shit pretty much the exact day I quit my job lol. I spent about 6 months just learning ROS, Inverse Kinematics, math for robotics, gradient descent and optimization, localization, path planning, mapping etc. I taught at a game development summer camp for a month and a half, that was awesome. Working with kids is always a blast. Also learned Rust and built a prototype for a multiplayer browser-based coding game I had been thinking about for a while. It was an excuse to make a full stack application with some fun infrastructure stuff.
https://ai-arena.com/#/multiplayer
The backend is no longer running, but originally users could see their territory on the galaxy grow as their code won battles for them.
For the current role, I really just got lucky. The previous engineer was on his way out for non-job related reasons. He had read a lot of the books I had (Code Complete, Domain Driven Design) and I think we just connected over shared interests and intellectual curiosity.
I think that in the modern day, so many people are really just in this space for the paycheck-- and that's okay! Everyone needs to make a living. But I think that if you have that intellectual curiosity and like making stuff, people will see that and get excited. It ends up being a blessing and a curse.
I have failed interviews because of honesty "I would Google the names of books and read up on that subject" or "I think if I was doing CSS then I would be in the wrong role" (I realize how douchey that sounds but I just was not meant to design things, I have tried). But I have also gone further in interviews than I should have because I was really engrossed in a particular problem like path planning or inverse kinematics and I was able to talk about things in plain terms.
I think it's easier to learn things quickly if they are something you're actually interested in, it becomes effortless. Basically I just try to do that so I can learn optimally, then I try to get lucky.
EDIT: Oh I just thought of more good advice. Find senior devs to learn from. They can be kind of grumpy in their online presence, but they help you avoid so many tar pits. I am in a Discord channel with a handful of senior engineers. The best way to get feedback is to naively say "I'm going to do X", they will immediately let you know why X is a bad idea. A lot of their advice boils down to KISS and use languages with strong typing.
In my case it took a good five years and a couple job hops to rebalance. But eventually you get back to a reasonable tech leadership role and back to making some of the bigger decisions to help make the junior devs' lives less miserable.
No regrets, but the five years it takes to rebalance can be pretty hard.
After realizing that, I decided I'd try as hard as I possibly could to never have to work at a job that I didn't like. I already didn't want kids so that part is easy. The other part of the equation is saving lots of money. I'm not an ascetic by any means, but I live well below my means on a SWE salary which means I can save quite a bit of money each year.
I also recognize that not wanting to go corporate severely limits my options down the line. But capitalism is all about making money for other people. If I can make someone a lot of money, they're not going to care about if I have the chops to stand up a Kubernetes cluster or write a Next.js app or whatever (I hope).
I don't think I'm super smart, I'd say I'm pretty average for this line of work. But I reckon that most SWEs are focused on learning new technologies to get to their next job, or are overly concerned with technical problems. I like to think that I am pragmatic enough about only doing things that are going to deliver business value to make up for being average in smarts.
Anyways, there's not really a point to this rant. These are just some thoughts I have had about optimizing my career for my own happiness, and how I hope I can stay a hot commodity even though I hate working in the cloud and my software skills aren't bleeding edge.
One thing though, I'd start with go. It's no more complex than python, more efficient, and most importantly IMO since it compiles down to binary it's easier to build, deploy, share, etc. And there's less divergence in the ecosystem; generally one simple way to do things like building and packaging, etc. I've not had to deal with versions or tooling or environmental stuff nearly as much since switching.
Fortunately, with managed DBs like RDS it is really easy to run individual DB clusters per major app.
This is rarely a problem when things are small, but as they grow, the bad schema decisions made by empowering DBA-less teams to run their own infra become glaringly obvious.
In the kitchen sink model all teams are tied together for performance and scalability, and some bad apple applications can ruin the party for everyone.
Seen this countless times doing due diligence on startups. The universal kitchen sink DB is almost always one of the major tech debt items.
Multi-tenant DBs can work fine as long as every app has its own users, everyone goes through a connection pooler / load balancer, and every user has rate limits. You want to write shitty queries that time out? Not my problem. Your GraphQL BFF bullshit is trying to make 10,000 QPS? Nope, sorry, try again later.
EDIT: I say “not my problem,” but as mentioned, it inevitably becomes my problem. Because “just unblock them so the site is functional” is far more attractive to the C-Suite than “slow down velocity to ensure the dev teams are doing things right.”
Full Stack is a lie, and the sooner companies accept that and allow people to specialize again, and to pay for the extra headcount, the better off everyone will be.
This is how you end up with the infamous "jira and confluence have two different markdown flavors" issue.
This will undoubtedly go over poorly, but honestly I think every data decision should be gated through the DB Team (again, if you have them). Your proposed schema isn’t normalized? Straight to jail. You don’t want to learn SQL? Also straight to jail. You want to use a UUIDv4 as a primary key? Believe it or not, jail.
The most performant and referentially sound app in the world, because of jail.
Uuids are really for external communication, not in-system organization.
They require periodic synchronization. What isn't a big deal at all and is required by many other database features.
PlanetScale uses int PKs [0], and they seem to have scaled just fine.
[0]: https://github.com/planetscale/discussion/discussions/366
[0]: https://www.percona.com/blog/uuids-are-popular-but-bad-for-p...
[1]: https://www.cybertec-postgresql.com/en/unexpected-downsides-...
[2]: https://www.2ndquadrant.com/en/blog/on-the-impact-of-full-pa...
DB team could act as an auditor and expert support, but they should never be fully responsible for DB layer.
That’s the point. Would you send a backend code review to a frontend team? Why do DBs not deserve domain expertise, especially when the entire company depends on them?
> they are not responsible for the whole product, just for the database
I assure you, that’s a lot to be responsible for at scale.
> DB team could act as an auditor and expert support, but they should never be fully responsible for DB layer.
Again, the issue here is when the DB gets borked enough that a SME is required to fix it, they effectively do become responsible, because no CTO is going to accept, “sorry, we’ll be down for a couple of days because our team doesn’t really know how this thing works.”
And if your answer is, “AWS Premium Support,” they’ll just tell you to upsize the instance. Every time. That is not a long-term strategy.
I wish and maybe there is a programming language with first class database support. I mean really first class not just let me run queries but almost like embedded into the language in a primal way where I can both deal with my database programming fancyness and my general development together.
Sincerely someone who inherited a project from a DBA.
Not quite embedded into the OS, but Django is a damn good ORM. I say that as a DBRE, and someone obsessed with performance (inherent issues with interpreted languages aside).
I have worked in many languages with many ORMs and this has been my personal favorite.
[0]: https://github.com/prisma/prisma/issues/5184#issuecomment-18...
One thing that has worked well for us is to alway include the top-most parent key in all child tables down yhe hierarchy. This way we can load all the data for say an order without joins/exists.
Oh and never use natural keys. Each time I thought finally I had a good use-case, it has bitten me in some way.
Apart from that we just try to think about the required data access and the queries needed. Main thing is that all queries should go against indexes in our case, so we make sure the schema supports that easily. Requires some educated guesses at times but mostly it's predictable IME.
Anyway would love to see a proper resource. We've made some mistakes but I'm sure there's more to learn.
With that said, this still sounds like a strange situation - most colleagues, acquaintances and people I consulted know they way around SQL and dropping down to 'dbset.FromSql($"SELECT {...' is very commonplace out of the need to use sprocs, views or have tighter control over the query.
But schema design is something else. I still take my time doing that.
Especially since our application is written with backwards compatibility in mind, so changing schema after it's deployed is something we try very hard to avoid.
But yeah, when hiring we require they are comfortable writing "normal" SQL queries (multiple joins, aggregation etc).
The neat thing is, you don't. Nobody ever avoids fucking up db design.
The best you can do is decide what is really important to get right, and not fuck that part up.
P.S. to the original person concerned about this though… for your own sake and your successors, please keep trying.
Just do the exercise of deciding what is really important first, so you can make sure you succeed for that stuff.
If you can't do something like determine if you can delete data, as the article mentions, you won't be able to produce an answer to how to deal with those problems.
Being shared between applications is literally what databases were invented to do. That’s why you learn a special dsl to query and update them instead of just doing it in the same language as your application.
The problem is that data is a shared resource. The database is where multiple groups in an organization come together to get something they all need. So it needs to be managed. It could be a dictator DBA or a set of rules designed in meetings and administered by ops, or whatever.
But imagine it was money. Different divisions produce and consume money just like data. Would anyone imagine suggesting either every team has their own bank account or total unfettered access to the corporate treasury? Of course not. You would make a system. Everyone would at least mildly hate it. That’s how databases should generally be managed once the company is any real size.
Decades of experience have shown us the massive costs of doing so - the crippled velocity and soul crushing agony of dba change control teams, the overhead salary of database priests, the arcane performance nightmares, the nuclear blast radius, the fundamental organizational counter-incentives of a shared resource .
Why on earth would we choose to pay those terrible prices in this day and age, when infrastructure is code, managed databases are everywhere and every team can have their own thing. You didn’t have a choice previously, now you do.
You DO have to share data in other ways, usually datawarehouse or services, but that is not the same thing.
I’m not saying literally every source of data has to be shared and centrally managed. I’m also not saying “rdbms accessed via traditional client and queried via sql” when I say database. I’m just saying a shared database of some shape is inevitable.
Also, operationally it’s not “semantics” at all. You don’t get into (many) operational problems with analysts sharing a datawarehouse. You absolutely do with online apps sharing a rdbms, they aren’t the same thing.
A data warehouse is a type of database and is does need to be managed. Your assertion that it is easier to manage is orthogonal to my assertion that there will always be a central database to manage in an organization of decent size.
1) We don't have any software, so we don't have a prod environment.
2) We have 1 team that makes 1 thing, so we just launch it out of systemd.
3) We have between 2 and 1000 teams that make things and want to self-manage when stuff gets rolled out.
Kubernetes is case 3. Like it or not, teams that don't coordinate with each other is how startups scale, just like big companies. You will never find a director of engineering that says "nah, let's just have one giant team and one giant codebase".
Kubernetes is appealing to many, but it is not 100% frictionless. There are upgrades to manage, control plane limits, leaky abstractions, different APIs from your cloud provider, different RBAC, and other things you might prefer to avoid. It's its own little world on top of whatever world you happen to be running your foundational infrastructure on.
Or, as someone has artistically expressed it: https://blog.palark.com/wp-content/uploads/2022/05/kubernete...
I have seen people rewrite Kubernetes in CloudFormation. You can do it! But it certainly isn't problem-free.
You’re right that if you use a cloud provider, IAM is something that has to be reckoned with. But the question is, how many implementations of IAM and policy mechanisms do I want to deal with?
Also, you can bundle your load balancer config and application config together. No written description of the load balancer config + an RPM file to a disinterested different team.
https://github.com/martinvonz/jj https://github.com/facebook/sapling
I see Kubernetes as one time (mental and time) investment that buys me somehow smoother sailing plus some other benefits.
Of course it is not all rainbows and unicorns. Having a single nginx server for a single /static directory would be my dream instead of MinIO and such.
I have spent vanishingly close to 0 hours on maintaining our (managed) kubernetes clusters in work over the past 3 years, and if I didn't show up tomorrow my replacement would be fine.
Admittedly, I was afraid of ever restarting as I wasn’t sure it would reboot. But still…
For that, you get automated backups, very simple read proxies, managed updates of you ever need them. You can vertically scale down, or uo to the point of "it's cheaper to hire a DBA to fix this".
We use this for our internal services at work, and the last time I touched the infra was in 2022 according to git
[0] https://github.com/gin-gonic/gin
[1] https://gist.github.com/donalmacc/0efbb0b377533232da3f776c60....
[2] https://docs.digitalocean.com/products/kubernetes/how-to/dep...
You also need to pay them which is an event.
- I want to spin up multiple redundant instances of some set of services
- I want to load balance over those services
- I want some form of rolling deploy so that I don’t have downtime when I deploy
- I want some form of declarative infrastructure, not click-ops
Given these requirements, I can’t think of an alternative to managed k8s that isn’t more complex.
Running off a couple of medium ( $3k/month each range ) RDS databases with failover setup. ECS for apps.
Databases looked after themselves. The senior people probably spent 20% of a FTE on stuff like optimizing it when load crept up.
Place before that was a similar size and no DBA either. People just muddled though.
My company uses redundant services because we like to deploy frequently, and our customers notice if our API breaks while the service is restarted. Running the service redundantly allows us to do rolling deploys while continuing to serve our API. It’s also saved us from downtime when a service encounters a weird code path and crashes.
We adopted TFC at the start of 2023 and it was problematic right from the start; stability issues, unforeseen limitations, and general jankiness. I have no regrets about moving us away from local execution, but Terraform Cloud was a terrible provider.
When they announced their pricing changes, the bill for our team of 5 engineers would have been roughly 20x, and more than hiring an engineer to literally sit there all day just running it manually. No idea what they’re thinking, apart from hoping their move away from open source would lock people in?
We ended up moving to Scalr, and although it hasn’t been a long time, I can’t speak highly enough of them so far. Support was amazing throughout our evaluation and migration, and where we’ve hit limits or blockers, they’ve worked with us to clear them very quickly.
I think the format of this is great. I suppose it would take a motivated individual to go around and ask people to essentially fill out a form like this to get that.
One suggestion if we're gonna standardize around this format. Avoid the double negatives. In some cases author says "avoided XYZ" and then the judgment was "no regrets". Too many layers for me to parse there. Instead, I suggest each section being the product that was used. If you regret that product, in the details is where you mention the product you should have used. Or you have another section for product ABC and you provide the context by saying "we adopted ABC after we abandoned XYZ".
I don't recommend trying to categorize into general areas like logging, postmortems, etc. Just do a top-level section for each product.
If anyone has newer posts like the above, please reply with links as I would love to read them.
https://world.hey.com/dhh/we-stand-to-save-7m-over-five-year...
https://world.hey.com/dhh/our-cloud-exit-has-already-yielded...
Related, looks like X is doing similar: https://twitter.com/XEng/status/1717754398410240018
Sounds like they experienced badly managed and badly constrained database. The described fks and relations: that's what the key constraints and other guard rails and cascades are for - so that you are able to manage a schema. That's exactly how you do it: add in new tables that reference old data.
I think the regret is actually not managing the database, and not so much about having a single database.
"database is used by everyone, it becomes cared for by no one". How about "database is used by everyone, it becomes cared for by everyone".
> Endorse-ish: Schema migration by Diff
Well that explains it... What a terrible approach to migrations for data integrity.
I never saw a successful fully automated one-way-of-doing process.
So every one needs to know every use case of that database? Seems very unlikely if there are multiple teams using same DB.
FKs? Unique constraints? Not null colums? If not added at the creation of the table they will never be added - the moment DB is part of a public API you cannot do a lot of things safely.
The only moment when you want to share DB is when you really need to squeeze every last bit of performance - and even then, you want to have one owner and severly limited user accounts (with white list of accessible views and stored procedures).
You don’t share a DB for performance reasons (rather the opposite), you do it to ensure data integrity and consistency.
And no, not everyone needs to know every use case. But every team needs to have someone who coordinates any overlapping schema concerns with the other teams. This needs to be managed, but it’s also not rocket science.
If DB is shared then data from different users is entered/updated through multiple transactions. So you cannot get anything better regarding consistency and integrity compared to multiple DBs and distributed TXs.
By introducing schema change coordination you will introduce enormous delays to almost any DB change. This is more realistic than everyone knowing each use case but less practical. Shared DB is an antipattern either way.
And by far "automate all the things" is probably my number one suggestion for DevOps folks. Something that saves you 10 minutes a day pays for itself in a month when you have a couple of hours available to diagnose and fix a bug that just showed up. (5 days a week X 4 weeks X 10 minutes = 200 minutes) The exponential effect of not having to do something is much larger than most people internalize (they will say, "This just takes me a couple of minutes to do." when in fact it takes 20 to 30 minutes to do and they have to do it repeatedly.)
Side node: There is a small typo repeated twice "Kuberentes"
> Not adopting an identity platform early on
The reason for not adopting an IDP early is because almost every vendor price gouges for SAML SSO integration. Would you say it's worth the cost even when you're a 3-5 person startup?
> Datadog
What would you recommend as an alternative? Cloudwatch? I love everything about Datadog, except for their pricing....
> Nginx load balancer for EKS ingress
Any reason for doing this instead of an Application Load Balancer? Or even HA Proxy?
Grafana Labs comes closest in terms of breadth but their DX is abysmal (I say this as a heavy grafana/prometheus user) Same comments about new relic though they have better dx than grafana. Chronosphere has some nice DX around prometheus based metrics but lack the full product suite. I could go on but essentially, all vendors either lack breadth, DX, or both.
The initial promise of "we'll take care of this for you, no in-house knowledge needed" has not materialized. For any non-trivial use case, all you do is replace transferrable, tailored knowledge with vendor-specific voodoo.
People who are serious about selling software-based services should do their own infrastructure.
No way. We used Terraform before and the code just got unreadable. Simple things like looping can get so complex. Abstraction via modules is really tedious and decreases visibility. CDKTF allowed us to reduce complexity drastically while keeping all the abstracted parts really visible. Best choice we ever made!
Many say in the database world, "use Postgres", or "use sqlite." Similarly there are those databases that are robust that no one has heard of, but are very limited like FoundationDB. Or things that are specialized and generally respected like Clickhouse.
What are the equivalents of above for Kubernetes?
Kubernetes is probably “use postgres”
It's just that, you should start with a handful of backed-up pet servers. Then manually automate their deployment when you need it. And only then go for a tool that abstracts the automated deployment when you need it.
But I fear the simplest option on the Kubernetes area is Kubernetes.
I shunned k8s for a long time because of the complexity, but the managed options are so much easier to use and deploy than pet servers that I can’t justify it any more. For anything other than truly trivial cases, IMO kubernetes or (or similar, like nomad) is easier than any alternative.
The stack I use is hosted Postgres and VKS from Vultr. It’s been rock solid for me, and the entire infrastructure can be stored in code.
I also stayed away for a long time due to all the fear spread here, after taking the leap, I’m not looking back.
The lightweight “simpler” alternative is docker-compose. I put simpler in quotes because once you factor in all the auxiliary software needed to operate the compose files in a professional way (IaC, Ansible, monitoring, auth, VM provisioning, ...), you will accumulate the same complexity yourself, only difference is you are doing it with tools that may be more familiar to what you are used to. Kubernetes gives you a single point of control plane for all this. Does it come with a learning curve? Yes, but once you get over it there is nothing inherent about it that makes it unnecessary complex. You don’t need autoscaler, replicasets and those more advanced features just because you are on k8s.
If you want to go even simpler, the clouds have offerings to just run a container, serverless, no fuzz around. I have to warn everyone though that using ACI on Azure was the biggest mistake of my career. Conceptually it sounds like a good idea but Azures execution of it is just a joke. Updating a very small container image taking upwards of 20-30 minutes, no logs on startup crashes, randomly stops serving traffic, bad integration with storage.
The industry has known this to be a stereotypically bad idea for generations now. It lead to things like the enterprise sevice bus, service-oriented architectures, and finally "micro services". Recently I've seen "micro services" that share the same database, so we've come full-circle.
Yet, every place I've worked was either laboring under a project to decouple two or more applications that were conjoined at the DB, or were still at the "this sucks but no one wants to fix it" stage.
How do we keep making this same mistake in industry?
https://fluxcd.io/blog/2022/11/flux-is-a-cncf-graduated-proj...
> Weaveworks will be closing its doors and shutting down commercial operations > Alexis Richardson, 5 Feb 2024
https://www.linkedin.com/posts/richardsonalexis_hi-everyone-...
If the project has legs, it's now under CNCF.
Is the project future at risk? https://github.com/fluxcd/flux2/discussions/4544
Even in a startup, it’s difficult to hire an expert in every platform that can maintain a robust, secure system. It’s possible, but not guaranteed, and may require a high pay to retain the right staff.
Many government agencies on the other hand are legally banned from offering a competitive wage, so they can literally never hire anyone that competent.
This cap on skill level means that if they do need reliable platforms, the only way they can get one is by paying 10x the real market rate for an over-priced cloud service.
These are the “whales” that are keeping the cloud vendors fat and happy.
Many of the decisions presented are not disagreeable (choosing slack) and some lack framing that clarifies the associated loss (Not adopting an identity platform early on). I think they're all good choices worth mentioned; I would have preferred a deeper look into the few that seemed easy and turned out to be hard, or the ones that were hard and got even harder.
It helps to hear the validation, although I think almost every decision has a dissenting voice in the HN comments.
There are some teams I work with that we'll never bother to make use Bazel because we know in advance that it would cripple them.
I wish my current company did this. It's infuriating. The other day, I asked a question about how to set something up, and a manager linked me to a channel where they'd discussed that very topic - but it was private, and apparently I don't warrant an invite, so instead I have to go bother some other engineers (one of whom is on vacation.)
Private channels should be for sensitive topics (legal, finance, etc) or for "cozy spaces" - a team should have a private channel that feels like their own area, but for things like projects and anything that should be searchable, please keep things public.
It'd be amazing if more people published similar articles and there was a way to cross-compare them. At the very least, I'm inspired to write a similar article.
I keep wondering when this is going to show up. We have a lot of service providers, but even more frameworks, and every vendor seems to have their own bespoke API.
That being said, Cloudflare is on the path to offering a great GPU FaaS system for inference.
I believe it’s still in beta, but it’s the most promising option at the moment.
So we got a $90k server with 184TB of raw storage (SAS SSD), 64 cores, and 1TB of memory. Put it on a 10GB line at our university and it is rock solid. We probably have less downtime than Github, even with reboots every few months.
Have some large (multi-TB) databases on it and web APIs for accessing the data. Would be hugely expensive in the cloud with, especially with egress costs.
You have to be comfortable sys-admining though. Fortunately I am.
I didn't understand this section. Ubuntu servers as dev environment, what do you mean? As in an environment to deploy things onto, or a way for developers to write code like with VSCode Remote?
Being able to write a bash script that runs on ever machine is nice.
As mentioned in the other comment, the most commonly used providers for terraform are "bridged" to pulumi, so the maturity is nearly identical to Terraform. I don't really use Pulumi's pre-built modules (crossroads), but I don't find I've ever missed them.
I really like both Pulumi and Terraform (which I also used in production for hundreds of modules for a few years), which it seems like isn't always a popular opinion on HN, but I have and you absolutely can run either tool in production just fine.
My slight preference is for Pulumi because I get slightly more willing assistance from devs on our team to reach in and change something in infra-land if they need to while working on app code.
We do still use some Pulumi and some Terraform, and they play really nicely together: https://transcend.io/blog/use-terraform-pulumi-together-migr...
Infrastructure should be declared, not coded.
Say what you want. The tool then builds that, or changes whats there to match.
I've tried Pulumi and understanding the bit that runs before it tries to do stuff and the bit that runs after it tries to do stuff and working out where the bugs are is a PITA. It lulls you into a false sense of security that you can refer to your own variables in code, but that doesn't get carried over to when it is actually running the plan on the cloud service (ie actually creating the infrastructure) because you can only refer to the outputs of other infrastructure.
CFN is too far in the other direction, primarily because it's completely invisible and hard to debug.
Terraform has enough programmability (eg for_each, for-expressions etc) that you can write "here is what I want and how the things link together" and terraform will work out how to do it.
The language is... sometimes painful, but it works.
The provider support is unmatched and the modules are of reasonable quality.
I understand, but I think they don’t have the luxury of not having a DBA. Data is important; it’s arguably more important than code. Someone needs to own thinking about data, whether it is stored in a hierarchical, navigation-based database such as a filesystem, a key-value stored like S3 (which, sure, can emulate a filesystem), or in a relational database. Or, for that matter, in vendor systems such as Google Workspace email accounts or Office365 OneDrive.
But if you only want to hire React developers (or swap for the framework of the week) then you'll likely end up with zero understanding of the DB. Down the line you have a mess with inconsistent or corrupted data that'll come back with a vengeance.
It's short-sighted for serious endeavors.
we have dotnet webapp deployed on Ubuntu and it leaves a lot to be desired. The package for .net6 from default repo didn't recognise other dotnet components installed, net8 is not even coming to 22.04 - you have to install from the ms repo. But that is not compatible with the default repo's package for net6 so you have to remove that first and faff around with exact versions to get it installed side by side...
At least I don't have to deal with rhel Why is renewing a dev subscription so clunky?!
If and when you start experiencing scaling problems (great!), that's the time to think about migrating to setting up infra.
Things largely look the same on the surface; this takes the most effect at the implementation-detail level, where adjusting and countercorrecting down the track is fiddly and uses an adrenally-draining level of attention span - right when you're at the point where you're scaling and you no longer have the time to deal with implementation detail level stuff.
You're on <platform> and you're doing things their way and pivoting the architecture will only be prioritised if the alternative would be bankruptcy.
It literally doesn't matter what service you're using at that point.
I don't see how you need to be "doing things their way" when that's all you have.
> Very intuitive to configure and has worked well with no issues. Highly recommend using it to create your Let’s Encrypt certificates for Kubernetes.
> The only downside is we sometimes have ANCIENT (SaaS problems am I right?) tech stack customers that don’t trust Let’s Encrypt, and you need to go get a paid cert for those.
Cert-manager allows you to use any CA you like including paid ones without automation.
Seems like yagni to me but please prove me wrong
[1] https://github.com/Azure/karpenter-provider-azure/pull/72
(source: On the team that is developing the provider)
I found this slightly ironic given there are ~50 headers in the article :)
I liked the format of the writeup
We have non-developers (artists, designers) on our team, and asking them to manage homebrew is a non-starter. We're also on windows.
We current just shove everything (and I mean everything) in perforce. Are there any better ways of distributing this for a small team?
Is it because most people are willing to pay someone else to manage monitoring infrastructure or other reasons?
something to keep in mind is that most companies are not like the folks in this thread. they might not have the expertise, time or bandwidth to build invest in observability.
the vast majority of companies just want something that basically works and doesn’t take a lot of training to use. I think of Datadog as the Apple of observability vendors - it doesn’t offer everything and there are real limitations (and price tags) for more precise use cases but in the general case, it just works (especially if you stay within its ecosystem)
This hits hard. Someone please take my (client's) money and provide sane GPU FaaS. Banana.dev is cool but not really enterprise ready. I wish there was a AWS/GCP/Azure analogue that the penny pinchers and MBAs in charge of procurement can get behind.
I wonder how much one should pay attention to future problems at the start of a startup versus "move fast and break things." Some of this stuff might just put you off finishing.
Any references to something like this with a Search slant would be greatly appreciated.
I have no first hand exerience wtih Okta, but everything I read about it makes me scared to use it. i.e. stability and security.
Both are more simple and do the same thing.
I think adding a DBA or hiring one to help you layout your database should not be considered a 'luxury'...
Maybe you need help with setup for a few weeks/months, and then some routine billable hours per month for maintenance / change advice.
Anyway, amazing write up.
Learning about alternatives to Jira is always good.
VPNs can be wonderful, and you can use use Tailscale or AWS VPN or OpenVPN or IPSEC and you can authenticate using Okta or GSuite or Auth0 or Keycloak or Authelia.
But since when is this Zero Trust? It takes a somewhat unusual firewall scheme to make a VPN do anything that I would seriously construe as Zero Trust, and getting authz on top of that is a real PITA.
What are they using for GPU bound services. Python?
No, just no. I see this cropping up now and then. Homebrew is unsafe for Linux, and is only recommended by Mac users that don't want to bother to learn about existing package management.