Grafana: Why observability needs FinOps, and vice versa
grafana.com
grafana.com
You can certainly spend more, but on-prem/self-hosted, time is usually the limiting factor, either directly or through opportunity cost. Contrary to popular belief, time does not always equal money, if you need more storage and the blocker is that you need to build out an S3 equivalent (rather than just paying for S3), then you'll be blocked by hiring, by hardware lead times, etc.
I also don't think that every organisation that needs file storage must build a storage solution that should compete the reliabiltiy and features of S3. Most of the times you can get by just fine at fraction of the cost.
Take file storage for example. Going from one server, to one with backups, to N, to big-N, are all points of inflection where significant engineering is required for on-prem/self-hosted file storage. With a cloud solution none of these are inflection points, none require additional work, they only require additional money.
Assuming you have infinite time, you can just funnel money into things like hardware upgrades and hiring engineers to build these things, but if you don't assume infinite time, time is often as strong or even stronger a factor than money.
At my last company we had a bunch of servers in colo, and could not throw money at solving problems there. Getting a new machine took 2+ days and a bunch of emails, not an API call. We moved to the cloud mostly because the opportunity cost, i.e. the time spent by engineers on toil scaling things on physical machines, was higher than the monetary cost that we could pay on a cloud provider.
This won't be the same for everyone, but the point is that money is roughly the only consideration in cloud, but not the only consideration on-prem, at least when you discount common factors between the two.
But sometimes you just need more compute and you are the type of organization that buys compute by the floor space and power consumption...
Regardless of the nuances of each situation, I think jsiepkes's comment meant to say that in the data center you can buy pretty killer hardware that will be totally overkill for the moment and won't require you to count active timeseries in order to not pay $300k a month for your metrics, and at the same time will last you for the next couple of years.
Also, for most companies, the next point of inflection will never come and this server will probably last them for a very, very long time.
I'm sharing my point of view as someone who works at an organization that took money as the only consideration and managed to grow over the years to now having to start taking both time and money into consideration because taking only money into consideration proves to be too expensive.
We recently switched from Grafana to Prometheus. Reason being that a license refresh took longer to process on their end. What happens when a license expires on Grafana? They fucking shut down all your shit cold turkey. Don't care if you're in prod or have a dedicated guy on their end for support or whatever. So you're happily churning along and then suddenly you're blind. Nice FinOps. With Prometheus there's a grace period where they'll happily overcharge you. But we've never had a product absurdly blow up on us like this before. It's truly mind boggling that they're out here talking about 'FinOps' now.
Datadog is expensive, but at least we were only making these decisions for the ~hundreds of custom business metrics, and not the ~tens of thousands of metrics from our infrastructure.
In my previous company I had a good setup for costs monitoring - including release to release comparisons, drill downs, statistics, etc.
After each release I looked at this data. It saved a lot of $, by simple fixes like "why we are calling this API twice?".
It also quite some issues that weren't strictly customer related, but weren't apparent from other type of data (you will always have some "unknown unknowns" in your monitoring, and costs data seem to be pretty wide net to catch some of those)
The game here is defining parts of the job away to be someone else's job or responsibility.
If it’s not got a familiar marketing thing going on, they’ll refuse to acknowledge it even if their devs are practically begging for it.
Or you change to a company with capable & technical middle management :)
It always blows my mind when people say "management does not have to be technical".
Later, FinOps role can evolve but expectations will mostly be reactive.
Lets see how long that is cheaper than just hiring 2-3 finops folk and just put them in every room where the software architecture is being designed for new services and make them drill down hard into team on what to avoid.
Not to mention it’s a better way to do things in the Single Responsibility Principle that most great teams follow
if everyone is responsible for cutting costs and optimizing.
Then no one is….
A high margin product might want to launch quickly and reliably, costs be damned. Another product may need to run as cheap as possible, even if it's unreliable and takes a long time to develop.
FinOps translates business goals into technical requirements. It's not just cutting everything down to as cheap as it can be.
How would you manage your cloud costs if you ran a company of, say, 4,000 engineers? Balancing the needs of delivery teams to build their technology with the needs of the business to manage costs. Do you think every single team should directly report their cloud costs to the CFO? Or at that scale does it make more sense to report costs to another individual? And when that needs to scale, maybe we give that individual a team?
At this point I'm waiting for someone to start flogging "CodeOps".
somehow i misread that as "massage your cloud costs", and it.. sticks..
I kind of agree that I don't love the term. That being said it has become the de-facto way that people refer to the space and practice. The FinOps Foundation hyped up the term and space quite a bit which they deserve credit for but do wish there was a better name :)
FinOps isn't a dedicated job but something a cloud engineer can do as part of its job. In the same way that DevOps doesn't need to be a dedicated function itself.
And as for the cloud... Yes that turned out to be a whole lot expensive for companies than predicted - and you share compute with anyone so that dualcore CPU isn't always that fast. But cloud is also flexible and that is where FinOps comes in.
Nothing to do with YC AFAIK.