Fly.io Infra log: week-by-week record of what the team does
fly.io
fly.io
I really don’t understand the point of the second half in publicly disclosing what each individual has done though. I hate the idea that team members would feel like they might have to rush stuff out so they look as productive as someone else, or that they’d feel not as good as others.
It wasn’t clear if it’s opt-in to have your status update posted publicly but I can imagine (and hope) it’s okay not to partake.
Rather than asking infra developers to do somersaults to make their work fit the mold for the "Fresh Produce" topic, we're just going to do some of that writing for them. We're going to keep doing this, and I expect it's going to make more sense as the weeks move on.
We certainly didn't submit this to HN; I think it looks super weird here, especially with just the first week of content on it.
I personally liked the way it's done (very transparent and interesting at the same time) and thought HN crowd would be interested in these too ! :D
Weekly updates by naming and shaming every team member would cause me so much anxiety even if they were just internal.
I’ve been on SRE/infra teams my entire career.
- Is it remote?
- What is the update cadence?
- Is this information shared with Manager and "manager communicates team status"?
That's not to say we don't have internal dysfunctions, or that we don't do things that annoy engineers. Just that that our dysfunctions are different.
Weekly feels right to me. We'll see how it goes.
My current experience is that once we started transition to SOC 2 compliance and brought in new processes everything crawled to a halt, we now have multiple layers of reviews and scheduling, and the bureaucratic process is followed at the cost of expediency and doing sensible things in a sensible way (i.e., everything must follow the process, no matter how trivial or significant.)
The very most important thing to remember about SOC2 is that your auditors are attesting that you meet a standard that you yourself set. If you set a standard for yourself that every single change anywhere is going to be reviewed in triplicate, that's what auditors are going to check you on. So the key is to be extremely deliberate about the standard you're setting in your initial Type I. Every auditor is going to want to see some kind of change review, but the purpose of that change review is something you determine, not them. For us, the SOC2-auditable change review process is about ensuring that people merging PRs aren't 3 raccoons in a trenchcoat; it's not a vulnerability management process.
The trap people fall into is that SOC2 is often the point at which they start paying attention to security process as a whole, and they led SOC2 lead security process for them. No! Death! Be thinking about security from day 1, and have clarity about what subset of your security process you want SOC2 to measure and track.
fly has a lot of issues with their platform.
from my perspective this reads like what I would usually do as a part of a development team to the strategic teams when things does not work.
the only issues is that I am a customer at fly, not a strategic relation - we simply move on the other platforms.
We only use fly for a dev / testing environment. Production and staging run on a more mature platform.
* something about infra after all, something to learn from, dm-clone - TIL
* largest team and just 6 persons...whoa!
* on "A Registry machine unexpectedly reached a storage limit, disrupting deploys that pushed to that registry for about 20 minutes." - quick reaction assuming it required some manual steps to do after issue has been detected
* naming is simple "infra engineers", not SRE, not DevOps, not DevOps Developer, not even plain Sysadmins
Added blog to my RSS reader, should be useful and fun
Edit: to be clear, I mean ping from my house, not another data center.
Like many others though I cannot live with the fly unbounded pricing model or difficulties with data persistence, otherwise I would absolutely have been there a year ago.
EDIT: In fairness to fly it appears they have updated both their pricing plans and data persistence situation since I was making these decisions.
> These reports criticize OVHcloud for having no fire prevention system and no power cut-off on the site, for using wooden floors, and for a free-cooling design that created airflows that spread the fire. The reports also say that water was detected near electrical systems before the fire broke out.
OVH lied to some customers about which data center held their backups, keeping them in the same data center as the source machines. Both their own backups offering, and when a customer managed their own backups and explicitly copied the data to another DC.
Oh, and OVH almost recovered a customer's data, but accidentally deleted it instead.
https://www.datacenterdynamics.com/en/news/ovhcloud-ordered-...
I would not want my name on that page if I worked there. Tech companies are weird.
The top portion of the page is great.
Would be interesting to know, for what exactly Fly is using OVH.