When costs are nonlinear, keep it small
jessitron.com
jessitron.com
For example, AWS free tier is a great bargain. AWS is a terrible value once you’re paying the price of a bare metal box each year (or sooner!), but then at Netflix scale it balances out again because Amazon knows that Netflix could easily talk Microsoft into doing the port themselves if it would actually make financial sense.
Most pricing is S-shaped, so it typically is best to be in the position of exploiting a loss leader or running at scale.
I’d imagine that anyone spending real money is always thinking of ways to get off of something like this?
It took work over several weeks to convince them that that was a ripoff. I think their business people eventually called up cloudinary and negotiated a better rate, but by that point I was annoyed by the whole situation. The site used cloudinary's HTTP based API from the browser. I configured cloudflare to redirect any request hitting example.com/images/... to act as a caching proxy to cloudinary's actual servers. Unsurprisingly, that one configuration change made their cloudinary bill drop to about 1/20th what it was. When the client was billed we got a panicked phone call asking why it was so low, and if something was broken on their site.
Anyway, tldr; lots of folks out there have no idea what a service like cloudinary should cost. Apparently more than enough to make cloudinary a profitable business.
You're probably off by about 2 orders of magnitude.
If your AWS spend is $10K per year or below, it simply isn't worth thinking about in terms of engineering time if you have real revenue.
Once your AWS spend crosses about $100K/year, it's time to pay attention. However, you probably want to be stingy about this as your engineer is in that range. Your engineers are supposed to be providing business value and looking at $10K/year costs probably isn't worth their time.
At the $1M/year mark, you already have an engineer on this full-time anyway so it makes sense to have some hardware you actually own in a colocation facility somewhere.
Its also easy to accidentally lock yourself in to bad infrastructure decisions early, that become extremely expensive to fix later. Mongodb might seem like a good place to start your project, but trust me - moving away from it down the road will be a nightmare.
Too much up front design will kill your velocity. And not enough up front design will cost you too. If there's a good rule of thumb here, I haven't found it yet.
Sure. Presumably, though, your engineering salary for a couple months had more business value than you would have saved, so it's still a net plus.
> Too much up front design will kill your velocity. And not enough up front design will cost you too. If there's a good rule of thumb here, I haven't found it yet.
In my opinion, err on the side of velocity. Most startups are worrying about staying alive.
If the startup survives to have the problem of ripping out bad decisions, consider that your engineering decisions did their job and chalk it up to the scars of battle.
"Good judgement comes from experience. Experience comes from bad judgement."
do you mean that for 100K/year you DON'T HAVE and engineer managing your account infra? Who does this in this case? CEO? This volume of spending implies quite an infra in AWS.
Is this generally true? Deploy as in going to production?
I join a web app project that didn't have any tests and the code was brittle, where changing one part of the app could break a completely different part without you realising. It was much more efficient (and less stressful) to batch up lots of small changes and do one thorough manual testing session of the whole thing while we gradually got the codebase under control.
It likely depends on how much automated testing you have, how bad it is if bugs go live and how quickly you need features live.
While manual testing changes the equation some (the testing effort is the same for each release; the more you can automate that, the greater gains), how easy to -fix- the issue becomes more complicated the more changes there are. 1 PR, you know what introduced the breakage. 30 far reaching ones, with interesting overlap between them? Good luck.
To be clearer, I mean there would be say 30 pull requests merged into "develop" and they would only be merged into "master" and pushed live after detailed manual testing. This is versus merging each pull request directly into "master" and going live immediately, where you wouldn't have time to do detailed manual testing every merge.
You would still have the git history in both approaches to debug.
You either test after each PR merge, or you test only when you 'craft a release'. Doesn't matter whether this merge is into a develop branch or straight into master; it's solely a question of when you test it.
If you manually test, testing only a batch of changes, together, is obviously easier, from a testing perspective. However, the effort to figure out and fix the issue can be very large. Plus, after fixing that one issue, you have to retest everything. So every bug requires retesting in any case (so while it's still a lower testing burden, unless you introduced 30+ unrelated issues, it's not as low as you think it might be on the surface).
Compared with testing (and at that point you might as well release if your pipeline and org let you) with each PR - any manual testing effort is obviously higher due to testing so often (whereas automated is no extra effort), but, figuring out what caused the breakage, and addressing it, is much, much easier, order of magnitude easier, to figure out, and fixing it is much less likely to introduce new issues.
Which is more to the original point; catching a bug earlier, allows the fix to be more targeted, which makes it more likely to be right, which reduces the likelihood of it making it to prod.
There is also, to the original point, a statistical thing to consider.
Let's say in 30 PRs, there is one that introduces a bug. Well, if QA is testing the one PR that broke something, they're more likely to spot it, since they're focusing on that PR. They're less likely if it's one of thirty sets of changes. But beyond that, let's say, in both situations, QA misses it. If you have one deploy, you've broken prod 100% of your releases (1 of 1 release). If you have 30 deploys, you've broken prod only ~3% of your releases. So immediately the original statement is validated; if your testing effort is the same (big if!), more frequent releases, with smaller changes, means the same or better percentage of deploying working code. Also, if you have to rollback with one mega release, you lose all 29 other changes; if you have to rollback with separate changes, you lose no other changes.
In addition this completely ignores that if you spend the same amount of time testing your testing per release drops when you increase the number of deployments.
No, you just automate your testing.
This does explain why it didn't work for you though.
CI/CD lets you iterate faster and it encourages best practices like proper testing. That said, if you do CI/CD without proper testing then you're going to break production.
It can also be a tough sell to make process improvements like this when the old way got them this far and they want new features now. I've worked with clients that don't even use source control, do all edits on production and there's no way setup to develop locally.
It would be great if HN had some content on «moving from random files on a network drive to git and basic automated testing» - it’s not an easy job and others would benefit from such writeups.
Did you miss a "not" here? For me an environment where my changes are deployed and used quickly is a lot more fulfilling than one where I work on something that may only see the light of day weeks or months down the line.
You think waiting days would be hateful. Having to wait months is more frustrating. And more risk for the business.
If you start with a project where code is brittle and there are no tests, it is indeed extremely hard and time consuming to move to CI/CD. Doesn't mean it's not a good idea, just that it won't be easy and it'll take time (you have to build the whole infrastructure of automated testing, canary deployments, etc. - but, it's a good idea to do that anyway! It'll gradually improve the velocity. The alternative is that you end up in a world where everybody is afraid of making changes and the simplest requirement ends up getting estimates like "3 months of work".
It sounds like a really bad idea. It's like saying: When juggling, it's best to start with knifes, because every mistake will mean you're going to cut yourself, and you will quickly learn not to make mistakes.
On the other hand when a process is run infrequently you tend to get surprised each time because something has changed or broken since the last time. Or you just simply haven't done the thing enough to observe the possible issues so every six months you discover and re-discover the problems.
Batching up many changes and features over longer and longer release periods brings so many issues, and definitely greater risks when actually deployed as the scale of change is larger.
What the article perhaps doesn't pick up on so well is that poor or difficult deployment processes lead to this behaviour of batching and putting-off the deployments. Do them less often, have less periods of service impact. Is the logic.
Which of course is one of the key points about DevOps, feedback loops, and CI/CD processes to simplify things.
If you can make deployment easier, you can do them more often. Making the changes smaller can help ease deployments, so can be a virtuous circle. Though not if deployments are not reliable or repeatable.
When I came onboard this project, I came into a situation where the client simply did not upgrade their libraries & nodejs version for a few years. When it came time to catch up with the latest version of Angular, the experience was painful. Three major version updates later, (version 9.x to 10.x), the Angular upgrade became unbearable & I moved the project over to Svelte & a pnpm monorepo. Now all dependencies are up to date, the architecture is improved, & page size reduced. It took a few months to hammer out all of the edge cases due to the complexity of the app but release is imminent.
"The old way was all over the place so instead of improving the existing thing, I completely rewrote everything from scratch and now everything is cool. Yay me/us!".
Why couldn't you just upgrade the dependencies once then set up the same CI/CD you're presumably using for Svelte so that you can them upgrade versions easily?
In software, you need to have code reviews, unit and regression testing, or you can fix one minor bug only to introduce a catastrophic fail.
Sublinear costs seem like something that you'd be best off delaying to fix, maybe forever, since their impact lessens the longer you wait.
Maybe this would apply when a major change will happen that would obviate the need for a fix: a planned building teardown causes a needed roof repair on the old building to have sublinear costs, or introducing a new subsystem that eliminates the old subsystem that had the outstanding repair.
Or aesthetics: a minor marring of a building facade matters when it's pristine, but if you wait longer, the more other minor marrings appear, the less that first mar individually matters to the value of the building.
If you need to make changes to the design of an injection molded part, you'd better make them all at once, because the cost of a new mold is $100k whether you make one change or seven.
Which is what economies of scale mean, just a huge range of sub linear costs all lumped together.
Except in the real world this is being optimized based on a forecast at every point in the day and every day of the year. You might not think of demand in those terms, but it’s a huge area of optimization at both large companies like Walmart all the way down to individual restaurants.
To see these opportunities you need data, but it is the most closely guarded secret of companies.
In the short term it gives them an advantage, but in the long run society as a whole could benefit tremendously from open data sharing.
Am I wrong about this?
Examples of sublinear costs:
- When you pay lower per-unit price by ordering large lots. But you can't just wait a long time and be able to afford a giant order; your first small orders may help you generate revenue that let you buy the later big orders.
- When your team is young perhaps you need 1 new laptop per new hire. A few years later, you need < 1 new laptop per new hire b/c some receive recycled laptops from departed employees. But you can't wait 3 years and then hire 100 people and 80 laptops.
- You're growing some infrastructure which serves/covers some territory. At first, all new customers are in new territory; the ratio of new infrastructure to new customers is high. Later, some portion of new customers are covered by existing infrastructure, and that portion grows over time. But even if you could afford to build everything at once, you may not know where to build most of it until you have a bunch of customers.
- I don't really know if this one is true, but a whole industry seems to believe that a burst of advertising effort all at once is more effective than a marginally greater volume of advertising spread over a longer period. But you can't do zero advertising for 5 years and then take over time square.