What does Unsplash cost?
backstage.crew.co
backstage.crew.co
At this scale Heroku is always going to be more expensive than AWS or Digital Ocean. What happens when the Heroku bill hits $10K/mo?
Once they hit the next S3 bandwidth tier (350TB), Imgix becomes roughly 50% more expensive than just putting images on S3.
Everyone has their own heuristic for these decisions, but once you get in the realm of "1 or 2 people's entire salaries of savings" that's when I start really thinking hard about changing infrastructure providers.
> $18k is a lot of money to spend each month. Understanding the scale of Unsplash though can help explain the costs.
But does it justify the cost? Are you able to attribute new business from Unsplash? Does it bring in enough money to cover the $18K AND the salaries of the people who mantain it AND the opportunity cost of what they could be working on?
Appreciate your thoughts.
This is stuff we've absolutely thought hard about. Believe me when I say that I would love to cut the costs by switching providers for certain things, but the tradeoffs at our current size aren't worth it in our opinion.
I'm working on another post which outlines the tradeoffs of the different services and why we chose the ones that we use. That's probably the missing piece here — we've thrown out a lot of numbers but didn't give the reasons behind it (we wanted to keep the article focused and short).
Re justify-ing the cost, we absolutely can justify the cost. We've thought about all of the things that we spend time on and we wouldn't do something unless we thought it was the best use of our resources.
We tried Imgix on a few other products where we had tens of thousands of images being uploaded each month and only seen a couple hundred times per image. That became prohibitively expensive — which is probably a similar situation that you ended up in (high number of `master` images to render ratio).
In addition to realtime resizing, we use:
- face detection - typesetting - overlays - cropping/point of interest cropping - color palette - exif/image metadata - client hints - automatic content negotiation
Pretty much everything except their watermarking endpoint haha ;)
For me the killer feature is being able to crop images such that a detected face within the image is centered.
[1]: http://thumbor.org/
1. At any scale pretty much AWS / GCE are less expensive than Heroku. They also come at an non-zero operational cost. Say, switching to Google App Engine (the cheapest from my understanding) -- would save them 75% of their bill (just back of the napkining). -- That's not much, compared to the cost of having someone migrate the platform over (multiple work-weeks)
2. S3 != Imgix. Imgix is a CDN + Image resizing service. At a couple employers ago we tried to run our own image resizing service in house, and it was a royal pain in the ass. I can go into this if you're interested.
The other thing is S3. My current employer uses S3 pretty heavily, and S3 is really not meant for end-user data access it seems. It suffers from lots of latency spikes, and a non-zero error rate. Slow, and unreliable sites make your users go away.
Ok, well, S3 + CloudFront just about is. Yes, I understand Imgix has other features like on-demand resizing. But it looks like Unsplash is a rails app–in which case a library like Refile [0] can handle that.
When deciding on making any changes you have to consider the cost/benefit. Say an engineer costs you $10k for a month (fully loaded cost). For them to spend a month on reducing costs then what kind of reduction makes it worthwhile? For me, I would want it to take a year or less to pay back. So for me, a developer would need to reduce the monthly bill by $1k in order to justify the effort of a month of their time. Maybe they have ideas to easily do that, or maybe not.
Over time, as the bill keeps rising, you eventually reach a point where the savings become greater than the developer cost and so you go ahead and do it. So just work out your numbers (engineer cost, minimum payback period) and the decision becomes easy.
One thing we also consider as well is the overhead of adding another person to our team. We're learning a lot as we go (we've never built anything like this), so we have to be very careful that each teammate we bring on is aligned on vision, has the tools and resources they need to be successful and make smart autonomous decisions, and can fit into the current team without disrupting too much of the other teammates. That means that we can't just double our team size overnight and stick two new people on optimizations and cost reduction, even if it financially and procedurally made sense.
By hosting on Heroku (or any other cloud hosting provider) you're saving on devops, but you're also paying over the odds for your hosting. However, if you had an option that retained most of the devops simplicity of Heroku, but also cut costs, evaluating this would make sense.
I would suggest that this option exists, and that option is application containers (Docker, etc...). If you can build your infrastructure around application containers, not only do you have the option of using dedicated hosting when you have access to the processing capacity to do so (and thereby save money), but you also have option to scale into the cloud if/when the local capacity is exceeded. A number of the cloud hosts support Docker, including but not limited to OpenShift:
https://blog.openshift.com/openshift-v3-platform-combines-do...
Furthermore, it's a change you can make gradually. Developers can start using application containers as development environments, and you can roll them out more broadly once the implementation issues have been ironed out.
Does this sound like something you would consider?
EDIT: Worth noting that Heroku also supports deploying using Docker containers:
At your scale, there are CDN providers that can get down to $0.05/GB or lower. That would cut your bill by $3500/month...roughly 20% lower than current. I assume moving your legacy CDN stuff at the same time might add to that savings.
Not sure how tightly bound imgix is to fastly, if fastly has some feature that's not available elsewhere, or if imgix would even be willing to entertain support for alternative CDNs.
Edit: Also curious if you've ever done any analysis to see if bots are using a significant portion of your bandwidth. High res images sounds like a popular target for leechy bots...maybe some savings to be had in detecting/blocking that?
132.2 TB of bandwidth through them is 1,712.196 per month---6 pops EU/US.
Is it as good as Cloudfront? No. CF has 32+ PoPs.
But you really need to question "do I need < 5ms latency to everywhere on the planet including Asia (huge costs) and Australia (insane costs) for my free service?"
I hope you think about this seriously.
Though if you think you need a service comparable to Cloudfront---contact High winds or even Akamai, they're much easier to negotiate with than CloudFront, and you will save money.
---
Aside: I run a media group with massive bandwidth requirements (1PB+ per site) so I've churned through quite a few "enterprise level" CDNs while building up traffic over the past 10 years.
You can't ignore an entire affluent continent and expect to be taken seriously. Australian bandwidth has dropped massively in cost since you last took a look at it, so there are no excuses.
For instance, our network servers 45% US/Canada traffic, 10% UK traffic, 15% German traffic, and 20% Korean.
Australian? That's 1% of 1%. We get more traffic from New Zealand than Australia.
Is there more of these types of blogs?
I always wonder how much companies / startups are spending for their infrastructure.
"Well, you can run some kind of little site/app for free or around $5-10 month. So, I imagine a much larger and more serious app would be like... $50/month? At most?"
It seems right. I don't blame them. We know it's not true.
"My nephew made a Tetris game in a week, so you should be able to make a billing system in... I dunno, two weeks? It doesn't even need sound effects!"
The only time I hit it on the head, is when I'm doing a project, the type of which, I've done before. And I tend to avoid those.
The problem is as you scale the lack of optimization scales with it. Ultimately a late optimization and migration affect more users, complexity etc..
I can't wait for the next post to hear about the details of the choices that have been made.