Scaling SQLite to 4M QPS on a Single Server (EC2 vs. Bare Metal)
blog.expensify.com
blog.expensify.com
Also, it's already been submitted a bunch of times:
https://news.ycombinator.com/item?id=23070888
https://news.ycombinator.com/item?id=22856746
What about:
Networking equipment
Backup power - UPS and generators
Internet backbone - Speed as well as reliability
Cooling and other infrastructure costs
Then there is capital lock in risk. Once you buy the above hardware, you are probably stuck with it for a few years.If you have stable load now for 25 big bare-metal servers, it may be cheaper to go bare metal. Otherwise, there is a lot more involved in the decision than going to dell.com and pricing a server.
You can get bare metal servers from a host on monthly or hourly terms.
You can get a 1/3rd rack cabinet at a colo facility and bring your own router, and use the colo's IPs
You can get a cage or a room at a colo and arrange for connecticity to your various upstreams and peers with your own ASN and IPs.
Only after all this is exhausted would I move towards and owned and operated datacenter where all the pieces are under company control and responsibility. (Although, if you rely on it, you're responsible for it, regardless of if you actually control it)
As far as reliability, you can't compare almost any COLO location with an AWS data center.
Why not? We had much less downtime than AWS' NoVa DC while coloing, and I don't think our DC was exceptional. AWS is probably much more competent, but they're also trying to build something much more complex than your average small-scale colo setup.
There's no super secret reliability technology. You can absolutely compare colo with AWS. Many colos may be less reliable, and some may be more reliable. You need to look at power redundancy: how many utility feeds, what kind of ups, generators, testing schedules, what the points of failur are, how they responded in the past, etc. Same for networking feeds. Most colos have simpler networks than AWS which reduces the possible failures, but they may be less redundant etc; they probably have better bandwidth prices though.
Anyway, if you want a reliable system, you're going to need to be in multiple locations, and if you're in multiple locations, you should be able to weather the inevitable power relay failures and border router failures, etc every couple of years.
There are plenty of hosting providers that will provide everything you mention plus the actual servers for a monthly fee without requiring you to buy anything.
There is no calculus that makes aws cheaper unless you need extreme variability or some of their managed services and even then, you can usually use those services in conjunction with your rented bare metals.
Trying to be helpful and not pedantic: "On premises" or "on prem" are good choices here, while "premise" means something else (e.g. "The premise of moving off-premises is that we can pay a little more in order not to worry about stuff that isn't our company's primary business.")
Very common, for sure. And although I think it's worth professionals being aware of the difference, I tend not to bother with less technical folks.
> I wonder if 20 years from now the word premise would have a different meaning.
Given that "literally" is now often used to mean "figuratively", it's absolutely possible!
The only reasons I've seen to justify putting something in the cloud are:
You're tiny - colocation is fine but generally speaking if you don't need more than a single-server's worth of performance and you've got a for-profit business, "the cloud" is probably easier.
Your workload is extremely bursty.
You're using the other stuff. AWS/Azure/GCP has all sort of services that require scale to be efficient - if you're subscribing to those services and are medium-sized (I know, that's not even loosely defined) - it probably makes sense to be in AWS/Azure/GCP.
If you don't need those services, or you're extremely large, you can almost always do it cheaper yourself. Whether you want to is another/completely different question.
The time from ordering to login is something between 15 minutes and 2-3 days (you know beforehand). You can cancel every month. Faulty disks are replaced 24/7 in 15-30 minutes.
Turns out this is available stock right now as the "unix-none" VFS implementation. We're using it in an almost-but-not-quite-POSIX RTOS that happens to be missing advisory locks.
See also https://sqlite.org/vfs.html
Cloud VM vs on-prem bare metal doesn't seem like a very fair comparison
> WHERE indexedColumn > RANDOM() LIMIT 10
That can be bound on RANDOM() performance or we can rely on a seeded "non-random" PRNG to produce a random selection vector without running random() for every row.
If that has a syscall into /dev/urandom for every row, then that sucks for sure & is not representative of the actual intention.
In Apache Hive, we have a similar argument about unix_timestamp(), which was made deterministic, but stateful (MapReduce failure tolerance demands that if a task fails while running a query like > unix_timestamp() + n, the next attempt will produce an identical output, even-though some time has passed between the attempts).
So for that, we do the unix_timestamp() replacement per-query, instead of per-row.
I assume that's what was done here.
At the end of the day, we want operating expenses, not capital expenses.
And if you purchase a ton of bare metal and suddenly your business needs change it will not be easy or quick to move that much metal or install new metal in place.
Stick to the cloud.
2 years ago we build thumbor-like apps in python using vibora and opencv. A single aws m8 large($80/month?) can serve hundred millions images resizing in a month.
At that time we build it with only 2 people, copy pasting code from internet in a few days and vibora’s github readme also state that its in alpha release. With $2K server, you need to cheat to make those inefficient implementation :)
Side note: Several weeks later I found out about thumbor and replace our ‘emergency work’ with dockerized thumbor
I also tinkered with thumbot in aws lambda. But the cost raised to $500/month. I think we did something wrong with aws rekognition.
There's plenty of hidden costs though of doing this yourself that need a certain level of scale to pay off. Like people to manage the hardware, that's a big one.
Still renting bare metal servers from e.g. Packet could make much more sense even for small companies than using AWS.
No. At one time over 10 years ago, I had a dedicated server with one of the hosting companies. They charged me $225/month.
I moved to AWS and my cost dropped to ~$30/month. Even now, 10 years later, my total AWS bill for primarily EC2 and RDS services that runs an application that pays my bills is ~$75 per month.
If my company suddenly went viral and had a bunch of new customers, I could scale up instantly.
Ultimately it's all capex. The choice is whether the capital asset is on your balance sheet or someone else's.
If someone else is tying up their capital in a depreciating and inflexible asset, they'll demand to be paid well for giving you the flexibility to avoid all of that.
So you can't really say opex is always better than capex, since it depends on the price you're willing to pay for that flexibility.