Also forgot to mention - for compute providers, other storage, bandwidth, etc look at their bandwidth alliance. Free bandwidth to/from any of the members.
How can I get a message queue on Cloudflare and have hundreds of servers read and write to it?
How can I get run a Lucene search index on Cloudflare and search against it?
Also R2 is still in beta development
However, as I said the vast majority of applications don’t involve message queues for hundreds of servers or Lucene. Many not even SQL.
I also pointed to compute and bandwidth alliance members that augment Cloudflare very well so you can build whatever you want while still taking advantage of the Cloudflare edge.
Or put Cloudflare as a CDN or augment with what they do have in front of whatever cloud you want. Azure and GCP are limited members of the alliance that offer reduced pricing egress to Cloudflare.
How hard is for them to build something similar to https://calculator.aws ?
Plus you get better disk access speed (NVMe) than something like an AWS EBS-backed instance. You can also order new instances within a matter of minutes.
You don't have to use them, but being aware of what is available there would be a really good idea. I think some workloads could easily be deployed there at a fraction of the cost of a cloud provider.
Those are considerably cheaper, suffer less from unpredictable performance, a.k.a. "loud neighbour" problem, and are less likely to make you lock into some API/feature that would become a huge headache if you ever decide to move.
The by far biggest cable goes from Finland to Germany. No idea what happens after that. I believe Hetzner uses it and I have never heard that they wouldn't have enough bandwidth. But I have neither inside knowledge nor do I use them.
Being in Finland myself I have better latencies from real cheap Scaleway in Paris than from far more expensive AWS in Frankfurt. So it's not always you get what you pay for. Still no drops from either.
Maybe Hetzner has more noisy neighbor issues because they are cheaper? The same could hold for Scaleway, but I have never noticed anything.
Packet loss can have numerous causes, not having enough bandwidth is one but faulty or misconfigured gear or misconfigured routing can also cause issues.
> Maybe Hetzner has more noisy neighbor issues because they are cheaper? The same could hold for Scaleway, but I have never noticed anything.
Yeah, in general, cheaper virtual servers are more likely to have these problems, just like cheaper bandwidth is more likely to be oversold. Low-cost providers can push it pretty far if most of their customers are hosting low-to-moderate-traffic websites or other HTTP services, which are pretty forgiving of steal time, low-level packet loss & latency.
I use scaleway privately only but very happy with them. I wish I had their kind of support from the vendors at work. You can often try out upcoming services for free, too.
Hetzner wants a copy of your ID which I was not willing to provide so I couldn't try them. Germans are normally really privacy-conscious so this surprised me. This was > 10 years ago though so perhaps they stopped this practice.
Heroku - For a more "abstracted" cloud that just works
Vultr/DigitalOcean - Cheap VMs, additionaly some services like managed DB and K8s
Hetzner/OVH - Cheap dedicated hosts
Netlify/Cloudflare Pages/etc - for static websites + functions/3rd party services for dynamic stuff
please don't.
I have nothing but the worst experiences with Oracle Cloud. Granted, I was always a little biased against Oracle as a company; but I have a friend working on OCI and assumed that since what I needed was really basic it would be fine.
It really wasn't, it was awful, confusing and opaquely and egregiously expensive.
That is true, but it is also very much an understatement of just how miserable every step of the OC experience is.
I was once forced by a customer to endure the torture that is OC, and the only thing that saved me was when Oracle's terrible management software just deleted our entire account and deployment one weekend. Gone without a trace, could not even log in anymore. Luckily, that was enough to persuade the client to move to Azure.
I've had previous clients whos entire account with production servers just get nuked without even so much as a curtsey email and OCI support WILL NOT HELP.
To be fair, I have been testing Oracle Cloud Infrastructure(OCI) on a paid account for about 2 years now and actively monitor server hosting forums.
During that time I have had no "abnormal" problems with paid production account.
However, it is well known that Oracle accidentally nuked free/trial accounts early on, and again recently accidental nuked free/trial accounts reserved public IPs. However, I know of no such events for paid accounts, and paid accounts on OCI should be production ready quality in my experience.
On the other hand support really is somewhat bad, but support is free and the cloud costs are much less than AWS, while having better products than similar priced cloud products such as Lnode, Vultr, Digital ocean.
You can bet your bottom dollar I trial things on lower tiers, and much later on I may very well take the things I learned into production. My production budget isn't as big as some, but its still a tidy 6 figure sum each year just for cloud hosted services, with no licences or wages in that figure. If my trial on your platform goes tips up without notice, I'm going to likely conclude you're not a particularly safe place for my production stack for whatever it is I'm working on. You should also count on me telling my contemporaries about my experience too.
The basic idea is simple, "give developers free personal resources so they might be inclined to start using your cloud for work related thing". Instead, the "free account" is basically just a guided tour on OCI incompetence before it boots you out onto the street.
On the surface they have given it some thought with setting up the system perfectly with the whole, "if your account is free you can never be charged thing" but they can't seem to stomach people actually using their "free account".
Out of the frying pan into the fire, eh?
Did you check out hetzner "cloud"? 2c/2g seems to be roughly 5$/Month, x3 would be 15$
But I guess you need to think about your backup plans anyway (depending on what service you run). So using a less reliable provider might just mean that you practice your backup plan slightly more frequently. Some OVH customers can tell you the story when one of their data centers burned down.
You get internal networking out of the box with per-app-region internal DNS entries, and with their CLI tool it's one command to SSH into your container or connect your laptop to the internal network.
As a nice bonus, their pricing is very competitive. You pay about as much as you would for the same size Linode or droplet (minimum is 256 MB RAM for $1.94/mo), which is a welcome relief vs other providers. Great free plan, too - 3 apps at 256MB RAM running 24/7 in perpetuity.
Although, my dealings with them have been entirely negative. Their strategy is to make a clone of every service AWS has, except the billing system which hides a lot of majorly expensive pitfalls.
Also, if you choose one of their mainland hosted zones, trying to get your data out is even more expensive than the other zones.
So in this case, "knowing more about" is really "tread carefully".
- no nonsense prices
- solid hardware, with good, consistent performance & low steal (even on shared nodes, but dedicated is there if you need it)
- great support (humans you can call)
- adding more and more managed services over the last few years
https://www.akamai.com/newsroom/press-release/akamai-complet...
Whether the future is consistent with the past on the points you listed is hard to predict so soon after they got acquired, but I suspect it will increasingly diverge over time.
* Alibaba Cloud
* AWS
* Azure
* Azure Stack HCI
* Baidu Cloud
* BYOH
* Metal3
* DigitalOcean
* Exoscale
* GCP
* Hetzner
* IBM Cloud
* KubeVirt
* MAAS
* Nested
* Nutanix
* OpenStack
* Equinix Metal (formerly Packet)
* Sidero
* Tencent Cloud
* vSphere
Disclaimer: I currently work at Oracle.
We tried to use it at work, it was the second-biggest disaster of a cloud provider we've ever encountered (Huawei is the first). Stuff is just broken so often it's hard to ever feel like you can trust it.
You can specify the default storage class to use the provider's CSI driver. The PVC is then specified and consumed the same way in a portable manner across providers.
You can also choose something specific like a mounted hostVolume which is not provider specifc.
Cutting edge stuff like NVMe-oF is even possible, the CSI ecosystem is very active. https://github.com/spdk/spdk-csi
For single server, I'd just use sqlite.
We at DigitalOcean tend to see a lot of SaaS builders on our cloud. We obviously are smaller as compared to the big 3, but in our experience most people starting out with an idea don't really need all the bells and whistles.
Price-competitive, performant, and easily-automatable IaaS (compute/network/storage), and some managed services on top of it, are sufficient for a vast majority of builders.
I think about this thread all the time: https://news.ycombinator.com/item?id=22310879
Here is some marketing material with more information: - https://www.digitalocean.com/blog/how-to-scale-your-saas-pro... - https://www.digitalocean.com/blog/forrester-total-economic-i...
If you need many of their services, AWS is very difficult to beat. Nobody else quite matches up with them on breadth offerings. If you don't, do not use AWS due to the cost and the mental overhead of management.
If you just need to spin up some servers and want them to be fast and cost effective, Hetzner is the current champ in the US market (their new US datacenter). DigitalOcean, Linode and Vultr are entirely reasonable options after Hetzner.
If you have a $50-$100 / month or less budget, know how to set up and secure a linux server, and want to stretch your dollars as far as you can with a top notch cloud provider, Hetzner wins at present. And then throw Cloudflare out in front of it until or unless you can justify paying for a service. Keep your costs low, keep your infrastructure simple, keep your runway long, focus on selling selling selling (customers).
For ease of use, I like GCP.
Their serverless support story (Oracle Functions) is pretty clunky, but for standing up instances, and the terraform provider is quite verbose, but when it works, it works in a sane and logical manner.
As an added bonus they will randomly nuke accounts and production servers without even a curtsey email.
I won't even take on clients if they are using OCI.
To be fair, I have been testing Oracle Cloud Infrastructure(OCI) on a paid account for about 2 years now and actively monitor server hosting forums.
During that time I have had no "abnormal" problems with paid production account.
However, it is well known that Oracle accidentally nuked free/trial accounts early on, and again recently accidental nuked free/trial accounts reserved public IPs. However, I know of no such events for paid accounts, and paid accounts on OCI should be production ready quality in my experience.
On the other hand support really is somewhat bad, but support is free and the cloud costs are much less than AWS, while having better products than similar priced cloud products such as Lnode, Vultr, Digital ocean.
HyperScaler are the three you mentioned.
Then you have Alibaba, Tencent and Baidu. ( All Chinese )
Oracle and IBM Cloud
OVH, Hetzner,
Linode, Digital Ocean, UpCloud, Vultr, Scaleway. Possibly Fly.io ( not sure )
They all own their at least part of the Datacenter, Network and their own hardware. You have other services like Render and Heroku.
I am not aware of any potential upcoming players in this space ( Tell me about it ). But I think Cloudflare will likely be the big disruption in the next 5 years. They will own most of what we call "Cloud" services and Network and leave only the capital intensive compute to others.
Then you have many players that are the size of Linode or DO but dont operate as "Cloud".
I am still sad that Heroku didn't improve or move up / down the ladder. They seems to have happy at where they are.
Looking at YCombinator's alumni, there are a few interesting ones - for example, CloudThread.io (no affiliation) tracks cost efficiency of teams and applications, providing "technical cloud cost unit metrics as a service."
My company, Usage.AI, automatically buys and sells Reserved Instances to cut EC2 costs by up to 57%.
Been using Tilaa (if you can tolerate being hosted in Netherlands) for nearly 10 years, I'm quite happy with them.
Vultr is not as good, but it's decent enough.
But what separates a cloud provider from a VPS/Hosting provider?
I might argue some kind of object storage, as that's the differentiating feature AWS had when it came out. In which case, I'm not sure what there is outside of AWS and GCP, sadly.
OpenStack is the open source democratization of the cloud. In Europe there are a dozen such providers, in the US there's only one it seems. Hoping they continue to grow and other hosters break away from basic VPS hosting and also provide OpenStack clusters.
Pure IaaS ( Infra as a Service ) players are:
Oracle Cloud Digital Ocean
It’s compatible with S3’s API buy way way cheaper. Free egress.
- Oracle Cloud - IBM Cloud - Alibaba Cloud
Why do you think it's important to know more than the one you run on?
We don't run on one yet -- we are at the stage where we are expanding my understanding of the cloud service offerings and am curious what is available with good reputation.
SaaS is at whiteboard phase; where we invest our engineering time will be important as we recognize the issues of lock-in for something like AWS.
The more important question you might ponder is how to validate the idea is something people both want and will pay for, before investing time & money in a good cloud setup. If you select smaller providers, expect to spend engineering time reproducing existing products in the big ones. Saving money on VM hourly rates often costs you more elsewhere.
The only time I recommend worrying about multiple clouds is when you have a "deploy to your cloud" product and strategy.
What's the SaaS going to look like technically? Lots of compute processing? Or maybe large volumes of data? Perhaps machine learning? Or is it just a simple website and app deal?
Then there's the question of the technical competence on the ground (and the opinions that tag along for the ride). On the one hand, what's the expertise level? On the other hand, what do people prefer using? (And separately, what to people have experience with?)
I wonder (I don't have enough context to competently move the slider to "proper accusation") if somewhere along the way a general search for un-turned-over rocks got specialized/narrowed into "the overhead's all in the sheet metal benders" (the hardware). If this *is* the case, the only correct course of action (IMO) is to wind the train of thought backwards (oohc oohc) back to that frame of reference, then apply https://en.wikipedia.org/wiki/Five_whys until you're somewhere alien and interesting.
(Reiterating the caveat at the start of the previous paragraph, this is all massive conjecture and just-in-case assumption.)
If there was in fact a point at which hardware was specifically called out as a primary focus of optimization, I would loudly note that this type of hyperfocus leaves space for entire forests' worth of trees to fall over without ever being noticed, in this case for all of the software. Not just certain bits of it but like the whole kit and caboodle, unnoticed.
A related alternate possibility is that hardware optimization might be being treated like an axis point to orbit around, which can contribute to seeing things like immovable Mt Everests, and the construction of great and confusing Rube Goldberg machines to work around... perceived resistance that isn't really there.
To offer a 180-degree counterpoint that sorta flies in the face of the abstraction-away you might be trying to do here, if you want to look at different providers, I would suggest adopting a pat-answer strategy of keeping the architecture cloud-agnostic *where reasonable to do so*, and further suggest trying multiple providers - how does their support help out when it's 3am and you have no attention span and you accidentally something ridiculously straightforward? What's the performance like relative to the workload? What do all the engineers that are going to be headdesking against this stuff every day think about different options? Etc.
It's quite possible this reply is entirely misguided, in which case please ignore. I'm still learning the art of reading intent through text, I have a long way to go.
Your mindfulness is quite appreciated and your answer is on point!
We are at the design phase at this point. Fairly data intensive and will definitely have ML components (but that could be centralized and managed with standard MLOps practices). But the data is more on the backend, forward facing components are pretty standard CRUD/API calls against pre-computed and quickly updating data stores.
User-facing data payloads are in design, but expected to be low-to-mid volume.
I recently realized (light bulb moment) that cloud platforms like AWS are awesome for prototyping and experimentation: you can spin up seriously wide+deep instances that bill $1/hr, crunch through a complex workload in a couple of hours, and use hundreds of GB of RAM connected to hundreds of CPU cores for the cost of a stereotypical cup of coffee. But that bottoms out after just a very short period of time; a literal couple of hours. I wouldn't be surprised to learn the price scales was intentionally tuned to feel accessible at the prototyping stage - this would take advantage of communicative blind spots on both sides of the management/engineering boundary during evaluation and testing, where management gives engineering a green light budget to play around and learn about the platform, the engineers get insightful field experience after not very much expenditure, it feels possible to do a lot of work without spending too much, and... woops, a short time later they're staring down the barrel of a deadline, the fastest solution is to use their newfound experience and familiarity... and the cycle repeats.
The theoretical blind spot here is of course in the learning/tinkering stage that is so hard to communicate to management - and so when engineering makes positive noises about having some level of understanding and confidence about how to do a particular thing on AWS, the nuance of the hair-thin line between "thing that is generally achievable" and "I've practically played with AWS' implementation of the thing and I understand how to do it in that sandbox" is lost. (And the cycle repeats.) What's the right word for something incredibly well-engineered for altogether depressing reasons? :(
(Of course manglement and disorganization can also be at fault here, although I wonder where the cultural/structural root cause was when NASA forgot the egress bandwidth costs of migrating to AWS :D https://news.ycombinator.com/item?id=22626097)
Maybe the moral of the above story/thought experiment might be to theorize that AWS's market dominance has had a significant impact on broad tinkering and experimentation, and adjust for that by explicitly authorizing a broad-spectrum tinkering budget - "go and play" meets "engineering feasibility study" or something. Take the brakes off just enough that the engineering opinions that eventually come back are the product of having had the chance to stare into the horizon blankly for a bit, that sort of thing. Hypothetically.
Regarding the data intensity bit you mentioned, I'm personally up to the "just buy the whole computer" point at the moment mindset-wise, since this presents the most cost savings where it's feasible to do so. One somewhat cute but relatively transferable/normalizable way of broadly articulating the hard differences is to find vendor workstation/server configurators, adjust all the settings until the big number with the $ in front won't go any higher :), then (try and) configure something comparable on EC2.
An HP Z8 workstation with 2x28 core Xeon 8280s, 3TB of RAM, 172TB of storage, and 3 NVIDIA GPUs will set you back $111,575: https://zworkstations.com/products/hp-z8-workstation/?config...
An EC2 x2iedn.24xlarge (wat) with 96 CPUs, 3TB of RAM, 2 1.4TB SSDs, and 50TB outbound bandwidth, will set you back
- $192,625 for 1 year at standard monthly rates: https://calculator.aws/#/estimate?id=e9d103d78312322c281bc0c...
- $185,436 if you buy 1 year in advance: https://calculator.aws/#/estimate?id=c57064223f0aa745207efcf...
- $231,575 if you buy 3 years in advance: https://calculator.aws/#/estimate?id=2795a93261ed9353668b1bd...
An EC2 p4d.24xlarge (again, wat) with 96 CPUs, 1.1TB RAM, 8TB NVMe, 8 GPUs and no outbound transit (I couldn't figure out how to add any) is $196,398 for 1 year at standard monthly rates: https://calculator.aws/#/estimate?id=ef23ee773d298c703015453...
So this specific class of hardware would pay for itself in just over half a year ((111/192)*12=6.9375 months), give you total monopoly over resource availability, and redefine the bandwidth usage situation. The question then is whether the same sort of wet-fish disparity also exists within the target performance window you'd be aiming for.
In practice, datacenter colocation only charges for rack U height, power consumption and network bandwidth, regardless of the cost of the hardware (very cool) - and hmm, now I'm wondering if insurance options offer different premiums for different types of colocation facility standards - and then the only ongoing costs are having solid sysadmins (oh hey, remember those? lol) and inevitable hardware replacement (something something ZFS hot/warm spares).
The infuriating twist (sigh) is the ML part, and whether you want to do that on GPUs or otherwise.
I read a highly impressionable comment chain from someone who'd had some great engineering experiences playing with Google TPUs a few months back (https://news.ycombinator.com/item?id=27728225), and I wonder if the status quo described there has shifted at all - perhaps the scope of access has been reined in somewhat, or billing have come in and poured a pricing bucket over the engineering parade. I'm not sure. But what's stuck with me, sadly, about accelerator technology like this is that the cloud-computing theme of normalization-of-centralization unfortunately means this sort of hardware is only network-accessible from a large vendor like Google; you can't buy them. Infuriatingly. So even if the amazing-sounding prototyping opportunity described in the above link is still a thing (oh hey... the prototyping stage... this seems familiar ......), at production scale the only question that's on the table is "what are the pricing tiers?" since there's no leverage to be had, you either want the service at the price it's offered at or you don't.
Hence my question about using GPUs, which are, arguably, only barely at the break-even point in terms of buying vs renting for a lot of workloads. (Probably in large part only because enough large-scale scenarios demand on-premises provenance.) Because your alternatives are HaaS (hardware-as-a-service) for whatever the cloud providers offer. Yay.
I'm not entirely sure what you were referring to in terms of ML centralization - whether from a hardware or software standpoint - but I definitely go for the the "build a moat around it" mentality myself. I don't have enough experience to have the first clue what that would look like in practice though - just that it's a giant rabbithole of risk. (And a nuanced one, since Google et al want to generate demand for their product/service! Just... within the comfort zone of rent-seeking. Adlskdfglkjsdflgd *headdesk*)
In terms of what sounds like front/mid-end caching, Cloudflare currently seems to want to seriously differentiate on egress bandwidth, making it legitimately hard to dismiss them as interesting; but I do wonder what they'll do once they have both engineering familiarity and account entrenchment. I recall how Amazon Cloud Drive initially launched as both free and unlimited, then abruptly announced $60/TB pricing after people had uploaded TBs of data (and that one guy who uploaded 1 PB of test images). Cloudflare are certainly smart enough to see that similar actions would rapidly reorganize their reputation in a hurry, so I do wonder what their long-term strategy is there. Distilling the value-add down to basically "200+ PoPs and the world's lowest latency", it's totally possible to sustainably reproduce a heavy subset of that capability using commodity replication in a few key locations; and that may well suffice, supporting the argument that the value-add is not a critical requirement.
Hetzner
At some point 'the other clouds' aren't as relevant.
Think about the following interactions instead:
- CNCF compatibility (interpret as deep and as wide as you like)
- Infrastructure vs. Platforms vs. Services
- Legal boundaries
- Locality (can interact with legal limits, but also latency, transfer costs)
- Scope of services vs. scope of what you actually need
A lot of providers are good at a thing, and bad at everything they tack on to it. Some providers are reasonably good at many things, but win on integration between those things. Others are simply too dissimilar to orchestrate, so either you'll have to bring your own orchestration or not use it in orchestrated scenarios.Instead of knowing about the clouds, know about requirements engineering. Fitting your needs and the services you pay for is way more important than the details of those needs and services.
If you just need some random compute (read: a shell into an OS, a complete VM, a container, things like that) and nothing else, do NOT use some cloud. It will require you to do a lot of other things as a side-effect of using those services at all, and will cost a lot for what you need.
On the other hand, if you need to be highly elastic, have completely managed RDBMs on-demand available and orchestrate networking, IAM, object storage, block storage, compute and ingress, do start out with a cloud.
Regardless of what you are building, make sure you know ahead of time if:
- What your scaling is going to depend on (usage, work hours peaking, seasonal peaking, tenants)
- What your scaling is going to be like (horizontally scale and spread the load? vertically scale for a few weeks until you reach the scaling limit and then rebuild the application the right way instead? deploy one instance per customer?)
- What availability rules are you going to have? (downtime? data loss? time to recover?)
- what legal limits will it have?
Example for a MVP SaaS: say you want to manage shopping lists for consumers, you might call it Shoppr and build a PWA and a app-wrapped PWA so you get immense reach. You mostly have front-end engineers, but you do know a bit about metrics and scaling.I'd say that means:
- Downtime for a few hours unlikely to tank the business
- Legal limits are basically just generic data protection
- Scaling is likely linear
- Since your data is mostly basic CRUD, any read-replicated system will do
This can be built using any stack, and as long as your persistence can keep up you're golden. Don't fuck it up with an ORM that doesn't know how to create the proper indexing rules on tables and you can easily get a couple of million customers on an IaaS-only provider that just has virtual machines or containers, and only has one flavour of persistence store. Plonk Cloudflare in front of it and done.You can make this infinitely more complicated, but as an example add this feature to make this entire setup suck and the entire infrastructure incompatible with the needs of the application: international receipt scanning to recommend/autocomplete shopping lists for customers. Suddenly your requirements are expanded with:
- Incoming upload queue
- Object storage for image blobs
- OCR or ML pipeline to process images
- ML or Analysis pipeline to make sense of the contents of the now processed/read images
- Instances or expansion for all of the above per region
To make your traffic bill not suck you'll need endpoints in most major regions, and you'll probably want to prevent cross-region transfers so you'll want multi-region compute. Since you don't really have to do the processing realtime, you'll be doing some queue work and perhaps have a DLQ that needs human intervention or QA analysis for product improvement. All that stuff also needs a 'control panel' for lack of a better word so you'll be adding backoffice systems too, and those will have a workflow that doesn't compare to consumers at all and should never share any interaction with them, so now your application tenancy requirements change as well, which flows down into infrastructure requirements. At the same time you'll also need training data or validation data, and you'll want to be able to do all of that elastically to not go bankrupt for paying for 100% of capacity that you'll use 50% of the time at best. Suddenly 99% of the vendors are unable to provide what you need and the 'big three' remain (well, not exactly, but for illustrative purposes this will do).Considering a good SaaS might grow, make revenue and be sold or be valuable etc. there will be requirements stacked on top of everything else about redundancy, durability and availability and those will be increasingly hard to guarantee in an IaaS-only provider or PaaS-only provider scenario.
When you 'start' with a SaaS, you'll need to know what you need now, and what you might need in the near future, and make sure that whatever you do now doesn't paint you into a corner within weeks. That means that setting up IIS on a Windows desktop by hand on a VPS at some IaaS hosting provider is highly unlikely to be a good place to start. On the other hand, an equally janky setup with a random Ubuntu VM where you manually install Docker and say, Nomad might actually not paint you into a corner too much since a container can easily be run on a container-PaaS and Kubernetes beyond that.