Scaling to 100k Users
alexpareto.com
alexpareto.com
I have fintech systems in production with 100k+ users with complex Gov app for entire country that runs on commodity hardware (majority of work done by 1 backend server , 1 database server and all reporting by 1 reporting backend server using the same db). Based on our grafana metrics it can survive x10 number of users without upgrade of any kind. It runs on linux and dot net core and Sql Server.
Most of the software is not multimedia in nature and those numbers are off the charts for such systems.
I also work with fintech systems built upon .NET Core and have similar experiences regarding scaling of these solutions. You can get an incredible amount of throughput from a single box if you are careful with the technologies you use.
A single .NET Core web API process using Kestrel and whatever RDBMS (even SQLite) can absolutely devour requests in the <1 megabyte range. I would feel confident putting the largest customer we can imagine x10 on a single 16~32 core server for the solution we provide today. Obviously, if you are pushing any form of multimedia this starts to break down rapidly.
Actually thinking about it- I have inherited plenty of performance problems at the application layer, but that's because the joins were effectively being done at the application level (and making huge numbers of database calls) when they should have been done at the database layer.
Now a WD Blue NVMe does 95k/84k r/w iops at $215 and that's just off their website (https://shop.westerndigital.com/products/internal-drives/wd-..., may be more depending on shipping method...)
That said, it's not a fair comparison and I wouldn't want to run a big service on a single sqlite/nvme setup for more reasons than are worth mentioning, but not prematurely optimizing can take you really far - scale and money - with good design.
It could be 10 users that each views a few web pages once every month or it could be 10 heavy users for an internal system that uses the system most of the working day and creates thousands of calls per user to the system each day.
It is as a lot of people have mentioned also very different with information that can be easily cached and things like monetary transactions where you don't want to see old data in reads and therefore often can't cache or use read replica db.
For systems with heavy users I have seen the need to use a few servers even with just 1000 - 2000 users. For a web system the same can happen if you have 1000 users active on the site at the same time (within a minute or so).
For our company, we had more than 500k users with 1 small nginx + 1 medium appServer(with autoscaling, though never needed it) + 1 small cache server and RDS till now. We just added a aws managed load balancer into the setup and think it might be overkill.
For a client with NewsFeed needs, I used a dedicated server(64GB, 2TB space) to run nginx, app, cache, huge elasticsearch and postgres. It was great(and cheap) option for an MVP and let them validate the product for few months with >10k users.
It was awesome to learn a few years ago, how much compute power we don't use.
What would a managed cloud like AWS or Azure provide besides higher expenses? Last I checked, AWS would be ~$550/month based on their calculator.
IT departments managed redundancy long before cloud was a thing, and it's certainly possible to achieve acceptable uptime without the managed services. I realize that I still rely on DO for their uptime, and I'm pricing renting rackspace for that reason. However, uptime is currently within acceptable levels (ie I haven't had to explain to the CEO why all his employees are standing around). We had tried 2 SaaS vendors before our current on-prem solution, and both had had downtime that cost us money. Both were hosted in AWS, so clearly that isn't the panacea for ultimate reliability. If our on-prem server goes down, I get a call. However, I got calls when our SaaS vendor went down as well; sitting on line with their tech support wasn't any great comfort.
The same graceful failover processes are accessible on machines you control, there's no reason that you have to run on one machine if you avoid the cloud. The biggest hurdle for me to implement this was database concurrency with multiple servers, but a cloud solution wouldn't have done anything to solve that problem.
[edit:] typos and miss-typed AWS pricing due to mobile keyboard.
Redundancy was around long before cloud compute was a thing.
per year? per month? per day? simultaneously? doing what?
it matters.
i ask this as someone who runs a $40/mo Linode with a debian/nginx/node/mysql stack that's definitely 20x over-provisioned for an e-commerce site with 10k daily visitors, 15 simultaneous backend users (reporting, order-entry, CRM, analytics) and 0 caching tricks. i could easily run the site on any 5 year old laptop with an SSD and 8GB RAM.
normalize/de-normalize when needed, understand and hand-write efficient SQL queries (ditch ORMs), choose small/fast libs carefully (or write your own), and you can easily serve 100k users per day on a single cheap VPS with no orchestration/replication/hz-scaling bullshit. definitely can't say the same about 200k simultaneous users - that would need proper hardware, but can still be a single server.
Monoliths Are the Future: https://news.ycombinator.com/item?id=22193383
Depends on which ORM you use I think. Basic ORMs like Dapper in .NET just map your SQL query result to a object model.
Raw SQL strings or a query builder?
I guess it depends on the daily traffic.
When you say "local mssql account is required" do you mean a user or some other entity on the linux client? Or do you mean an mssql user that the client with authenticate as?
Using .net core also.
In general: I agree, there's rarely a case for really using the cloud. Page loads of my E-commerce project are 8ms for basket, I'm wondering what should kill it first on a big load, even without caching. Probably the database, not sure yet.
[1] https://www.microsoft.com/en-us/sql-server/sql-server-2017-p...
1: https://gist.github.com/majkinetor/877d5174ba322fbb808cc47a8...
Do not invest time making sure your service runs for $6 a month if it can run for $50 with 0 hours invested. Invest that time talking to customers and measuring what they do with your service.
Most times a few customers pay for the servers.
This is just a friendly reminder. I see a lot of comments talking about running backends for cheap.
Fast-forward two days later, his service and a competitor were both featured on Product Hunt. He's now making a profit on the service, as he managed to scale it up very fast, while the competitor buckled and completely lost momentum.
If you're talking about spending _a long time_ preparing a perfect infra, then your argument makes sense. Spending a few hours? It's both a great learning exercise and can literally save your project, so why not?
There’s other ways in other platforms as well (eg if you’re using Kubernetes)
* do i need to factor in experience with VMs?
* do i need to factor in OS and networking theory learnt in college?
* do i need to factor in high school algebra?
* do i need to factor in learning the english language?
Of course you need to factor in learning k8s. Before, it was learning VM's+ {ansible, chef, salt}, before that it was prob something else.
In practice - the application running in the pod has to be aware of this, and the most intensive part are rarely the bottleneck. Most of the time this is an architecture issue, not a resource issue. This takes time and experience, and overhauling a platform to remove a bottleneck is usually very painful if you have a bit more complex setup.
I am sorry but you are really talking about 2 days of time interval and projecting predictions based on that? Unbelievable...
So what's to stop the competitors to do the same thing that your friend did to invest 2-5 hours and catch up on the third day???
I wish I was good at ascii art, then I would draw a nice facepalm here.
I’m not convinced that you need to superscale your infrastructure first. I think it’s normally a waste of time and money. But for the example listed this is a likely benefit.
Pride, a refusal to accept vendor lock-in, misaligned incentives, sunk cost fallacy, etc, etc.
Welcome to the web?
That's an incredible story! Could you please link to the two products on PH?
The few hours you used on infrastructure can be better used fixing a bug, polishing/adding a feature or even giving yourself a break so you can be better focused the next day.
A Heroku-like platform will literally do the scaling for you. The non-financial cost is that you need to develop your application in-line with their framework/platform. If you make this decision at the start, this cost in practically nil.
With the risk of sounding like a broken record: yes in such a case it is better but that is often not the case at hand; this exact argument is used for spending $500/mo after optimizing the software vs $50k/mo autoscaling with ‘0 hours’ invested (between ‘ because ofcourse it takes a lot of time to even get that working, but, for many programmers, it is apparently easier work?).
A few weeks ago I commented here while optimizing a Laravel cloud install, now I am working on a Clojure one. Client is spending $28k/mo on aws, especially Dynamo and the rest on ELB.
Rewriting to Postgresql standard and the dynamo part to postgres columnstore and adding proper indexes has the lastest stress tests down to a few 100$/mo they will spend when they launch this.
The $28k is spent using exactly your argument, and like most of these projects like this I do, it was quite a lot less than 28k (1 month hosting) to optimize this (which had us rewrite a lot of spaghetti from dynamo to psql).
So yes, in some cases you are right, I would say when cloud hosting pops over 3k (especially if sudden), I would hire a me to have a bit of check to see if you are not burning money for nothing.
I am not saying Dynamo could not be used or tuned for this work, but Postgres was a better fix and more far more portable ofcourse.
So I read your comment as $28k dynamo to $100 postgresql.
Let's put your statement on another perspective : it is worth investing a few days/weeks of work of a low power machine (a human) to reduce significantly the long-term power usage of a high-powered machine (a computer or cluster thereof).
That seems a sad view on the value of a person.
Shouldn't the whole point of technology and infrastructure be to allow a person to use more resources so as to better[0] allocate their limited time?
1 kwh on the power grid produces 0.9884 lbs of greenhouse gas: https://carbonfund.org/calculation-methods/
0.1163 kWh of human energy produces 0.64 lbs of greenhouse gas: https://www.washingtonpost.com/national/health-science/runni...
If we even humor this sentiment (I don't buy that we should, 5$ vs 50$ of compute is not why we're struggling with climate change. Compute period is 10% of all electricity production, we're not going reduce that with a few hours of optimization here and there.. It's way too late to be taking such half-measures seriously), the math doesn't work out.
-
Thinking about climate change on a personal level is positive, but I also feel our efforts should be grounded in reality, not just things that feel good.
The hours you spend optimizing your bootstrapped service to reduce it's CO2 footprint... could be spent in plenty of other activities that actually reduce your carbon footprint.
We have people who don't believe it's real (which is silly) and, at the other end of the scale, people who believe we can actually fix it in 50 years (which is just as silly, even 100 years is silly). People are convinced they "know" the "truth" without even bothering to throw a few numbers at a spreadsheet to see if what they think they know aligns with any imaginable version of a non-science-fiction reality.
I am very concerned that politics and ignorance is driving this far more than real science.
5$ to 50$ is not 10 servers on Heroku. In fact, it's not even one server, you'll be sharing resources at that price point.
Let's say you generate approx. 500 lbs of CO2 a year (based of figures for a half desktop PC running 24/7 because you're only getting 2 cores and 1GB of ram)
At this point you're thinking, 500lbs?! That's insane!
But 40 hrs (1 week) is a lot of time, and 500lbs of CO2 is less than it seems.
If you spent 40hrs spread out over an entire year air drying your clothes a total of 20 times, you'd save over 2 Tons of CO2 a year. (You can do the math for a dishwasher if air drying doesn't work where you are)
If you live in a cold climate, the EPA estimates you can save 15%, or almost 1000 lbs of CO2, by weather proofing your home, easily accomplished in 40hrs
-
You might say "I already do all these things!", but the point is there are so many ways to convert time to CO2 savings.
We spend a lot of CO2 trying to save time you could say.
Optimizing your bootstrapped service is not one of the places I would use CO2 expenditure as reasoning in the slightest.
This comment reads you didn't read my reply, humans could make 0 CO2 and spending a week optimizing 10 servers and not make a dent compared to simple lifestyle changes.
I used 1 server because that's the scale thread was about ( saving <50$ of spend on Heroku)
I know you tried to make it about 10 servers to force a point, doesn't end up changing much though...
You found stats on that? I’m surprised
2% projected to reach 8% in a decade (https://fortune.com/2019/09/18/internet-cloud-server-data-ce...)
So using your math, you’ve got humans, which have high environmental costs in order to... well... live... wasting their life and consuming tons of resources doing something that was done better for cheaper.
Humans cost tons of money to operate. More than data centers or anything else. Don’t waste them on stupid projects like writing shitty versions of AWS.
Your own logic should lead you to conclude all the people wasting expensive human lives reinventing javascript frameworks, deployment systems, cloud orchestration systems, database systems and AWS... these are the folks doing true harm to the environment. It is much better for the planet to lock yourself into AWS, Azure, or google cloud and exploit the shit out of everything they do than it is to piss away incredibly expensive resources building your own.
And I am more than happy to boldly assert if you are working on a project that aims to re-invent AWS for your company... you are a waste of human capital.
A startup is a business; a side project may be pure hobby.
Most software engineers make plenty to budget for infrastructure at a side project.
Tech has progressed really far and there are tools like Netlify for hosting that would replace 90% of the non-DB parts of this. Cloud providers have also grown drastically and so again a lot of this would/could look a lot lot different.
Fwiw original deck from Spring of 2013, delivered at a VC event and then went on to be the most viewed/shared deck on Slidehare for a bit: https://www.slideshare.net/AmazonWebServices/scaling-on-aws-...
thanks, - munns@AWS
Fwiw, you are almost certainly shooting yourself in the foot by avoiding vendor lockin at stages before 8 revenue figures per year. Your engineering takes longer, is more brittle, and because you’re only using 1 vendor actively, your solution is still vendor locked-in.
Love, ~ Guy who learned his lesson many times
Writing service API integration code instead of code that interfaces directly with the technology that service is doing makes code quite brittle. If/when the vendor deprecates the service, introduces backwards incompatible changes, or abandons development of the product, you're left on the hook to engineer your way out of that problem. Often times that effort is equal to or greater than the effort of an in-house solution in the first place.
I had the same mentality as you until this happened to the SaaS product I work on for a few different services. Now at very least I try to make sure solutions are cloud agnostic.
Actually, we did once and we found that our abstraction was so tightly coupled to the underlying API that we had to remake it anyway. The core concepts between those APIs were just too different.
And I’ve had at least 2 cases where our attempt at being vendor agnostic made the integration completely fail and never work right. To the point the vendor told us “You’re holding it wrong, please stop”
removed Uber as its not 100% cloud (or at least wasn't in the past)
With managed services it looks a world different.
"Let's work together and design a system that scales appropriate but isn't overbuilt. Let's start with 10 users".
Then we talk about what we need and go from there. The end result looks a lot like this blog post, for those who are qualified.
Heh, you're being modest. I'm sure you've dealt with far more complex distributed systems than the hypothetical one in the blog post.
If you can scale to 100K users, you can probably learn the rest to scale to 100M users.
The reason punting on this is a good idea is because you can get pretty far with vertical scaling, database optimization, and caching. And when push comes to shove, you are going to need to shard the data anyway to scale writes, reduce index depths, etc. So a re-architecture of your data layer will need to happen eventually, so it may turn out that you can avoid the intermediate "read from replica" overhaul by just punting the ball until sharding becomes necessary.
And then you have to rearchitect your data layer under extreme duress as your databases are constantly on fire.
So you really need to find the balance point and start doing it before your databases are on fire all the time.
With that, you can start planning your rearchitecture if you're running out of upgrades, and start implementing when your servers aren't yet on fire, but are likely to be.
Today's server hardware ecosystem isn't advancing as reliably as it was 8 years ago, but we're still seeing significant capacity upgrades every couple years. If you're CPU bound, the new Zen2 Epyc processors are pretty exciting, I think they also increased the amount of accessible ram, which is also a potential scaling bottleneck.
But that's not how the real world works. The databases don't just slowly get bad. They hit a wall, and when they do it is pretty unpredictable. Unless you have your scaling story set ahead of time, you're gonna have a bad day (or week).
Usually, databases are pretty good at running up to 100%, though. And if you started with small hardware, and have upgraded a few times already, you should have a pretty good idea of where your wall is going to hit. Some systems won't work much better on a two socket system than a one socket system, because the work isn't open to concurrency, but again, we're talking about scaling databases, and database authors spend a lot of time working on scaling, and do a pretty good job. Going vertically up to a two socket system makes a lot of sense on a database; four and eight socket systems could work too, but get a lot more expensive pretty fast.
Sometimes, the wall on a databases is from bad queries or bad tuning; sharding can help with that, because maybe you isolate the bad queries and they don't affect everyone at once, but fixing those queries would help you stay on a single database design.
It can be an easy fix (buy more memory), but the first time it happens it can be pretty mysterious.
Your will implement new features, add new tables and columns, indexes, etc which will affect your data layer.
So just splitting up the heaviest part (eg. Catalog) into "a microservice" would be easy while I add nginx as load balancer. I already separated domain vs Integration Events.
Both now use events in memory in the application, I only need a message broker like NATS then for the integration events.
It would be a easy wall ;). I have multiple options like heavier hardware, splitting up the db from application server or splitting up a domain bound api to a seperate server.
As long as I don't need multimedia streaming, kubernetes or implement Kafka the future is clear.
Ps. Load balancing based on tenant and cookie would be a easy fix in extreme circomstances.
The thing I'm afraid for the most is hitting the identity server for authentication/token verification. Not sure if it's justified though.
Side note: one application has an insane amount of complex joins and will not scale :)
That will run Stackoverflow's db by itself for reference, along with sensible caching (they're very read-heavy and cache like crazy). Here's their hardware for their SQL server for 2016:
2 Dell R720xd Servers featuring: Dual E5-2697v2 Processors (12 cores @2.7–3.5GHz each), 384 GB of RAM (24x 16 GB DIMMs), 1x Intel P3608 4 TB NVMe PCIe SSD (RAID 0, 2 controllers per card), 24x Intel 710 200 GB SATA SSDs (RAID 10), Dual 10 Gbps network (Intel X540/I350 NDC).
https://nickcraver.com/blog/2016/03/29/stack-overflow-the-ha...
We did use memcached where we could.
Truthfully, unless you're working on some kind of non-transactional problem like analytics, even assuming you will need to shard the data or scale out reads ever due to user activity is borderline irrational unless you have extremely robust projections. The database will be the last domino to fall after you've added sufficient caching and software optimization. It's so far down field for most projects (and the incidental complexity cost so high) that my personal bias is that even having the conversation about such things on most projects isn't even worth the opportunity cost vs talking about something else.
Even then, the first thing to fall over will probably be write heavy analytics-like tables that are usually append only due to index write load. Out of the box, you can often 'solve' this by partitioning the table (instead of sharding.) In modern DBs, this is a simple schema change.
Having many servers gives you redundancy and horizontal scalability, but also comes at a high complexity and maintenance cost. Also, with many machines communicating over the network, latency and reliability can become much harder to manage.
Most smaller companies can probably get away with having a single powerful server with one extra server for failover, and probably two more for the database with failover as well. I think this would also result in better performance and reliability as well. I’m curious to know whether the author tried vertical scaling first or went straight to horizontal scaling.
And as a side note, anyone who is using up that kind of data is not going to be able to afford cloud egress prices unless they are making a mint on those users. Saturating a 10gbps connection would cost you around $450 an hour at AWS rates.
The primary difference is that this post tries to be more generic, whereas the original is specific to AWS.
The original, for what it is worth, is far more detailed than this one.
>> This post was inspired by one of my favorite posts on High Scalability. I wanted to flesh the article out a bit more for the early stages and make it a bit more cloud agnostic. Definitely check it out if you’re interested in these kind of things.
With 10 users you don't "need" to separate out the database layer. Heck you don't need to do that with 100 users. Website I ran back in 2007-2010 had tens of thousands of users on a single machine running app, database, and caching fine.
Users are actually a really poor way use for scalability planning. What's more relevant is queries/data transmission per interval, and also the distribution of the type of data transfers.
I'd say replace the "Users" in this posts to "queries per second" and then I think it's a better general guide.
Let's say that maybe 10% of your users are on at any given time and they each may make 1 request a minute. That's under 200 QPS which a single server running a half-decent stack should be able to handle fine.
We are however actually thinking about switching to a new dedicated server at another provider (Hetzner) where we are looking at having the Web server and the DB on the same server, however the new server will have hugely improved performance (which is sometimes needed), still at a reduced cost compared to the DigitalOcean setup.
The thing we are doubting is if having a managed db is worth it. The sell in is that everything is of course managed. But what does this mean in reality? Updating packages is easy, backups as well (to some extent), and we still do not use any standby nodes and doubt we will need any replication. So far we have never had the need to recover anything (about 5 years). Before we got the managed db we had it in the same machine (as we are now looking at going back to) and never had any issues.
Any input?
And while Hetzner customer support is generally excellent, in my experience, their handling of DDoS incidents will generally leave your server blackholed and sometimes requires manual intervention to get back online.
This is something you need to account for in terms of redundance if you are planning to expose your application directly to the net without any CDN/Load balancer/DDoS filter in place.
From my experience it makes sense to work with a data centre that is less focussed on a mass market but allows for individual client relations to mitigate risks like that. I love Hetzner for what they are and do host some services with them, but I wouldn't build a business around services hosted there.
And this not only goes for Hetzner but pretty much any provider whose business model is based on low margin/high throughput.
Also want to throw in there that it is important to not only compare specs, but to also compare hardware. If DO has newer chips and faster RAM, then you will take a performance hit moving to the new provider even if the machine is beefier.
Here's their pitch: easy setup & maintenance, scaling, daily backups, optional standby nodes & automated failover, fast reliable performance including SSDs, can run on the private network at DO and encrypts data at rest & in transit.
For example, I would be surprised if they noticed that your IOPS was high and you needed to upgrade the storage/disk components. (That would be cool if it's the type of thing they offer).
The old Stack Overflow podcast was also very instructive. They went a very long way on a single server and had the Reddit founders on the show to talk about their scaling during their process of adding a second box. This was on servers of the mid-aughts, running ASP.NET.
So, in the ‘spirit’ of this article, would that not be one of the first things to implement with the system?
Wait until you have added a caching layer and sharding the DB, to begin implementing logging?
I may not be reading this correctly.
I could see the case being made for distributed tracing, but having a logging strategy that can also scale and be flexible seems really important, to me at least.
App load: |User| <-> |Cloudfront| <-> |S3 hosted React/Vue app|
App operations: |App| <-> |Api Gateway| <-> |Lambda| <-> |Dynamo DB|
Add in Route53 for DNS, ACM to manage certs, Secrets Manager to store secrets, SES for Email and Cognito for users.
All this will not cost a whole lot until you grow. At that point, you can make additional engineering decisions to manage costs.
In the context of a start-up, cost is a big factor and then perhaps (hopefully) handling growth. You could start small and refactor apps/infrastructure as you grow but I am unsure how one could afford to do that efficiently while also managing a growing startup.
On the selling soul to cloud provider, I don't see it like that. I have a start-up to bootstrap and I want to see it grow before making altruistic decisions that would sustain the business model.
Once you are past the initial growth stage, there are many options for serverless, gateway, caches, proxies that can be orchestrated in K8 on commodity VMs in the datacenter. Though this is where you would need some decent financial backing.
(I am not associated with Amazon, Google or Azure. I do run my start-up on Azure.)
As patio11 would like to remind us all, we've got a revenue problem, not a cost problem. [0]
Sadly, anything more in-depth than that, you'll need to sign an NDA with AWS to learn anything about the performance limits of their services (eg Redshift), and you won't get that unless you're already a big customer there. Azure's not going to be falling over themselves to let you know where they fall short, either. This is vendor lock-in, and is why there are so many free cloud credits to be had to startups.
This is also a reason I believe SaaS companies will find it is harder than they realized to arbitrage between clouds, and business models based on that may not be able to get that right.
Using one of those is where you'll spend most of your operational time if you really need that level of scalability. Most people don't, but the options are there if you really need them.
If you are happy with your storage layer, which most people are, the rest scales horizontally pretty easily. And there are plenty of free things you can use to get what a cloud provider gives you.
> App load: |User| <-> |Cloudfront| <-> |S3 hosted React/Vue app|
The CDN is always going to be tough to replicate on your own. In the end, latency is bounded by the speed of light, so you can only bring your files closer to your users. I wouldn't expect you to build one of these yourself; just buy one until you're the size of Google.
> App operations: |App| <-> |Api Gateway| <-> |Lambda| <-> |Dynamo DB|
Your app should be designed to scale horizontally; don't keep any state in your app, delegate it to your storage layer so you can scale a CPU-intensive app up across multiple servers.
There are quite a few API gateways around; Ambassador comes to mind but there are a million. I personally use raw Envoy for everything. I was load-testing my website the other day and pushed 5000qps through it from my cable connection before I decided "it's probably fine". (I started dropping frames on the Twitch stream I was watching, though ;)
There are plenty of "serverless" frameworks that emulate what Lambda does. knative comes to mind. I have not experimented with them in depth, but am intrigued by the idea. (I am more intrigued by turning config files into webassembly-compiled programs, to make existing apps more configurable at runtime. This is like serverless, but less general.)
> Add in Route53 for DNS, ACM to manage certs, Secrets Manager to store secrets, SES for Email and Cognito for users.
CoreDNS scales nicely and has an API. cert-manager is an open source way of obtaining certificates (though it's tightly coupled to Kubernetes); either ACME (letsencrypt) or your own root CA. There are a bunch of free software secret managers; Vault, bitnami-labs/sealed-secrets, etc. I personally use git-crypt ;)
Email deliverability is always going to be an issue. Like the CDNs, you might want to delegate it while you're small. Use anything except Mandrill.
Another good use-case is to store checkpointing information. Say, you've processed some task and would like to check-in the result. Either the information fits the 400KB DynamoDB limit or you use DynamoDB as a index to a S3 file.
You could do those things with managed or self-hosted RDBMS, but DynamoDB takes away the need to manage the hardware, the backups, the scale-ups, and the scale-outs, reduces ceremony whilst dealing with locks, schemas, misbehaving clients, and myraid other configuration knobs whilst also fitting your queries patterns to a tee.
KV stores typically give you consistent performance on reads and writes, if you avoid cascading relationships between two or more keys, and make just the right amount of trade-offs in terms of both cross-cluster data-consistency and cross-table data-consistency.
Besides, in terms of features, one can add a write-through cache in front of a DynamoDB table, can point-in-time-restore data up to a minute granularity, can create on-demand tables that scale with load (not worry about provisioned capacity anymore), can auto-stream updates to Elasticsearch for materialised views or consume the updates in real-time themselves, can replicate tables world-wide with lax consistency guarantees and so on...with very little fuss, if any.
Running databases is hard. I pretty much exclusively favour a managed solution over self-hosted one, at this point. And for denormalized data, a managed KV store makes for a viable solution, imo.
While this could be ok for some apps, I think for most use cases it's really bad and ends up being more trouble than what you save on ops in the long run, especially considering options like Aurora that, while not as hands-off as Dynamo, are still pretty low-maintenance and don't limit transactions at all.
Between my issues with AWS currently and the exterior look of Amazon, I'm skeptical AWS is a good solution.
One of the other major advantages of cloud is that you can save a lot in support staff. Compare the wages of even 1 decent sysadmin looking after your own hardware compared to several thousand dollars of AWS and it's still loads cheaper. Hardware upgrades, OS updates etc. are often automatic or hidden.
Many of those steps are reduce scalability if they're applied prematurely, splitting out the API and database layer at 1000 "users" is going to use more resources serializing things across the network than keeping it in process would. Same for seperating out the database, it's great if you need it but there's a cost if you don't. I worked on one system where we pulled out the API layer after realizing that this was where ~50% of our CPU time was being spent.
It also seems to focus on vertical layering more than horizontal splitting, being a photo sharing website I would have thought there was a lot of CPU intensive photo manipulation or something they can split off to services on the side that doesn't need to be done in real time.
Clearly a million users on Facebook is much heavier than a million registered with online banking and who only use it once a month.
1) I don't really get the "100 Users: Split Out the Clients".
It definetly helps in terms of understanding your customer profiles, whether they prefer the mobile app or the web interface for instance, and it might help from a usability point of view, but how does this help scalability per se if the API layer stays the same?
Splitting client happens, obviously, at the client level..
2) Also, I don't understand "This is why I like to think of the client as separate from the API."
Who considers the client and the API as the same thing?
You can consider the API as the client of the DB, sure, but why would you mix the user client and API together?
3) Caching: here I lack some knowledge. "We’ll cache the result from the database in Redis under the key user:id with an expiration time of 30 seconds". I assume that every access to Redis will not refresh the cache (aka reset the counter), otherwise you could potentially never get an updated data, right?
1. Should I interpret each X number of users as stated in the article to be "simultaneous" or "generally active over the past Y amount of time" or even "total user count ever?" "Unique per day?"
2. What do you think in general about skipping the step "100 Users: Split Out the Clients" if one is reasonably certain to not want or need multiple clients? It would seem as though this could keep the deployments and testing simplified until later in the growth stage, as more code can be deployed/tested as a single bundle. But also, I want to be sure I'm not missing something by just trying to justify my own interests.
I figure if we did get to a thousand or ten thousand users, we could recruit the DevOps talent to pull this off before the bills kill us!
How do you plan on recruiting DevOps talent? Do you have funding to pay them?
* I'd imagine the website layer is frequently static (html/js) and could just be hosted on s3/cdn. One part of scaling avoided.
> This is when we are going to want to start looking into partitioning and sharding the database.
You have to be at pretty huge scale before you really need to consider this. A giant RDS instance, some read replicas, and occasional fixing of bottlenecks will go a lonnng way. And scaling RDS is a few clicks. By the time you need to start sharding, you can probably afford a dedicated database engineer, or at least I'd hope.
I have built a very fast and efficient CPU-only neural TTS engine in Rust/Torch JIT that is the synthesis of three different models. I've got a bunch of celebrity and cartoon voices I've trained. The selling point is that this runs on cheap, commodity hardware and doesn't require GPUs. I can easily horizontally scale it as a service.
I've currently got it running in a Kubernetes autoscaling group on DigitalOcean, but I'm worried about the bandwidth costs of serving up potentially thousands of hours of generated audio. I haven't thrown any real traffic at it beyond load testing, but I think it can survive heavy traffic. The thing that worries me is the bandwidth bill.
Does anyone have experience with other hosts that are cheap for bandwidth intensive apps? Are there hosts that provide egress bandwidth on the cheap for dynamically generated (non-CDN) content?
Subsequent to this, I would really like to sell or monetize this app so I can fund the R&D / CapEx intensive startup I really want to undertake.
Who might be the market to buy a TTS system like this?
I was thinking Cartoon Network might want "Rick and Morty" TTS, but despite my engineering to scale this and make it sound really good, I doubt they'd pay me much for the product. I suppose $2M would give me runway to hire a few engineers and buy a lot of the equipment I need, but I have no idea who would pay for this.
Glass for optics is surprisingly expensive, and beyond that I have other extremely high R&D costs.
Alternatively, I also have a "real time" (~800ms delay) neural voice conversion system. I thought about running a Kickstarter campaign and selling it to gamers / the discord demographic. It's relatively high fidelity with no spectral distortion, and I have a bunch of hypothetical mechanisms to make it an even better fit.
I've also thought about slapping a cute animation system on top of my TTS service let people animate characters interacting. (Value add?) An earlier non-neural TTS system I built before the last presidential election cycle had something like this, but more primitive: http://trumped.com (The audio quality of this concatenative system is absolute garbage. The new thing I've built is unrelated.)
I'm curious what kind of hardware can sustain 100k concurrent connections these days.
Add large request/response sizes or CPU/RAM bound operations and your servers can very quickly reach their limits with far fewer concurrent requests.
Architecture is a big picture task since you have to consider the whole system before implementing part of it, otherwise you end up having to start again.
It teaches good business practice of tackling 1 thing at a time & not over-engineering.
But simply using read replicas, caches, CDNs, etc. does not mean things will scale.
Actually, often this will break your app unless you write concurrency-safe code. Learning and writing concurrent code is how you scale, the "adding more boxes" is just the after effect.
For instance, we have about 10M+ monthly users running on GUN (https://github.com/amark/gun), this is because it makes it easy/fun to work/play with concurrent data structures, so you get high scalability for free without having to re-architect.
But learning something new is never an excuse for shipping stuff today. Ship stuff today, you can always learn when you need it.