Multi-Tenant Architectures
blog.codonomics.com
blog.codonomics.com
It's also amenable to distributed processing if you use something like Citus.
I'm really looking for an alternative to Citus, though. Citus itself is a bit tricky to use, and the SaaS version of it is owned by Microsoft, which means Azure-only. Also, Microsoft makes the SaaS version insanely expensive. If they had Citus-like features on Amazon Aurora I'd be there in a heartbeat.
The move from AWS to Azure came with a reduction of IOPS per instance. We have not had the same level of performance since migrating to Azure.
I am considering running our own Citus instances.
https://influitive.io/our-multi-tenancy-journey-with-postgre...
Some things were mistakes but others look like pretty fundamental flaws. Performance is a problem, database changes are a problem, and you aren't able to query across tenants.
I have recently been using row level security with a transaction middleware where I set the tenant ID.
Nice article regarding it - https://aws.amazon.com/blogs/database/multi-tenant-data-isol...
So, it works.
We're wanting to move off that architecture to something more future proof, but it's not our biggest pain point at this point in time.
Can you use SET LOCAL ROLE <user> on each transaction ?
It adds a layer of security so it might prevent some bugs leading to exploits. But in itself is not enough to rely on to separate tenants.
It looks like you could also use SET SESSION AUTHORISATION for this but I haven't used it so I don't know how this works with data access/pooling
I think for this use case security is focused on accidentally returning the wrong tenant's data (fully or partially)
Yes, but typically not across tenants. Maybe the flaw is only exploitable to admins of each tenant and they shouldn’t see other tenants data.
Customer A is running the new schema, so gets one cluster. Customer Z will get migrated in the next hour, and then none of the old system is running.
Once it was available on MSDN site of Microsoft, but I can't find it on microsoft.com anymore.
I think it really depends on the kind of application and the kind of user / customer you're dealing with. I'd probably lean towards "single" database with horizontal sharding .
But we can only do that because they pay a lot of money for this stuff. Our “shared, multi-tenant” environment is an order of magnitude cheaper.
Considering how long end exhausting the sales process are with enterprise customers I think the price is fair, even though the extra cost to operate that kind of customer is negligible.
As soon as you hear any of these, it's the hint you're dealing with enterprise and you have to up the price tenfold.
On the other hand, it takes a lot more effort to manage releases. I guess it scales in the same way Microsoft’s enterprise business scales; hard labor on top of a scalable platform.
Rule No. 1, never shard by account_id, Pareto is just waiting you kick you in the ass. Always shard by whatever gives the best distribution of workload.
There are more rules of course.
I can also anecdotally confirm the statement in another user's comment about 1 tenant using 60% of everything. And it's not even a large airline! This one relatively small carrier uses more of our resources than AA, UA and a few of the other top 10 global carriers combined.
The code was written with open sourcing it in mind actually. The amount of effort involved pushed me to select License Zero effort and specifically the Prosperity Public License https://prosperitylicense.com (License legal TLDR is “open” but if you use it to make money you have to pay me something I agree to) I’m hoping to do a significant refactoring before I mark the 1.0 version. I have some ideas that may be vastly cleaner but really requires nailing the whole thread local / async context vars lifecycle so I’ve kind of tried to limit who might start relying on it based on how extensively tested it is.
I agree that it is a completely valid approach, especially if you can afford it, but in my opinion it is just a single tenant application that can be easily spun up. For example, what if you want to manage all of your "tenants" within an admin interface. With a separate instance for all your tenants this is not really possible without additional development or a completely separate admin application.
What tends to happen is a first immediate client is needed along with a couple of sales demo clients. This buys time for the business to see if it’s viable before we bring on complexity.
But I agree clones of a core system isn’t multi tenancy.
For what it’s worth I tend to tell clients we won’t be sustainable after 4-5 of these such clones.
For people considering the clone approach let me caution one issue that comes up about 90% of the time. A client will as for a custom feature and be willing to pay for it for that “one” clone. These divergences are allowed to happen because we aren’t multi-tenant yet. You need to work hard to educate during this phase so you don’t create too much work down the line when you finally consolidate and refactor.
Exactly! That's the point.
Common usage 95% of of tenants is 20% of your traffic.
1% is 60% of traffic. Rest is the Rest.
Balancing is always weird. 4000 tenants can share a 3 servers and be fine. 1 client might actually need dedicated hardware.
Turns out AWS gets painful to manage beyond 1000 tenants. You start running into various account limits.
Devops it so that you can create new tenants with a script (in fact a script that is run by the signup codebase) that starts them out in cluster A.
Migration to cluster B would definitely be involved if you're not using a shared database - so potentially be ready to know your potential clients up front to get them on the correct cluster ahead of time.
I’ve often had teams of 2-6 to handle every technical aspect of a company with millions of visits. That includes feature development and support. K8 seems to need a large team just to keep it running.
At one point we had a dozen different Amazon accounts to keep everything separate. Thankfully the business model was crap, so if resolved itself.
Did Amazon mere said, "you are doing it wrong", without an explanation? In retrospective, do you think you could have done better? I'm looking to hear the lessons you learned, so may be, I don't end up making that mistake ;)
S3, I think they did up the bucket limit on request, but maxed out at 1,000. That was years ago though.
If extensions are rare, you can keep a switch for each and separate upgrades on model 4. However, if you want changes on switch, then you have to keep old code in conditional branches of your code forever. High cost of ownership.
Solution I implemented on my project (IoT) was shared apps, shared database. When we were starting the project, we decided to use a column as a discriminator and designed the system so that developers for entities that need to be tenant specific, just need to extend an abstract class, the rest of the system detects when you are trying to save or load such an entity and in those cases it applies a filter or automatically assigns the tenant ID. This means that normal developer can work just like he would on a single tenant application. I feel this is pretty normal stuff.
- hybrid model where paid / larger clients get dedicated hardware for perf reasons
- YAGNI -- will you even need multitenancy?
- business + legal considerations; which industries have legal or regulatory requirements not to intermix
- sharing and permissions -- what happens the first time someone needs to share a doc cross-account?
- tools and codebase strategies for verifying permission model on shared arch
Shared application scenarios can bring headaches when different customers want different application behavior implemented in upgrades.
Being forced to flip the switch for all of your users at the same time can make that an awfully big and scary switch you're about to touch.
Just because you're SaaS doesn't mean your clients are...
It's significantly better for us as a team to be able to focus on projects and improvements that benefit everyone. I've noticed the business starts to suffer when we focus on individual clients with narrow needs. Any one client only brings in 1/4000th of our revenue, so spending 1/4 of our development capacity on just something for them is a huge opportunity cost when we could be using it on something that benefits all our customers.
It was a different ballgame when I worked on on-prem software that cos millions and we only had 10 customers with a few customers sharing the same version.
Long ago I was the product manager of a complex enterprise platform which had been heavily customized for one of our large banking customers. They hosted the database on SQL Server shared clusters and much of the application backend in VMware instances running on "mainframe-grade" servers (dozens of cores, exotic high-speed storage). The hardware outlay alone was many hundred thousand dollars, and we interfaced with no less than 5 FTE's who comprised part of the teams maintaining it. Ours was one of a few applications hosted on their stack.
Despite repeated assurances of dedicated resource provisioning committed to us, our users often reported intermittent performance issues resulting in timeouts in our app. I was the first to admit our code had lots of runway remaining for performance optimization and more elegant handling of network blips. We embarked on a concerted effort to clean it up and saw huge improvements (which happily amortized to all of our other customers), but some of the performance issues still lingered. Over and over again in meetings IT pointed their fingers at us.
Eventually we replicated the issue in their DEV environment using lots of scrubbed and sanitized data, and small armies of volunteer users. I had a quite powerful laptop for the time (several CPU cores, 32GB RAM, high-end SSD's in RAID) and during our internal testing I actually hosted an entire scaled-down version of their DEV environment on it. During a site visit, we migrated their scrubbed data to my machine and connected all their clients to it. That's right, my little laptop replaced their whole back-end. It ran a bit slower but after several hours the users reported zero timeouts. This cheeky little demonstration finally caught the attention of some higher-up VP's who pushed hard on their IT department. A week later they traced the issue to a completely unrelated application that somehow managed to monopolize a good chunk of their storage bandwidth at certain points in the day. Our application was one of their more-utilized ones, but I bet correcting this issue must also have brought some relief to their other "tenants".
I know this isn't a perfect example, but it demonstrates how architecture encompasses a whole lot more than just the DB and apps. There's underlying hardware, resource provisioning, trust boundaries, isolation and security guarantees, risk containment, management, performance, monitoring and alerting, backups, availability and redundancy, upgrade and rollback capabilities, billing, etc. When you scale up toward Heroku/AWS/Azure/Google Cloud size I imagine such concerns must be quite prominent.
If they insist that the problem is our software, we tell them that we will begin troubleshooting, but if the error is outside of our responsibility, we will bill for all the hours used. 19 out of 20 times we restore the old version and benchmark them against each other and the performance is comparable. At that point they go back and recheck something, and it turns out they allocated resources differently or another application was upgraded too.
Nothing worse than having to spin up more and more instances and manage more and more hardware/virtual hardware/services as your customer base grows.
I would incrementally bring tenants together into a single DB.
> Suitability: This is the best way to begin your SAAS platform, for product-market fitment until stability and growth.
Uh what? How is this even feasible when you get to 1000s of clients?
It all depends on your customer/business model. If you're expecting to get to 1000s of clients quickly (i.e. within 12 months), then this will be an ops nightmare, unless you have really good automation.
In terms of security and privacy, isolating an individual tenant from others wasn't so much of a concern as each tenant was a customer within the organization with the same data classification level. So from a security perspective, we were "okay".
Where this gets interesting is that one tenant would suddenly decide to push a massive volume of data. Now processing events within a specific SLA was a critical non-functional requirement with this system. So then our on-call engineers would get alerts because shared messaging queues were getting backed up since Mr. Bob had decided to give us 3-5x his typical volume.
The traffic spike from one customer, which could last from minutes to hours, would negatively impact our SLAs with the other customers. Now all the customers would be upset. ^_0
Being internal customers, they were willing to pay for the excess traffic, but we didn't really have the tooling to auto scale. Our customers also didn't want us to rate limit them. Their expectation was that when they have traffic spikes, we need to be able to deal with it.
Now – we didn't want to run extra machines that sat idle like 90% of the time. And when we had these traffic spikes, we'd see our metrics and find maxed out CPU, memory, and even worse, we'd consume all disk space from the log volume on the machines filling everything up. The hosts would become zombies until someone logged in and manually freed up disk.
There were a few lessons learned:
1. Rate limit your customers (if your organization allows).
2. If your customers are adamant that in some instances each month, they need to be able to send you 5x the traffic without any notice, then you can't just rate limit them and be done with it. We adopted a solution where we would let our queues back up while some monitors would detect the excessive CPU or memory usage and would start scaling out the infrastructure. Once our monitors saw the message queues were looking normal again, they'd wait a little while and then scale back down.
3. When you're processing from a message queue, you need to capture metrics to track which customer is sending you what volume. Otherwise, you can have metrics on the message queues themselves and have one queue per customer.
4. If it's a matter of life and death (it wasn't, but that's how one customer described it), something you can do is stop logging when disk space usage exceeds a specific amount.
5. Also – when you have a high throughput system, think very carefully about every log statement you have. What is its purpose? Does it really add value?
This works really well to stop clients from doing this sort of things in the first place. Client devs get 429 errors for spamming the shit out of your server, they add some sleep to spread out the requests a bit. Everybody wins.
You will never hear any complaint about it, it's infinitely easier for the customer to add a sleep than to figure out how to contact your support.