What's Wrong with Amazon's DynamoDB Pricing?
bailis.org
bailis.org
Someone who wanted an interesting research project might be able to pull back the covers on DynamoDB a bit by issuing a large number of requests, very carefully measuring the response latency curves for consistent and inconsistent reads, and looking at what model fits them best.
>Someone who wanted an interesting research project might be able to pull back the covers on DynamoDB a bit by issuing a large number of requests, very carefully measuring the response latency curves for consistent and inconsistent reads, and looking at what model fits them best.
This is possibly interesting, but why shouldn't Amazon just provide this data? The data would useful to developers who otherwise have to guess.
And finally, why not allow users to place this cost on the write path instead? "W=3" seems like a reasonable option (given that consistent reads may be unavailable anyway).
Well, it's not just the number of disk I/Os which increases -- you've also got the network traffic and CPU time associated with each copy of the read request you send out. Given that (a) DynamoDB is using SSDs, and (b) Amazon seems to love using slow languages, I'd guess that the CPU time is what ends up dominating their cost.
why shouldn't Amazon just provide this data?
Amazon is very secretive. And I can't really blame them; after all, why would they want to subsidize their competitors' development?
And finally, why not allow users to place this cost on the write path instead? "W=3" seems like a reasonable option
I'd guess they wanted to be able to tolerate partitions without sacrificing availability. (At least for the common case where "partition" means that nodes are completely offline and are both unable to communicate with other nodes and unable to receive incoming requests.)
-You have to provision throughput for each table individually, so you basically have to pay for the maximum throughput you expect for every single table all the time (you can only reduce your throughput once per day as far as I understand). This means that adding a new table can be pretty expensive (even if your total throughput isn't increasing at all).
-Each unit of throughput gives you one 1kb write/read per second. If you exceed your throughput for a second, the call fails. This means that if you want to support the ability to write 50kb (like a "notes" field or something), you need to constantly pay for 50 write units even if you won't ever realistically use that much. And you have to do that for every single table
As the OP points out, it's all about FUD. I'm so terrified of exceeding my throughput allotment that I'm forced to pay for an order of magnitude more resources than I will actually use. This seems to go against everything AWS is about.
At the same time, clearly this is not the case for something like EC2, where the marginal cost for another machine isn't necessarily 0, but it is true for Dynamo in some sense. Are we in a world where one can't choose to mix the model?
That said, what provisioned capacity means in DynamoDB is pretty opaque. Great point about the utility of latency information for developers.
Conflating resource usage, pricing, and consistency using a single metric (here, price) seems confusing, especially without the right metrics to guide devs.