If you are doing INSERT in batches with one million rows, it will give
SELECT formatReadableQuantity(1000000 * 100 / 0.0125)
8.00 billion
inserted rows per dollar. Pretty good, IMO.If you are doing millions of INSERT queries with one record, without "async_insert" setting, it will cost much more.
That's why we have "write units" instead of just counting inserts.
As I've said several times in this thread, I understand why you don't count inserts or rows. What I don't understand is what unit a WU does actually correspond to. In particular I don't understand its relation to e.g. parts or blocks, which are the units one would focus on optimizing self-hosted offerings.
For those complex pipelines you may find more useful to run tests during trial. Data distribution, partitioning and so on can change actual cost significantly so estimates can be too pessimistic or optimistic
Right, that's exactly what I don't want to deal with. Unless I have even just a ballpark estimate of complex pipelines both before I commit to any sales crap and afterwards when we're designing new pipelines, it's just not an option for us at all. I have no clue if it's going to cost us $10, $100, or $10000.
Check out the response below that has a reference to some of our billing FAQs.
Is one block on one non-partitioned non-distributed table one write unit? What about one insert that's two blocks on such a table? What about one block on a null engine with two MVs listening to insert into two non-partitioned non-distributed tables? What if the table is a replacing mergetree, do I incur WUs for compactions? etc.
My worry is that it is essentially 1 WU = 1 new part file, which I understand makes sense to bill on but is tremendously intransparent for users - at least I have no clue how often we roll new part files, instead I'm focused on total network and disk i/o performance on one side and client query latency on the other.
For example, I just checked that uploading 1.1GB example table(cell_towers with 14 columns) cost me 0.38 write units.
Nonetheless you can't insert a-whole-file-and-just-that-file in less than one write.
This doesn't imply to me that each individual INSERT costs 1 WU, but that it could be fractional. I guess it depends on how you read it?
(See https://news.ycombinator.com/item?id=33081099 for the original wording.)
There's no way to think about what an actual write unit means. You could measure the costs on a sample workload, but that's far from ideal. Some transparency here would be nice.
I understand the answer is complicated, based on hairy implementation details, and subject to change. Give me the complexity and let me interpret it according to my needs.
Working on updating the FAQ and tooltips now and sharing your feedback. <3