655 karma · joined August 31, 2016
https://blog.acolyer.org/2018/09/26/the-design-and-implement...
If the sharding key matches (or is a subset of) a join or group-by key, then identical values are local to a single shard, which can be processed independently.
This type of thing is typically done at large granularity (eg one shard per MPP compute node), but there are also benefits down to the core or thread level.
Another tip is that if no shard key is defined, hash the whole row as a default.
I did learn a lot from watching MDS people try to beat it (in the end I'm also looking for what should come next), but mostly it was confirming the article and RDD. What they didn't know about data warehousing they also didn't know about performance or price-performance or selecting tools or managing projects or vendors, so costs exploded.
These folks were hilarious because they kept insisting that Vertica is not "modern", while it beat the pants off them with basic columnstore stuff.
I find it hard to imagine obtaining much bias from a random hash seed in a large group of small-scale users, but I haven't looked at the problem closely.
What became funny was the series of "tasters" who claimed they had figured out the formula with other ingredients like cinnamon, etc.
I likely need to read the paper linked, but it's common to have an MPP database lose a node but maintain data availability. CAP applies at various levels, but the notion of availability differs:
1. all nodes available 2. all data available
Redundancy can make #2 a lot more common than #1.
While there are all kinds of issues it can't handle for CRUD, such as locking and small updates, it's also expensive to do small reads due to the reliance on auto-scaling instead of timesharing. If your app pageloads and issues dozens of queries to populate its widgets, SF will be both slow and expensive.
We've been meeting online, and I'm interested in setting up physical venues if anyone has one.
The PWL Discord: https://discord.com/channels/1025104619975737446/10251046205...
What I'm curious about is whether Neon can run pg locally on the app server. The company's SaaS model doesn't seem to support that, but it looks technically doable, particularly with a read-only workload.
Snowflake compute instance types cost about $.30-.40 an hour on EC2, so it's quite a markup.
As far as security I do believe they allow the customers to set their own storage keys, so there may be some isolation from a global breach.
The pros:
-- Samples are small and fast most of the time.
-- can be used opportunistically, eg in queries against the full dataset.
-- can run more complex queries that can't be pre-aggregated (but not always accurately).
The cons:
-- requires planning about what to sample and what types of queries you're answering. Sudden requirements changes are difficult.
-- data skew makes uniform sampling a bad choice.
-- requires ETL pipelines to do the sampling as new data comes in. That includes re-running large backfills if data or sampling changes.
-- requires explaining error to users
-- Data sketches can be particularly inflexible; they're usually good at one metric but can't adapt to new ones. Queries also have to be mapped into set operations.
These problems can be mitigated with proper management tools; I have built frameworks for this type of application before -- fixed dashboards with slow-changing requirements are relatively easy to handle.
> OLAP databases are not good at serving large volumes of requests with low latency within milliseconds.
Just use Vertica lol.
> The Nitro Cards are physically connected to the system main board and its processors via PCIe, but are otherwise logically isolated from the system main board that runs customer workloads.
https://docs.aws.amazon.com/whitepapers/latest/security-desi...
> In order to make the [SSD] devices last as long as possible, the firmware is responsible for a process known as wear leveling.... There’s some housekeeping (a form of garbage collection) involved in this process, and garden-variety SSDs can slow down (creating latency spikes) at unpredictable times when dealing with a barrage of writes. We also took advantage of our database expertise and built a very sophisticated, power-fail-safe journal-based database into the SSD firmware.
https://aws.amazon.com/blogs/aws/aws-nitro-ssd-high-performa...
This firmware layer seems like a good candidate for the slowdown.
Also, both features use Nitro SSD cards, according to AWS docs. The Nitro architecture is all locally attached -- instance storage to the instance, EBS to the EBS server.