I’m also curious what application domains are storing close to petabytes on a single machine with only a few terabytes of dram.
I’m also curious what application domains are storing close to petabytes on a single machine with only a few terabytes of dram.
This has become a great emerging market for new database tech, tremendous amounts of money are being spent on these data models and traditional data infrastructures are very poor for this. The size and velocity of the data implies that you need a single database instance that can run ingestion workloads concurrent with analysis workloads.
The research on caching algorithms for ultra-dense storage is new, it is literally just starting to be worked into production designs, so I wouldn't expect references. Most advanced database research is done outside academia, so a significant fraction is published long after implementation or not formally published at all. Publication is not the objective per se and there is little time or incentive to do that work. Database research has a discovery problem, I hear about most interesting new research via informal community channels and much of that never shows up in literature.
The fact that the research is being done outside of academia and not being published is pretty sad imo. Hopefully those in the know decide to publish at least tech-reports/whitepapers once they have a share of the market.