306 karma · joined December 15, 2015
Previously:
CTO of 3dc.io CEO of Callista Software eerik.kivistik@callista.ee
* JuiceFS - Works well, for high performance it has limited use cases where privacy concerns matter. There is the open source version, which is slower. The metadata backend selection really matters if you are tuning for latency.
* Lustre - Heavily optimised for latency. Gets very expensive if you need more bandwidth, as it is tiered and tied to volume sizes. Managed solutions available pretty much everywhere.
* EFS - Surprisingly good these days, still insanely expensive. Useful for small amounts of data (few terabytes).
* FlexFS - An interesting beast. It murders on bandwidth/cost. But slightly loses on latency sensitive operations. Great if you have petabyte scale data and need to parallel process it. But struggles when you have tooling that does many small unbuffered writes.
You don't actually directly charge for storage itself, so I assume this a "bring your own s3 bucket" type of deal, correct?
How long does data, that is no longer being accessed sit in the cache and count towards billing?
As for availability, are you in the process or do you have plans to also support Google Cloud?
* What does "$0.05 / gigabyte transferred" mean exactly. Transferred outside of AWS or accessed as in read and written data?
* "$0.20/GiB-mo of high-speed cache" – how is the high-speed cache amount computed?
Out of curiosity, why did you choose EFS, it's insanely expensive at even modest scales?
This is what it looks like from the data. The key here is "diverging paths", which means different parts of the bridge were moving in different directions, putting strain on it.
While knowing this information is useful, most services fail in different domains and problems way before you reach that point. I'm not sure people really comprehend how hard you can hit a single machine before you need to distribute a workload.