Can anyone speak to quality/reliability of other object storage providers that have S3-compatible (including presigned URL) APIs? S3's pricing is absolutely ridiculous by comparison, but they have the reliability argument on their side...
Can anyone speak to quality/reliability of other object storage providers that have S3-compatible (including presigned URL) APIs? S3's pricing is absolutely ridiculous by comparison, but they have the reliability argument on their side...
Well, unless your S3 buckets are in us-east-1. For some reason Amazon keeps having issues with S3 in that region.
Since the storage costs appear to be the same between Spaces and S3 ($0.02/GB/month) and neither charge for inbound transfer, I'm assuming your problem is with the outbound transfer pricing (S3 charges 9x what DO charges) and/or the per-request pricing. GCP's Regional Cloud Storage has the same storage pricing, even higher outbound transfer pricing, and the same request pricing. I haven't looked at any other providers, but if you want reliability, you're going to have to pay for it.
As this is mostly a side project, I think I can live with being a little adventurous.
I also looked at B2 [0] once or twice. The price is great, but the traffic cost (egress from GCE) renders it unusable for us.
But, you can build a poor-man's CDN -- varnish caches on DO/Linode/whatever where you get multiple terabytes of bandwidth for a small VM. So, you use the best object storage provider, but move most of the bits cheaply using Varnish + Route53 geo-dns.
There are actually a few options in egress land for us if cost is your primary concern.
If you're doing serving over http(s), you should probably be using Cloud CDN with your bucket [1] or put one of our partners like Cloudflare or Fastly with CDN Interconnect [2]. Both of these get you closer to $.04-$.08/GB depending on src/dest.
If not, and you don't care that we have a global backbone, you can get a more AWS-like network with our Standard Tier [3] (curiously with pricing squirreled away at [4], I'll file a bug). The packets will hop off our network in a hot potato / asap fashion, so you're not riding our backbone as much.
[1] http://cloud.google.com/cdn
[2] https://cloud.google.com/interconnect/docs/how-to/cdn-interc...
[3] https://cloud.google.com/network-tiers/
[4] https://cloud.google.com/network-tiers/pricing#standard_tier...
It's still way too expensive. And Cloud CDN isn't an appropriate tool for my use case. I really do just need a bunch of egress from a single location that isn't insanely expensive. $0.085/GB is in that insanely-expensive tier, for me.
I'd be happy to talk further via email; it's not secret, just not public.
Your bandwidth pricing is a joke. Yes, you got a nice network and yes you pay premiums to get transit of providers that are "hard to work with". And yes, you have dark fiber between your locations, which is costing a lot of money, but even considering those facts you are still charging at least 10x as much as your bandwidth should cost your customers.
How have you even calculated those prices? "Let's look at AWS and make it even more expensive"?
Not by enough, but they are.
For distributed object storage: I have also used MooseFS,LizardFS distributed object storage and MooseFS,Lizard runs very steady on production work loads. Steady as setup and then no ops issues.
Also to the short list is BeeGFS, BeeGFS is created by Fraunhofer is seriously fast distributed file system.
However, improperly designed/architected you will end up with serious scaling issues.
It's important to actually test things.
There is no other software, open source or otherwise that works quite as well as Ceph for providing durability and scale.
ScaleIO gets high marks for block storage performance compared to Ceph. It's not quite as durable and lacks some other features, but people seem to like it.
These sound like... problems with Ceph.
Sane deployment, management, and troubleshooting are critical features of any distributed system.
If they’re not there the system isn’t “ready”.
Personally, I feel like Inktank was pretty close and when they were acquired by Red Hat, progress seemed to slow to a crawl.
I haven't run the last couple of versions. It could be much improved. Certainly bluestore is promising for performance.
Salesforce runs several large Ceph clusters, and they have a dedicated team to run it. If you can't invest in the employees, you should invest in commercial support.
Salesforce also commits a lot of updates and patches back to the Ceph community
It has erasure coding as well. You could deploy on bare VM's with local storage in any cloud provider and have no dependency on network blocked storage.
With k8s 1.10 you get persistent local storage as well, as such you could probably build a fairly highly available system. Pro tip: do it in GCP as they have nice local SSDs you can attach to any instance. They're 375GB, 25k IOPS, and $.08GB, way cheaper than AWS I2 instances.
Right now the leading (uncomfortable) solution is probably DigitalOcean Spaces and a little bit of prayer.
[0] https://storpool.com/ [1] http://blog.inoreader.com/2018/03/success-story-inoreader-op...
This is the part that ruled it out for me