S3for.Me – cheap alternative S3 cloud storage
s3for.me
s3for.me
Yours: http://www.s3for.me/images/s3forme1.png
Salesforce: http://blog.database.com/wp-content/uploads/2012/11/intro-db...
Your S3: http://www.s3for.me/images/s3forme8.png
S3 from 2011: http://themetest.hollywoodtools.com/files/2012/06/s32.png
You: http://www.s3for.me/images/s3forme6.png Unbrekable IT (2011): http://www.unbreakableit.com/uit/wp-content/themes/unbreakab...
Your roadmap: http://www.s3for.me/images/s3forme4.png Oblaksoft (Nov 2012): http://www.oblaksoft.com/wp-content/uploads/2012/11/cloud-ch...
Albeit this following one could be a stock image: Your keyboard: http://www.s3for.me/images/s3forme5.jpg Cloudonlinebusiness (2010): http://www.cloudonlinebusiness.com/wp-content/uploads/2013/0...
Edit: Your TOS is Hetzner's almost verbatim. https://www.hetzner.de/en/hosting/legal/agb
Also, if you go to http://rest.s3for.me/ (the URL you use to calculate uptime) the error message in the document tree says 'The AWS Access Key Id you provided does not exist in our records'.
You mentioned S3 stands for Storage Should be Simple, what does AWS stand for in your company? Unfortunately cloud software isn't my forte, I would love to know what that acronym stands for.
My naivete leads me to believe that your status check is hosted with Amazon, and that your uptime checker would in fact be checking Amazon's uptime but I am surely mistaken.
We use this URL to monitor service uptime: http://rest.s3for.me/check/test.html
AWS stand nothing for our company, we don't use this acronym, but for Amazon it stands for "Amazon Web Services".
Here is our public status check: http://host-tracker.com/website-uptime-statistics/11891746/l...
HTTP/1.1 200 OK
Date: Sat, 07 Sep 2013 20:37:05 GMT
Server: Apache
x-amz-request-id: 1378586225522b8e71190b5
x-amz-id-2: storage1-1.s3for.me
ETag: "444bcb3a3fcf8389296c49467f27e1d6"
Last-Modified: Mon, 31 Dec 2012 20:09:52 +0000
Content-Length: 2
Content-Type: text/html
okWhy is your key id called 'AWSAccessKeyId', why does your error message say 'AWS Access Key Id does not exist in our records' why is your error message identical to the letter of what S3 would return. Surely the error message text isn't required for protocol validation.
This looks like you just took an open source S3 REST API clone (there are many) and stuck it on a Hetzner server without bothering to change any variable names.
To me there are a lot of questions, the most obvious tell is that I highly doubt Salesforce would use a stock image as the logo for their cloud database solution considering the level of investment they made in it.
This was the first thought - to take an open source S3 REST API clone, install it on Hetzner servers and work this way, but it is not the case. Any of the available solutions fit to us for different reasons. All core software is build by our team. We use Open Source software a lot, but the core of S3For.Me was developed from the first to the last line by our team.
I've checked Salesforce site and do not see anything similar to our logo. It will be replaced in the nearest future anyway.
Also, the math they use to get their 99.9999% durability is a bit sketchy - that 1% failure rate for a HDD isn't independent of its age, stress level, temperature, batch number, etc., and likely doesn't include the odds of corruption as opposed to failure. At scale you can't simply rely on your vendors' claims.
For 99.99999% durability their way of doing things is good enough to protect from HDD failures. Of course, there are other failure scenarios.
From who exactly?
Although Microsoft amusingly did quite poorly in two such cases over the years: http://en.wikipedia.org/wiki/Microsoft_vs._MikeRoweSoft http://en.wikipedia.org/wiki/Microsoft_v._Lindows
>We guarantee 99.9999% storage durability and 99.99% availability over a given year. Each byte is stored on three separate servers in two datacenters at the same time to achieve this level.
99.99% means you are violating your promise if you are down for 53 minutes over a given year[0]. What happens when (not if) this happens? If you would credit me back a couple of bucks and call it even, the whole four nine thing is useless. I don't understand the promise of high availability. On one hand it seems difficult to achieve and on the other hand it seems there is not much penalty for failing your promise.
[0] (3.15569e7*(100-99.99)/100)/(60)
As noted elsewhere this is hardly the only source of potential data loss however.
Large companies on the other hand need to have the reliability and security of a big name behind the storage. Why would they even consider moving to a no-name company that hasn't proven itself? Sorry, it just seems like the wrong niche to step in to.
I'd like to use DigitalOcean or Linode VPS machines for processing and as web servers, but then I need a large object store like S3. Rackspace and AWS have both object storage and VPS, but their VPS machines are underpowered for the price.
So ideally I would use Rackspace-Files or Amazon-S3 for storage, and DigitalOcean for number crunching, but I'd get killed on transfer rates from S3 to DigitalOcean (for example). Amazon and Rackspace have a trump card with free data transfer within their data centers. You have to use their slow machines to realistically use their object stores in this type of a use case.
So that's the 'S3' problem I need solved, which still isn't quite what S3ForMe is going after. Extremely cheap data transfers to/from major VPS providers.
Back to pgrove's point, though — there are trade offs with using a dedicated server. The storage isn't as reliable as S3 and the server itself is likely less reliable than a VPS (since some hardware problems on a VPS can be solved by migrating to new hardware and only involve a small downtime, rather than the large downtime of restoring from backup).
A major advantage to dedicated servers that I don't hear much about is simplicity. Dropping everything on a dedicated server means I can use local services like the file system rather than having to deal with S3 and latency and bandwidth costs.
Edit: I just realized that the idea was to use Rackspace files + one of their servers rather than S3 + EC2. Somehow I missed that bit and was thinking about just storing everything on the dedicated server.
I can't imagine why it is marginally cheaper.
Pro: - Servers are located in Europe (WOOOW, fuck spying govs)
Not so good: - I think you had a great idea, got one (maybe 2,3) developers working on it for some time and rolled out this product. Which obviously is quite unfinished (on the business point of view, i have no hint on how the technical side is working as i never used it).
So, question for you guys: How the hell invoicing is working? As you could guess most of us work/own companies, we need freaking invoices for every cent we spend. No billing, no invoices, nothing like this in the private area, ideas?
That said, if i get (as an example) a search warrant from a US court and i'm based in EU, i will simply ignore it. This is quite enough.
This is not mentioned on the site's FAQ.
What is S3for.me's policy regarding NSA warrantless surveillance? Will S3for.me release a NSA transparency report, like Yahoo?
http://en.wikipedia.org/wiki/NSA_warrantless_surveillance_%2...
Recent developments from Mr. Snowden show, that BND regularly sniffs and provides data to the NSA.
If you want to find out more about this, try to find a translation of the TKÜV (Telekommunikations-Überwachungsverordnung). It was established in 2002 http://www.heise.de/newsticker/meldung/BND-ueberwachungsschn...
Update: http://www.bmwi.de/BMWi/Redaktion/PDF/Gesetz/TKUEV-deutsch-e...
The obligated party shall tolerate the installation and operation of equipment of the Federal Intelligence Service in its premises which shall only be installed and serviced by Federal Intelligence Service staff specially authorised and shall meet the following requirements:…
Otherwise good stuff!
BTW: websites hosted by Germans are required to have an imprint ("Impressum") with full contact data incl. phone number.
Starter questions:
How does data placement work? How is data checked for correctness? How do you do listings at scale? During hardware failures, do you still make any durability or availability guarantees? How do you handle hot content? How do you resolve conflicts (eg concurrent writes)?
> How is data checked for correctness? With checksum. Once the data is uploaded by user he will receive its checksum and must compared it with local checksum to make sure that it was correctly transfered and stored. The same checksum is used to ensure server-side data correctness.
> How do you do listings at scale? There is a trick - we support only one delimiter (/), this means that we can use very simple listing algorithm which scales very easy.
> During hardware failures, do you still make any durability or availability guarantees? Yes, all data is split in server groups by 3 servers each. If one of 3 servers will fail, this group will still running like nothing happened, some running requests may fail though. If 2 servers will fail at the same time, then this group and all data in it will be put in read-only mode to avoid any possible data damage.
> How do you handle hot content? It is cached in RAM by OS, we do not perform any additional measures. OS does a pretty good job.
> How do you resolve conflicts (eg concurrent writes)? Some conflicts are resolved by the software if possible. Unrecoverable conflicts are returned back to user with HTTP 400, 500 errors to make him know that something is wrong and he must run request again. For concurrent writes we use simple rule - the last one wins.
1) Server groups of at least 3 mirrored servers, with a max of 5TB.
This seems like an interesting design choice. What do you mean by "at least"? Does this mean you'll have some data with more replicas? Are these server pools filled up and then powered down until they are needed? How do you choose which server pool to send the data to? And since you have a mirrored set of servers, when do you send a response back to the client?
Is the 5TB number something that is a limit for the storage server (ie 15TB total for a cluster of 3)? That seems rather low. It also doesn't divide evenly into common drive sizes available from HDD vendors today. So what kind of density are you getting in your storage servers? How many drives per CPU, and how many TB per rack? Since you're advertising on low price, I'd think very high density would be pretty important.
2) You say you split data into smaller chunks if it crosses some threshold. Let's suppose you split it into 1MB objects, once the data is bigger than 5MB. And each 1MB chunk is then written to some server pool which has replicated storage (via the mirroring). How do you tie the chunks back to the logical object? Do you have a centralized metadata layer that stores placement information? If so, how do you deal with the scaling issue there? If not, another option would be to store a manifest object that contains the set of chunks. But in either case, you've got a potentially very high number of servers that are required to be available at the time of a read request in order to serve the data.
Just as an example (and using some very conservative examples), suppose I have a 100MB file I want to store and you chunk at 10MB. So that means there are now 10 chunks replicated in your system, for a total of 30 unique drives. Now when I read the data, your system needs to find which 10 servers pools have my chunks and then establish a connection with one of the servers in each server group. This seems like a lot of extra networking overhead for read requests. What benefits does it provide that offset the network connection overhead?
And what happens when one of the chunks is unavailable? Can you rebuild it from the remaining pieces (which would essentially be some sort of erasure encoding)?
Overall, the chunking and mirroring design choices seem to me like they would introduce a lot of extra complexity into the system. I'd love to hear more about how you arrived at these choices, and what you see as their advantages.
In order to not make my long post even longer, I'll not pursue more questions around listings, failures, or hot content.
2) We have a centralised replicated metadata layer which is stored on the same servers as the data itself. All object chunks are stored at one servers group at the same time, so there is no need to connect to multiple servers to serve a file, it is enough to connect to one server from server group to get all the data. Metadata may be stored at different server group though. All chunks are replicated to 3 servers at the same time using a sequential append-only log to ensure that all servers has the same data. This may introduce replication lag and if it is too big for some server then it is removed from the server group until replication lag back to normal (1-3 seconds usually).
Actually, it is much simpler than I explained, data layer with replication, data consistency and sharding is completely transparent to the application layer and it is really-really small and simple. Email me at support@s3for.me and I'll share with you software details and you will understand how simple it is.
Amazon (standard): $0.095/GB/mo |transfer out: $0.12/GB/mo
S3for.me: $0.060/GB/mo |transfer out: $0.04/GB/mo
Also I'd consider changing the font/color on s3forme so that it 'pops' a bit more compared to your competitors. As is my guess is the user's eye is drawn to Rackspace.
I can't think that those prices/durability promises are realistic unless you're operating at massive scale.
like a ddos or a very popular video
how many servers can transmit one file?
I can't find anything on their website that lists physical contact information.
Right now you are a unknown player with a unknown product running on top of a bargain-basement server provider.
I wish you the best of luck but I'll be surprised if you don't slam into a growth wall and have to raise prices.
I'm sure you still have valid reasons why you rolled your own, and I would be interested in hearing about them.
Personally I think that we made a right choice and I don't regret about it. We learned a lot about cloud storage, scalability, performance, possible problems and I believe that this knowledge is very important for every team who works in cloud storage business.
Which is interesting, because Hetzers standard bandwidth overage seems to be €2/TB (~ $0.26/GB)
Based on my experience this should be really good.
I wouldn't want to promise that.
It looks like an interesting storage service, but like others mentioned, you should consider renaming it. I was convinced you were reselling Amazon S3 storage when I visited the site.