Hetzner Object Storage
docs.hetzner.com
docs.hetzner.com
I remember the data loss from OVH where they put backups in the same building as the primaries and people only found out about this once a fire took out the whole building:
https://blocksandfiles.com/2023/03/23/ovh-cloud-must-pay-dam...
Isn't that exactly what Availability Zones are for? They're physically separate[0] datacenters and each one contains a copy of each S3 object (unless using the explicit single-zone options)
It's also straightforward (although not necessarily that cheap) to replicate S3 objects to another region
OVH’s failure was a single building. That’s the problem with a lot of server hosters - even Google has their availability zones all co-located in the same building, so a physical event like a fire could take down an entire region. AWS has AZs in physically separate locations, each with 1+ separate DCs.
I take the typical formulation (e.g., [1]), and translate it into:
- Keep 3 copies of your data: production + 2.
- Keep 2 snapshots, stored separately.
- Keep 1 snapshot on-hand (literally in your possession), or with a different provider.
It's great to see more options for "different provider". If I were an early adopter, I would either target this as my on-Hetzner snapshot of my Hetzner-hosted data, and replicate it elsewhere; or I would consider trialing it as a backup to non-Hetzner data. To your point, though, I'd probably wait on the latter until it has gone through some growing pains.
[1] - https://www.backblaze.com/blog/the-3-2-1-backup-strategy/
S3 egress, on the other hand, is so expensive you can often justify putting a writethrough cache for your entire working set at Hetzner...
edit: apparently not writethrough: https://developers.cloudflare.com/r2/data-migration/sippy/
Also R2 is extremely basic compared to S3. For example only trivial permissions are supported and auth tokens are bound to an actual account.
However they also have an option to copy files that are added to Tigris, to S3 automatically [1] (`--shadow-write-through`). I asked their founder if it's okay to use it as an extra redundancy continuously instead of a one-time migration, and they said they have no issues with it.
[0] https://www.tigrisdata.com
[1] https://www.tigrisdata.com/docs/migration/#starting-the-migr...
This is a massive claim that unless they're storing everything in memory I don't know how you can get close to Redis speed.
Even the flash-based Redis alternatives can't get close (but are good enough)
another story, a small company I worked for contracted with a small data center. They did things right and ran a tight ship. However something happened like a lightning strike or so other electrical fault and a critical component welded itself into a position they could t get out of. I wish I could remember the details better but they were down multiple days.
https://www.switch.com/switch-shield/
The Switch one I'm in (Grand Rapids Pyramid), the servers are about 3/4's underground.
From what I see, this actually might be a great backup/recovery solution for S3 in my case at least.
A great writeup of the incident and its context, three years afterwards: https://www.datacenterdynamics.com/en/analysis/ovhcloud-fire...
Shame their control panel is massively behind their competitors, it's the worst cloud control panel I've had the experience of using, even worse than Azure and Google.
It's slow slow slow slow slow and horribly buggy, did I mention it's slow?
From https://docs.aws.amazon.com/AmazonS3/latest/userguide/DataDu... :
> S3 Standard, [...] redundantly store objects on multiple devices across a minimum of three Availability Zones in an AWS Region. An Availability Zone is one or more discrete data centers with redundant power, networking, and connectivity in an AWS Region. Availability Zones are physically separated by a meaningful distance, many kilometers, from any other Availability Zone, although all are within 100 km (60 miles) of each other.
rclone to somewhere else at least.
And 3-2-1 rule for backups if you're serious.
The trick is not to put all eggs in the same basket and it's easier to do that when you're paying much less than you'd pay on providers like AWS.
Benchmark finished! block-size: 4.0 MiB, big-object-size: 1.0 GiB, small-object-size: 128 KiB, small-objects: 100, NumThreads: 6
+--------------------+--------------------+------------------+
| ITEM | VALUE | COST |
+--------------------+--------------------+------------------+
| upload objects | 207.02 MiB/s | 115.93 ms/object |
| download objects | 405.78 MiB/s | 59.14 ms/object |
| put small objects | 278.4 objects/s | 21.55 ms/object |
| get small objects | 498.5 objects/s | 12.04 ms/object |
| list objects | 40977.06 objects/s | 14.64 ms/op |
| head objects | 1995.4 objects/s | 3.01 ms/object |
| delete objects | 312.1 objects/s | 19.22 ms/object |
| change permissions | not support | not support |
| change owner/group | not support | not support |
| update mtime | not support | not support |
+--------------------+--------------------+------------------+
Currently have a project using OVH, which is slower, but than I guess there's currently less load on Hetzner. Benchmark finished! block-size: 4.0 MiB, big-object-size: 1.0 GiB, small-object-size: 128 KiB, small-objects: 100, NumThreads: 6
+--------------------+-------------------+------------------+
| ITEM | VALUE | COST |
+--------------------+-------------------+------------------+
| upload objects | 80.82 MiB/s | 296.94 ms/object |
| download objects | 274.61 MiB/s | 87.40 ms/object |
| put small objects | 31.6 objects/s | 189.98 ms/object |
| get small objects | 145.1 objects/s | 41.36 ms/object |
| list objects | 6293.39 objects/s | 95.34 ms/op |
| head objects | 177.7 objects/s | 33.76 ms/object |
| delete objects | 136.6 objects/s | 43.93 ms/object |
| change permissions | not support | not support |
| change owner/group | not support | not support |
| update mtime | not support | not support |
+--------------------+-------------------+------------------+Wonderful tool. Useful for figuring out where some of these cheaper S3 alternatives make their cuts.
+--------------------+-------------------+------------------+
| ITEM | VALUE | COST |
+--------------------+-------------------+------------------+
| upload objects | 16.72 MiB/s | 956.98 ms/object |
| download objects | 109.01 MiB/s | 146.78 ms/object |
| put small objects | 15.3 objects/s | 261.99 ms/object |
| get small objects | 88.8 objects/s | 45.05 ms/object |
| list objects | 6115.93 objects/s | 65.40 ms/op |
| head objects | 164.3 objects/s | 24.34 ms/object |
| delete objects | 157.9 objects/s | 25.33 ms/object |
| change permissions | not support | not support |
| change owner/group | not support | not support |
| update mtime | not support | not support |
+--------------------+-------------------+------------------+I can see that Hetzner is starting with WORM capabilities [0]. I wonder if this nascent product is successful they'll consider some of the features that other providers offer such as mutating objects, storage policies, tiering, and conditional writes. I appreciate you've got to start somewhere and this looks terrific for an opening gambit. Kudos Hetzner.
[0] > Object Storage is mainly used to store and share data as it is not possible to edit any data that you uploaded to a Bucket (objects are immutable). So the main purpose of Object Storage is "WORM", which is short for, "Write once, read many [times]".
My only hope is that they won't make it more complex than it should be in an effort to match AWS.
You can see it in the minio repo and issues .
There are a _lot_ of challenges in keeping the clusters secure, reliable, and performant. Make sure you have systems or tools in place to prevent abuse. Be aware of the little nuances of Ceph, like what time lifecycle policies kick off, or when dynamic bucket resharding will kick in (and block client writes!).
If possible, conduct extensive failure testing in a lab environment under simulated load to see how your clusters will really behave when it eventually happens. Triple check all of your tunables and your pool configuration. Some things like erasure coding profiles are set in stone, and once you have customer data on your clusters, there is no turning back.
GCP: $0.023/GB/month (standard storage, beyond 5 GB free limit)
Azure: $0.019/GB/month (hot, first 50 TB/month)
Cloudflare: $0.015/GB/month (beyond 10 GB free limit)
Scaleway: ~$0.013/GB/month (single-zone)
Backblaze: $0.006/GB/month
Hetzner: ~$5.49/month + $0.0054/GB/month (beyond 1TB free limit)
Egress costs vary though.
Is it typical to have such per-bucket limits? This means that you might have to take them into account when deciding on how to organize the data to avoid bottlenecks.
Those that do likely will already have an actual sales rep and the means to get these restrictions lifted, with a direct line to engineering.
Hetzner certainly markets to a different crowd, which I suspect is why these limits are being put in place ahead of time. It's very difficult to design around noisy neighbor, and will take some battle tested time in production to get right for the "average Joe" style customer who won't engage engineering resources to work with a provider to avoid hot spots and impacting other customers.
My suspicion was also that they're intentionally targeting a different crowd than S3 and focusing on their niche here.
[1]: https://docs.hetzner.com/cloud/technical-details/faq#what-ki...
I am not sure about the exact specifications anymore, but they are building dedicated servers for this with custom chassis where each server has a ton of drives and I think they each had 40 GBits networking. These are special servers that are not available to customers directly.
Amazon S3 pricing looking more and more sane.
It does not tell me if I should count hours in hour-of-the-day, or lapsed time. Also the example tries to demonstrate a case of "you don´t have to pay extra", but then falls silent. Nice, not the info I am looking for.
What about envisioning a customer who asks «what am I going to pay? Specify it right now, right here».
30 * 24 * 0.0081 = 5.83
28 * 24 * 0.0081 = 5.44
Admittedly I got my money's worth, for about 3 EUR for the very basic plan of 250 GB storage, I get upload speeds of about 6-9 MB/s and download speeds of about 25 MB/s, which is okay for what I'm trying to do. That said, it doesn't seem like there's a way to create additional users with access to only specific buckets not all of them for the same service and the overall offering does feel a bit jank (e.g. when you don't have an active service, you can't log into Contabo at all and while their Object Storage site is nice, the regular VPS one feels functional but dated).
What I'm saying is that when it comes to budget options, Hetzner will probably do great, since they have a good track record and it's not like there are that many other alternatives out there!
Of course, I could have also gone with just self-hosting with MinIO, Garage or SeaweedFS, as long as the VPS that I'd get would also have enough storage. It's nice that there are self-hosted options, too, I almost dread when I see bespoke object storage solutions at work, either developed before S3 was a thing or because the devs just didn't know or had a case of NIH, so I have to look at how a bunch of blobs are passed through servlets and serialized/deserialized, as well as deal with custom metadata and permission mechanisms.
It seems the Hetzner's docs say they use MinIO to expose their S3 managed service. I guess everything will be the same as MinIO upstream features and contigurations
[0] https://docs.hetzner.com/storage/object-storage/overview#obj... [1] https://docs.ceph.com/en/latest/radosgw/s3/
* Couldn't tell you, as they flagged my account and asked me for an ID Document I couldn't provide, Maybe I could have sorted it with Customer Service, but instead I just picked a different german hosting
I had some servers that I had autopay setup for I thought. But the bill went over as unpaid and they shut off the servers. Then I noticed & tried to login to their web account system but my account was locked. Support wanted me to send a wire transfer. I begged and begged, told them how it would cost as much as by debt, and they opened my account for ~36 hours, but after some delay, and I was away vacationing. I sent them the wire transfer almost two months ago now & still nothing from their accounts/billing people, even though the transfer had the account & payment #'s. It's been so frustrating, & so unnecessarily over the top bad & they seem to love making it worse every step. What the hell Hetzner?
I started using their services since last year for some of the operations, including business, and it's been rock solid. I plan on increasing usage, but will hold on this new object storage offering.
Perhaps you heard less about them because they're based in Europe and thus less hype sorrounds them?
So far I'm happy with the service. Linode wasn't bad per se, Hetzner just offered more for less. Their admin console is a bit more spartan in comparison to Linode's, but that's not a place I spend much time anyway. If Russia glasses Finland and Germany I'd have to find another provider but I'm pretty sure I'd have bigger problems at that point than where to park my wordpress blahg and irc session, lol.
You see, I host a cloud storage service on Hetzner. I have 12 PETABYTES of user data stored there. The data is spread over 120 of their SX type storage servers. Last week Hetzner sent me an email saying they were closing my account for an unknown reason. I have repeatedly been asking about the reason of this sudden closure, but they won't tell me. Meanwhile I have a huge problem, because I need to move 12 PB to a different hosting provider, and they only gave me until the end of November. That's an almost impossible deadline for setting up such a large storage cluster. Especially considering that Hetzner's 1 Gbps port speed makes it impossible to transfer the data in less than three weeks.
Don't trust Hetzner, guys. They screwed me over real bad here, and it seems like they will be taking the hardware that I have been renting there for 10 years for themselves now. This is incredibly scummy behaviour.
But don't you have that spread over 120 servers?
perharps if you moved all servers to 10G and payed for the egress over 20Tb they would reconsider your account?
They still unreliable pulling stuff like this but difficult to find these hardware options off the shelf, ready for order, at this rpice
I considered that and even suggested it to them, but they ignore all my remarks and keep repeating that I have two months to leave.
Hetzner does not host user facing servers for me. It's all behind caching nodes which are hosted by a different company. Hetzner would not know what is hosted on pixeldrain.
Frankly, your service looks a bit suspicious (free and cheap file sharing). You seem to be aware of the potential for abuse - your DMCA/abuse page even says there are a "large number of abuse reports pixeldrain receives every day".
It sucks that Hetzner is closing your account. But, they are known to be pretty conservative and your users have almost certainly violated their T&C many times even if you are making a good faith effort to prevent and respond to abuse.
I concede that my service has been abused a lot. I am working very hard to clean up my act. I have been implementing better content moderation, content scanning and streamlining my dmca handling process. If hetzner has a problem with any of that they should have just contacted me instead of pulling the plug like this. Very unreliable company.
I've been wondering for well over a decade why Hetzner didn't get its ass in gear and start offering AWS-like services. Instead all they've offered for years is remote control of VMs. Even now, this new S3 offering is very little, very late -- 18 years after Amazon first offered it and took off like a rocket.
The only odd thing about the timing here is what was keeping them.
Because unlike the hyperscalers, they are not natively a software company but a sysadmin company.
Besides the normal support channels, try reaching out on their subreddit and also post in the customer forum - to make this more public.
If they don't want your business, fine. But giving you only a couple of weeks to move 12 PB is unreasonable.
That they don't communicate well with you really sucks, and I will read your blog post. But please stay on the sane side of reasoning.
I can't think of a reason to cancel my account except that they want to free up resources for their own storage service. I am using a lot of disks and bandwidth. They might have been looking for a reason to kick me out and the launch of their own storage service could have been the final straw.
Keep in mind that the account cancellation mail arrived almost exactly one week before this announcement. They are keeping their lips tight about the reason of my account cancellation so I can only speculate.
By "operator pushdown", i mean any ability to filter or map over the contents of the object on the server side in some way, sending only the results over the network to the client.
For example, say you have a huge CSV file of customer orders in a bucket. You might want to find the timestamp of all the orders which included a particular product. If all you can do is stream the whole file, then you need to do that, just to pick out a few timestamps. But you could imagine a kind of request where you say "only give me lines where the product ID is P01234, and only send the timestamp column". Perhaps you would express that as a pair of regular expressions, or a sed program, or a Lua script, or maybe the server would understand CSV and let you write something a bit like SQL. There are all sorts of ways it could be done. Providing a fully general way might be tricky, but it wouldn't need to be fully general to be useful.
I appreciate that if you want to do this sort of access frequently, you should probably be using a database, not object storage. But it seems like a very useful feature to layer on top of object storage, and one that feels like it should be fairly cheap to execute - the server has to do a small extra amount of computation, but then needs a lto less network bandwidth.
Despite it being an awesome feature I've been itching to use, I've never actually found a use for it beyond messing around. Most places where S3 Select might make sense seems to be subsumed (for my uses) by Athena. Athena has a rather large amount of conceptual and actual boilerplate to get up and running with, though, S3 Select requires no upfront planning beyond building a fancy query string (or using their SDK wrappers)
Where S3 Select is likely to become fiddly is anywhere multiple files are involved. Athena makes querying large collections of CSVs (etc) straightforward, and handles all the scheduling and results merging for you.
:(
But you and patrickthebold are spot on in pointing out Athena. I've always thought of it as a database you load via S3, but of course it's equally a tool for querying data in S3.
Kind of simple and not that useful for dynamic queries but quite good for queries where you know the queries beforehand and can index accordingly.
https://docs.aws.amazon.com/AmazonS3/latest/userguide/range-...
I run minio on hetzner and wouldn't send stuff across some other network to backblaze.
We use it for multiple use cases and all of them require lots of retries and error handling
essentially they're build for backups use case, with lots of spinning rust, any kind of "working data" easily underperforms
backups work fine because not time sensitive and retries handle the problems
How much debt do they have on their balance sheet?
This is one of the rare cases where AWS is actually cheaper for us. It's probably more worth it if you have a lot of data. The only thing we're missing now is SES, and then we've fully migrated away from AWS :)
Keen to see what is available to protect public buckets though to prevent huge bills from malicious actors
They know what they’re doing and they’re doing it well, cheap is more side effect than primary goal.
> Our S3-compatible Object Storage provides you with storage capacity for saving data in "Buckets". Any data you save in your Bucket is saved in a Ceph cluster.
from https://docs.hetzner.com/storage/object-storage/overview#obj...
I'm speculating but those are at least the two I can think of that aren't explicitly linked to speed equivalency of a basic filesystem.
Up to 10 TB per object
Up to 1 GBit/s bandwidth per Bucket
Up to 1024 operations/s per Bucket
Up to 100 TB per Bucket
Up to 100,000,000 objects per Bucket
Up to 100 S3 credentials across all projects
Up to 10 Buckets across all projects
btw backblaze is the best object storage offer on the planet at this time(i have research all options). second would be wasabi.
Not knowing the pricing difference between the two and assuming they are similar, I would favor Backblaze as it would allow me to exceed the limit if I needed. Based on how you framed it, I would expect that with Wasabi you might hit a hard limit.
Wasabi doesn't have an egress hard limit and doesn't charge for egress. You get consistent pricing, and if you need to recover everything, it's not an issue at all, and you won't pay for it.
"The reasonable use egress policy indicates that if your monthly downloads (egress) are greater than your active storage volume, then your storage use case is not a good fit for Wasabi’s free egress policy, and we reserve the right to limit or suspend your service."
If you're using Wasabi for normal backup data storage, you shouldn't worry about egress. It is meant to prevent malicious users from uploading data and using up all the egress bandwidth, for example, a 500GB user egressing 5TB with public access or using Wasabi as a dump point to upload in one location and download in another region 1:1 ratio. As your storage goes up, your available consistent monthly egress goes up. It only becomes an issue if you abuse the account by uploading/downloading in a 1:1 ratio on a consistent basis.
What happens is you get an email from support asking if something changed in your use case. If so, they will help troubleshoot it(Think of a CDN scenario, where the CDN gets misconfigured). You also have an egress monitor for suspicious activity in case you aren't normally downloading all your data, and then you see a rise in egress https://docs.wasabi.com/docs/en/whats-new?highlight=egress#e....
In 2023 Veeam backup offloading to Wasabi S3, Wasabi "messed up their catalog" and lost data: https://forums.veeam.com/object-storage-as-backup-target-f52...
2023, files missing from bucket, Wasabi support not replying for days, then said they "had system maintenance": https://old.reddit.com/r/msp/comments/13dqhgr/wasabi_storage...
In 2021 Wasabi migrating databases, lost customer data: https://forums.veeam.com/object-storage-as-backup-target-f52...
I've been hit by something like that and had to re-push data to them.
It's nice to see more competition in this space.
I'm a big Hetzner fanboy, quite sad that pricing is't that competetive...
Well - not that well known.
I can find references to less egregious incidents - i.e. they ban running actual crypto software. Do you have any links I could read to find out more about the more extreme end of your claim? i.e. " storing files that have the word "bitcoin" or any other word that Hetzner does not approve of"?
If this is true then it's a reason for anyone to avoid Hetzner as accidentally triggering it would be possible on almost all websites. If you're exaggerating for dramatic effect then that's fine I guess - I'd just like to find out the real state of affairs.
so if you have some kind of service/api/trading bot/scraping/monitoring hitting a somewhat related to crypto node you may find that you will have no routing from hetzner core router to said server
and support is as friendly as other comments describe :)
In my experience, the support is not friendly, but it is no nonsense. Every time I've reached out, they responded quickly, tersely, and took appropriate actions. While I don't personally mind formalities, there's something to be said for their efficacy.