Backblaze seemingly does not support files greater than 1 TB
wadetregaskis.com
wadetregaskis.com
I've been trying to replace all my various backups (Time Machine, Backblaze, CCC) to use a single tool (Arq - https://www.arqbackup.com/)
The Backblaze client is the next to go. To be honest, I haven't had too many issues with it, but the restore interface in particular is pretty poor and slow.
Of course, being bad at being bad could just be a different kind of bad. A good BadChunkRecord explains the problem with the chunk. A bad BadChunkRecord might have too little information to be a good BadChunkRecord. A bad BadBadChunkRecord could be an unhandled exception with no other details and the fact that a ChunkRecord is even involved is assumed and therefore questionable.
> Error: /Users/tzs/Library/Biome/streams/restricted/ProactiveHarvesting.Mail/local/91257846325132: Failed to open file: Operation not permitted
but when I check that file I have no trouble opening it. I can't see anything in the permissions of it or any of the directories on the path that would permit opening it.
Then I'd have to search the net to find out what the heck that is and whether or not it is safe to add an exclusion for it or for one of the directories on the path.
I eventually figured out that before searching the net what I should do is create a new backup plan and take a look at the exclusions in that new backup plan. Often I'd then find that there is a default exclusion that covers it. (In this particular example ~/Library/Biome is excluded by default).
When they update the default exclusion list that is used for new backup plans it does not update the defaults in existing backup plans. Evidently either Biome did not exist several years ago when I made my backup plan, or it was not a source of errors and so was not in my default exclusions.
So now I occasionally create a new backup plan, copy its default exclusions, delete the new backup plan, and then compare the default exclusions with those of my backup plans to see if there is any I should add.
The documentation implies a 10TB limit for large files on B2.
First - it's sensible to have limits to the size of the block list. These systems usually have a metadata layer and a block layer. In order to ensure a latency SLO, some reasonable limits to the number of blocks need to be imposed. This also prevents creating files with many blocks of small sizes. More below.
Second - Usually blocks are designed to match some physical characteristics of the storage device, but 10 MB doesn't match anything. 8MB might be a block size for an SSD or a HDD, or 72MB=9*8MB might be a block size with 9:16 erasure coding, which they likely have for backup use-cases. That being said, it's likely the block size is not fixed to 10 MB and could be increased (or decreased). Whether or not the client program that does the upload is aware of that, is a different problem.
But I doubt they're getting much support burden from this; the product targets home users, providing fixed-price single computer backups. This isn't their B3 blob storage offering.
The most bloated triple-A games out there are only ~100GB. The entirety of OpenStreetMap's data, in XML format, only reaches 147 GB [1]. The latest and greatest 670-billion-parameter LLM is only 600 GB... and they ship it as 163 x 4.3GB files [2]. Among PC gamers [3] about 45% of people have less than 1TB of hard drive space, total.
IMHO it's very unlikely they have a big support burden from this undocumented limitation.
It's clearly a bug though - trying to upload the same file in every single backup, failing every time? There's no way that's the intended behaviour.
[1] https://planet.openstreetmap.org/ [2] https://huggingface.co/deepseek-ai/DeepSeek-R1/tree/main [3] https://store.steampowered.com/hwsurvey
Seems like there is a bug somewhere - whether artificially imposing a file size limit, or in not imposing the limit correctly.
”Should” most certainly. However, often client teams don’t work so closely with backend teams. Or the client team is working from incorrect documentation. Or the client team rightfully assumes that when it reports a file size before uploading, the backend will produce an acceptable error to notify the client.
What else would I assume?
If there's a 1TB limit I would expect that to be described as "create files up to 1TB in size".
I guess different kinds of people are drawn to different careers, and people with loose morals, hubris, and a propensity to lie are drawn to marketing. Don't even need an incentive to lie, it's just who they are.
I think you will find that there are people that see "limit <humongous>" and pick the "no limit" or no limit stated option.
> people with loose morals, hubris, and a propensity to lie are drawn to marketing
You are jumping the gun here.
I've had many issues with it, and poking around at its logs and file structures makes it pretty obvious that it's very bad code. There's numerous different inconsistent naming conventions and log formats.
I once tried to edit the custom exclusions config file (bzexcluderules) and gave up shortly after because the format of the file is both completely nonsensical and largely undocumented - whoever implemented didn't know how to do proper wildcard matching and was likely unaware of the existence of regular expressions, so they came up with an unholy combination of properties such as ruleIsOptional="t" (no idea what this means), skipFirstCharThenStartsWith=":\Users\" (no comment on this one), contains_1 and contains_2 (why not contains_3? sky's the limit!), endsWith="\ntuser.dat" and hasFileExtension="dat" (because for some reason checking the extension with endsWith isn't enough?), etc.
What are a couple that you particularly recommend for this? I've been using B2 after abandoning the disaster that SpiderOak has become (was a long term user and stuck out of habit), and not exceptionally happy with B2 either.
Since 2 minutes isn’t long enough to upload 1 TB, it’s either looking at the blocks out of order and skipping ahead to block 100,001 for some reason or noticing that’s the first block that hasn’t been uploaded yet.
Another comment suggests 10 MB isn’t a fixed limit. In that case, I wonder if 100,000 blocks is the limit and the intention of the designers is that users with such large files would increase their block size. If so, that should be documented of course, and a scenario. Although it’s still a bit strange that it would apparently try to upload the 100,001st block and not immediately warn the user that the file size is incompatible with their block size.
I just found this out the hard way.
Mac: https://www.backblaze.com/computer-backup/docs/configure-exc...
Windows: https://www.backblaze.com/computer-backup/docs/configure-exc...
(Disclaimer: I'm a BB customer with 70TB backed up with them.)
Maybe you mean that it doesnt do image-level backups by default?
Now imagine how much data the Large Hadron Collider can generate in a second.
I recently was fixing my own tool here and had to shufle some fields. I settled up at 48bit file sizes, so 256TB for single file. Should be enough for everyone ;)
Having a 1TB file sucks in way more ways than just "your backup provider doesn't support it". Believe me.
One example: A Git repository that I created contained files that are above GitHub's file size limits. The reason was a bad design in my program. I could fix this design mistake, so that all files in the current revision of my Git repository are now "small". But I still cannot use GitHub because some old commit still contains these large files. So, I use(d) BitBucket instead of GitHub.
Of course, all the commit identifiers change.
These old files nevertheless contain valuable data that I want to keep. Converting them to the new, much improved format would take serious efforts.
Are you really sure you cannot rebase/edit/squash the commits in that git repositories?
Yes, commit hashes will change, but will it actually be a problem?
This was just one example. Let me give another one:
In some industries for audit reasons you have to keep lots of old data in case some auditor or regulator asks questions. In such a situation, the company cannot say "hey sorry, we changed our data format, so we cannot provide you the data that you asked for." - this will immediately get the company into legal trouble. Instead, while one is allowed to improve things, perfect capability to handle the old data has to be ensured.
You can do that on shared repos too but that would cause a bad headache to other maintainers.