Backblaze submitting names and sizes of files in B2 buckets to Facebook
twitter.com
twitter.com
This should be a warning to every developer, when you integrate with any third party - don’t just copy the default snippet of code they recommend, read the docs thoroughly and then test/monitor what is being sent back to the third party. And if you are writing a third party widget, be considerate and make the defaults the least aggressive possible when it comes to accessing user/site/server data.
Trust, but verify. Especially FB.
I see a lot of comments about a paid service using tracking pixels. I understand the backlash but this is also how companies that are taking your hard earned money improve and iterate on their products. I know people don’t like tracking but usage analytics are invaluable to improving the product. There’s only so much customer engagement you can do, and often that leaves really unexpected outcomes in the dark (if you asked most customers, they’d ask for a faster horse).
The issue here is that Backblaze injected untrusted code into their core product which by its very nature is extremely sensitive. I’m not really sure what they were thinking. Rolling your own tracking is a pain, but some security contexts demand it.
In the particular case of Facebook many of us simplify to "don't trust if not absolutely necessary" ;-)
In the particular case of Facebook many of us simplify to "don't trust"
There is no upside for Facebook to have anything about me.
But I do make choices where I can.
If you “trust Facebook to fuck [you] over” then you’re not trusting Facebook. You’re expecting Facebook to fuck you over and minimising your exposure to that process.
Indeed, if one didn’t have forbearance of (say) a software vendor or perhaps a dependency then one would be entirely free to not use that vendor’s product or that particular package. Similarly, if you don’t swerve right to prudentially make space for an approaching vehicle to serve into your path then you’re showing them forbearance and trusting them (or swerve left, depending on one’s location on the Earth what side of the road is legally mandate).
So, yep: those are perfect examples of Trust as embodied in the definition I presented above. And of course it lies at the root of initiatives such as Trusted Computing Initiative & cetera.
I wasn't aware that Facebook offered a Google Analytics style service, I was only aware of the Facebook Analytics for your Facebook page.
Can you provide links to this service, as my Google-fu doesn't return anything.
My guess would be that developers installed a single Google Tag Manager script and left the tracking to the marketing/analytics team. From that point they manage what third party scripts are added to the site, not engineers.
This situation cases a tunnel vision which can unfortunately lead to situations like this.
What analytics do you think they can do using Facebook that they cannot do from their own httpd access logs, a technology that's 30 years old at this point?
If you went into a retail store and an employee followed you around the whole time with a notebook and stopwatch writing down everywhere you walked and every product you looked at, you would rightly be creeped the fuck out and tell him to stop.
This is exactly what online tracking is, but done virtually.
I hate to break it to you, but retail companies are doing this today using security camera footage, to figure out what parts of the store customers start at, or spend the most time in..
https://www.nytimes.com/interactive/2019/06/14/opinion/bluet...
For the former purpose, it would generally be sufficient to inform visitors with a sign on the entrance with legitimate interest clause; for the latter example IMHO the only practical compliant solution would require anonymization of the data, so you could make and store density data iff you don't have any way to tie them back to customer identities including the purchases they made, which is a key difference from the facebook example, which (as far as I understand) uses unique IDs to link the conversions to specific FB accounts.
I will offer you another perspective to consider.
Technology is extremely subversive in that it bypasses all of our brain's instinctual responses. Someone or something monitoring and tracking you should be setting off warning sirens in your brain. At best they are trying to study you, at worse they are trying to exploit or harm you.
Through hundreds of thousands of years of evolution our brains have built up warning systems to make us feel fear and unease when we realise we are being tracked. But since humans have spent 99.99% of evolution entirely in the physical world these systems have no concept of the digital.
The reason you feel extremely uneasy when being monitored by a person, but not when being monitored by a computer system that is collecting the exact same information (or more), is that your subconscious brain doesn't understand computers.
>The reason you feel extremely uneasy when being monitored by a person
I don't though. If someone walked up to me and asked me what my favorite color was and they wrote it down I don't feel any negative feeling.
Even if it was a true thing just because we have a warning system that doesn't mean that something is actually bad. Tracking just gives people more information to allow for better decisions to be made and it can make things more efficient.
It's most exemplary case is the use done by the security teams in most IT shops: "Yes, we did put endpoint protection on your laptop, it's not that we don't trust you -- but you know what they say...". I wish they would simply say:"Of course we don't trust you, not on this. You are going to click on that shady link or download and run random executables from the interwebs. But we do trust you're good at your job".
It's essentially "audits instead of pre-approval".
But the level of irresponsibility with customers' most private data from a company who's MAIN JOB is to protect it is absolutely shocking. Yes, it's the freaking pixel that does the tracking, BUT IT'S THEIR RESPONSIBILITY TO KNOW WHAT THE HECK IS HAPPENING ON THEIR WEBSITE. Don't they have any sort of vulnerability assessment or security code review? Their reply tweet is almost as infuriating as how something like that could even happen to begin with. Like yeah, sorry, ain't our fault! It's astounding.
I considered using backblaze a couple of times and now I'm very happy that I eventually didn't.
I bet people just don't realize that the frontend code could end up being a source of major data confidentiality vulnerability. The threat modeling, auditing etc. usually just concentrates on attack scenarios involving the backend to save money and keep frontend development a bit lighter on the security review process side.
Does not make it excusable of course, just means their threat modeling was inadequate. But probably explains how this was able to sip into production.
Facebook want data on what actions users took before signing up, which users actually signed up and started paying, and how that relates to revenue. This UI is exactly where they can determine these types of actions.
Whether this actually makes Facebook better at marketing or not is a good question.
Killer mistake of your revenue, or killer mistake of what we'd prefer apps to be?
I haven’t seen any evidence that they do. The last time I brought up their history of bad security practices on HN, one of their co-founders decided that the correct course of action was to come on here, accuse me of being a bad actor, and repeatedly make up quotes I didn’t say.[0] All because I tried to warn others in the community that something just like this was likely to happen again. And now it has. So, you know.
As far as having an axe to grind goes… if wanting to protect others is grinding an axe, I guess I’m guilty of that. I don’t feel like a handful of topical messages warning people of a legitimate and clearly ongoing problem is some abusive behaviour on my part, but maybe I’m wrong. I’m happy to learn from others’ perspectives, since I’m sure I could be a more effective communicator.
[0] I suspect he was the one who replied to my vulnerability report since the same attitude was on display in those messages too, but I don’t know since that account was just named “bbqa”.
Same. I ended up going with spideroak because it was too hard to figure out what other providers were even promising for privacy, and it’s been fine. I’m thinking maybe of trying rsync.net with duplicity one day
Unless your business model is literally collecting all personal and private information, then don't do it.
Here's the thing though, being a third party, due diligence means you constantly have to check what they're doing.
Because otherwise how do you know when they suddenly change their script to do something entirely different?
This is why having any 3rd party scripts on a dashboard service like this is in my view entirely inexcusable.
When I go to visit my cloud-based dashboard for Acronis' backup service[1] there's one single domain involved, and that's how it should be.
https://developer.mozilla.org/en-US/docs/Web/Security/Subres...
Or is there a way for the server to specify all resources must have SRI?
You can combine CSP it with SRI hashes and also report violations to the backend.
This trend of dynamically linking to other people's shit needs to stop.
Even though it's inconvenient maybe we should treat it as just another 3rd party dependency that needs to be downloaded, screened, and then used from the internal store. Pretty dangerous to dynamically load a script from a site like facebook.com.
(a) Sandbox it.
(b) Have them sign a contract they won't do so but only do these N things, and if change those without telling you, sue them.
How much "due diligence" should be required to know that adding anything from facebook INSIDE your backup / file storage product is a bad idea... hell adding anything from facebook at any product should be considered a bad idea...
It's certainly a problem, but the biggest problem is simply that Backblaze doesn't need to integrate Facebook pixel to their web interface, especially when users are logged in.
There is absolutely zero benefit to this for their users, but I also fail to see what it could bring to them as a company??
They have now created a major PR problem for themselves, and what did they get in exchange?
(Ok, as the saying goes, there's no such thing as bad publicity. But still.)
It's disappointing that this is what the internet has become, and I'm happy to see issues like this brought forward. I can see value in feeding back conversion events to an ad network, but the "just let us run code on your page" style of integration needs to stop. Give developers some API to explicitly send such an event, if they really need to.
If you want to retarget me on FB, fine, do it on a marketing pages of a website, but don't let the trackers anywhere close to my private data.
Our permission schemes are way too broad in general. I have done some tests with Zapier and IFTTT. For every integration you consent with them seeing ALL your data, being able to modify and so on. You also don't get a log how often the permission was used.
This means that any network emissions you add to your site would have to be categorized by URL or domain instead of relying on shaky special-cased integrations that can fail as soon as an API changes.
– Facebook's pixel will, by default, attach click listeners to the page and send back associated metadata. This simplifies implementation needs, but can create unintentional information leaks in privileged contexts. To disable this behavior, there's a flag[1] you can use. After which, you can manually trigger FB Pixel hits and control both when they're fired and what information is included in them.
– There's a feature for the FB Pixel called Advanced Matching[2] that allows you to send hashed PII as parameters with your FB events. "Automatic Advanced Matching" can be enabled at any time via a toggle in the FB interface. I believe that setting autoConfig to false as mentioned above will similarly prevent Automatic Advanced Matching from working (since it disables the auto-creation of all those listeners to begin with). When manually triggering pixel calls like above, you can use this functionality via "Manual Advanced Matching"[3].
As a general rule, I'd strongly encourage anyone implementing a Facebook pixel to also include the autoConfig = false flag. This makes it work like most other pixels, where the base tag just instantiates an object. After which, hits only occur when explicitly defined in the site code, and include specifically what details you include in it. That way you're fully aware of the scope of data disclosure happening and any need from marketing to include sensitive (or potentially sensitive) information in these calls has to be explicitly requested (and theoretically vetted) as part of the standard dev process.
[1] https://developers.facebook.com/docs/facebook-pixel/advanced...
[2] https://www.facebook.com/business/help/611774685654668
[3] https://developers.facebook.com/docs/facebook-pixel/advanced...
This sentence, your thesis, is absurd. Where does it say they make money from selling data?
Furthermore, "I guess they can’t afford to cover costs given current prices" is a really strange foundation to leap from. Do you have any facts? Or are you just speculating on BackBlaze and making wild assumptions?
> now that other like-minded crazies can find each other faster than ever
You are right, there is nowhere else online where "crazies" gather - not 4chan, not reddit, not voat, not twitter, just Facebook?
But the previous two points make no difference on why tracking is baked into an internal portal so I speculated as to the reasoning as anyone would do — it’s not a stretch to think that data owners could sell customer data to an aggregator like Facebook for an additional line item of revenue.
Because it does? There are valid reasons for using the pixel in advertising.
What would be their motivation for just giving their customers data to Facebook for free?
This is plain wrong. One of the reasons I was a fan of B2 was their tech and how they achieved such low costs:
https://www.backblaze.com/blog/design-thinking-b2-apis-the-h...
https://www.backblaze.com/blog/backblaze-and-cloudflare-part...
While this FB pixel debacle is obviously a very big screw up, it's pretty much a "screw up" and unintentional from what I understand so far. And they have fixed it already which is a positive step towards redemption.
From my speculation - the screw up seems to have happened from including the googletagmanager. They probably only wanted it to stay on the home page of B2 (for ad conversion tracking if I were to guess), not on the dashboard itself after login. The screw up caused it to be on the dashboard too.
But for those who are implementing FB Pixels, I wanted to put out some potentially useful information that can help protect against unintended data disclosure, after mentioning the auto-listener behavior in a reply to another comment and being met with surprise[1].
ie they don't really want your truly-security-critical customer data. But if they can boost their conversion rate with sites like dogfoodreviews.com by 5%, and the price is sending backblaze.com's fantastically-sensitive paid customer data into an unsecured data path, they will absolutely do it.
Comparable to the absolute havoc that Zoom wreaked on browser security to save one click on starting a call.
> ie they don't really want your truly-security-critical customer data.
It's both. It eases implementation with a one-and-done snippet, and then slaps a user-friendly GUI on the other side for marketers to sort through the firehose and use what they want.
While making it also trivially easy for marketers to toggle a button that OKs the turbo-boost mode that siphons up (hashed) sensitive customer information, which can then be used to claim credit for additional conversions by cross-referencing the (hashed) PII siphoned up against what Facebook has for those exposed to your ads.
> Why is that? Perhaps they get business value out of it?
Oh, they get value out of invading my privacy? Carry on then!
Ignoring the situation and context to make a comical statement doesn't really add anything to the discussion.
And all targeting Ads is bad. Because by their definition, All ads are tracking ads.
This mentality also fits the current Internet and Twitter narrative. Especially true when it is from Facebook. Which happens to be pure evil on HN, twitter sphere and MainStream Media.
Ads will at best waste my time/bandwidth/processing power and at worst compromise my privacy and/or convince me to make a bad financial decision.
I don't see why this mentality is wrong?
One thing if those micro-payments end up enriching the creators, but another when blackholed into the coffers of a few tech cos.
> but the fact that others don’t allows them to micropay for a lot of content with their time.
I'm not objecting. I do the same. But in what sense are you paying? I'm sure I'm not.
The creator isn't getting any benefit from it, but I'm still paying.
John Gruber has an ad at daringfireball.net, currently for a company called Simris. IIRC the ad is pure text, not loaded by a script, and does not track you. Other blogs (usually security professionals ime) have text-based ads that are probably part of the theme in a static site generator.
I know all of what you said, but somehow HN still thinks Ads are bad.
Either way, if you must do ad tracking, do so on your homepage. Once the user is logged in and has paid you money for a service there shouldn’t be any ads nor tracking.
Yes, that's covered by this being a mistake in implementation as I said.
> "there shouldn’t be any ads nor tracking"
Again, based on what exactly? Finding new users that are similar to your existing customers is a completely valid strategy.
Most people in this thread are making wild statements from the typical emotional/outrage driven pile-on when anything happens.
> Finding new users that are similar to your existing customers is a completely valid strategy.
But this can be achieved with tracking in the homepage without embedding trackers in the actual product right next to sensitive data?
> Most people in this thread are making wild statements from the typical emotional/outrage driven pile-on when anything happens.
This doesn't make these statements any less valid though? Most people are indeed outraged that a paid professional product is ratting them out to Facebook which makes total sense as nobody would've expected that.
Yes, it was an implementation mistake. How many times do I have to repeat that? See, this is the outrage that doesn't even read the actual comment.
What on earth does “valid” mean here? It’s certainly not acceptable (to me as a customer) if it involves exposing your existing customers to these risks. Those ends can not justify those means.
Customers weren't intentionally exposed to that risk nor was it part of a trade-off, it was an implementation mistake for many reasons, something I've repeated 3 times now. What is so complicated to understand here?
Under your definition, there can never be mistakes at all; which is impossible.
I'll resist the temptation to draw a parallel between "advertising services from Facebook or other companies" and a crime syndicate.
If you have a real rebuttal against advertising then reply with that instead and we can discuss how technical implementations can be fraught with security mistakes and errors, regardless of industry or product.
Why not? Technical implementations can always be discussed separately from the context they're used in, and even your extreme example of guns has perfectly valid uses in the police and military. Yet you're making the strange comparison to being "a gangster". Why? What's the point of this convoluted analogy?
> "sharing PII with Facebook is something most web sites should avoid, not something they should do properly"
Again, why? You seem to claim a lot without any basis. Data has valid uses, and being used properly is foundational to providing privacy.
You seem to believe that this particular breach is accidental, but reckless incompetence on Backblaze's part isn't much better than deliberate disregard for user privacy: any online service from Facebook should raise a red flag.
Are there no browser level protections for this type of thing? I thought CORS was supposed to prevent these activities from happening.
As for GTM – a deployed container is self-contained. If you don't want to expose your site to third party code, but want to use GTM as a convenient control plane for configuration of tags and tagging rules, you can do that. Instead of using the standard snippet that loads the container from Google, you can just grab the generated javascript file for the container after a new deploy and self-host it. It gives you the convenience of GTM (central control plane for tagging-related stuff, versioning and commenting, etc) but without the security exposure of embedding externally hosted scripts.
[1] https://developers.facebook.com/docs/facebook-pixel/advanced...
Here we are talking about a tracking _script_ embedded in the page and sending to Facebook everything the user does (“standard or custom events triggered by UI interactions”).
Using only a pixel to track how users move around the app wouldn’t have landed Backblaze in as much hot water. Instead, it looks like the Facebook _tracking script_ (automatically) exfiltrated sensitive data like file names, and that crosses a limit.
As I mentioned in my original comment, the tracking scripts are more than just generator functions for the image pixels. They also do stuff like browser fingerprinting and cookie management[1], and ensure these things get tacked onto generated pixel calls. This improves the fidelity of the data sent back to Facebook, but ultimately it all boils down to image calls with tracking data tacked on as query parameters to the call.
The reason Facebook (and others) don't recommend doing this is because
– As you mentioned, they have way more freedom to do what they want on the page when you load their actual script. So of course that's going to be their preference.
– Advertisers use these pixels for attribution purposes, but ad networks also use the opportunity to further fingerprint and profile users for targeting within their platform.
– The tracking script abstracts away the actual tracking protocol being used (i.e. the query parameters and their associated values). Which helps ensure calls are made correctly, as well as provides flexibility to make changes in the underlying protocol while retaining a stable interface via the JS SDK.
- Takes care of things like generating a unique user id, looking for and saving Facebook Click IDs when seen on incoming traffic, and tacking those values onto pixel calls when they occur.
Any user ID can actually be used, so long as it's unique (and Facebook's methodology is documented and easily replicated in [1], if you want to be consistent with the SDK). And persisting a query parameter into a cookie is actually more robust if done by a first-party script, since ITP has made the lifespan for cookies written by third-party scripts so short.
As long as your custom image generator accounts for those two components (generates a client id if none exists and persists + includes a fbclid if seen on incoming traffic), you will get close to parity with the JS tracking library as far as attribution in Facebook Ads without any need to load third party scripts from Facebook (or other advertisers). Which, as an advertiser, is the only part that you care about. What isn't at parity is all of the secondary fingerprinting that ad networks do, but that's the ad network's problem and preventing that shady shit from happening on your site is the precise reason you'd want to roll your own tracking calls to begin with.
[1] https://developers.facebook.com/docs/marketing-api/conversio...
Server side analytics will prove much more powerful and opaque when it gets integrated deep enough into web dev stacks to work properly.
The upside of third-party trackers is that you can completely block all of them by just blocking third-party javascript. What are we going to do once all of this tracking code starts getting served from the first party domain instead? Or even served inside the same source files as site code?
I imagine we will start seeing a new class of privacy extensions that behave more like anti-virus. Checking for known hashes of tracking scripts, monitoring for certain patterns of behaviour during execution.
Personally, I haven't seen a desire in companies to skirt GDPR. Rather companies just want to be compliant and not have to worry about data breaches or reputational damage from their marketing tools. This example with Backblaze is exactly what companies are trying to avoid.
0: https://developer.mozilla.org/en-US/docs/Web/Security/Subres...
[0] https://addons.mozilla.org/en-US/firefox/addon/facebook-cont...
Don't give data to Facebook lmao.
Corpweb should be as static as possible, except for whatever third-party JS the marketing professionals think is necessary. It's their job, they know what's best.
Your app should have zero third-party JS except for technical analytics (New Relic, Datadog, whatever).
(This distinction can be fuzzier for free services, and for consumer stuff with non-sensitive data; Backblaze is neither of those.)
Also, what would you use to monitor user behavior to improve your product if third party tracking is frowned upon?
If you turn on demographic information in GA that explicitly also allows the data from your instance to be shared with Google’s advertising products.
> The proposed solution you mention would include plenty of folks who don’t convert.
The one bit of data I currently have going from the app to corpweb analytics is precisely this - associating a conversion with a website user. That conversion info is sent with hand-coded triggers, the relevant third-party libraries are self-hosted (a good thing generally), and code doesn't call out to it when the user is outside of the billing/subscription flow.
> Also, what would you use to monitor user behavior to improve your product if third party tracking is frowned upon?
I'm open to A/B-testing stuff - the important part in my mind is less specifically about third-party tracking, and more about ensuring that the product/engineering team makes decisions about the product. Pulling third-party code of any kind, but especially tracking code, is a process full of footguns that should be under the control of people who know what they're doing and are empowered to say "no".
I would still be very cautious about giving analytics code full access to all user activity.
* user testing, user interviews
* (there's probably a fancy word for this) generate usage metrics from production data. If the product is e.g. a To Do app, you could measure "engagement" by counting how many To Do items each user has created
* first-party tracking – analyze access logs, self-hosted analytics
That made me LOL. You'd like to think they know best. Most of the time, they do things because other people do things, but don't truly understand what the true ramifications are of their requests.
Marketing: "What's the big deal? All you have to do is add the 2 or 3 lines of JS to the site."
Devs: "Do you know exactly what that will do?"
Marketing: "It'll give us all sorts of useful metrics for free"
Devs: "Do you know if it is secure or will cause our site to become less secure due to vulns in the included JS? Will it cause the site's performance to become sluggish where we will get blamed? Do you know exactly what data is being collected, and will it affect any of our other obligations of maintaining this site?"
Marketing: "Um..., that's your job. We just want the data"
The product site itself (usually app.example.com, but Backblaze seems to use secure.backblaze.com) actually contains customer data in the browser context, is under much higher base resource loads from its core functionality, and is used repeatedly in workflows where poor performance is painful to the user.
No one gives a shit if your pricing page takes 500ms to load instead of 100ms, or if a dozen social media companies who already know where you work learn what kinds of professional products you're looking for.
They do care if a file listing takes longer, a recipe opens slower, if word frequencies in confidential data are leaked to the world.
In most paid products, the non-paid part of a site has radically different performance and security requirements from the paid part, and forcing one to be built to the requirements of the other (in either direction) is wasteful, or dangerous, or both.
To be fair, it's just as true for developers as it is for the marketing guys
> It's their job, they know what's best.
That's like sending an alcoholic down the spirits aisle.
Re: the more substantive legal points - there are off-the-shelf solutions (CMPs, for example) and easy checklists for complying in the setting of a static public-facing website. The web designers and brand managers I've worked with are more than capable of meeting clear industry standards.
It's inside a web app, where customer data is on the page and in JS scopes, where the product team is essential in safeguarding customer data.
Consent Management Platforms, things like Cookiebot?
In my experience, blocking 3rd-party HTTP requests, cookies, LocalStorage access etc. before consent is given – is easy in simple cases, but can quickly get technical and tricky.
In my experience, the people implementing these often don't understand enough about the technology to be allowed to implement this. I wonder if "easy to use" tag managers are to blame, by allowing non-experts to add JS and other includes to webpages without process or scrutiny.
Check a few big brand-name websites, and look at whether they place (third party) cookies before the CMP has even been interacted with.
I can think of some major high street labels where the consent prompt is mere theatre.
I strongly disagree. Marketing professionals often lack technical understanding and are superficial about the consequences. This mentality is how you end up with engineers working for companies doing morally questionable choices because they just want to be a cog in the system instead of being concerned about the direction and how the company do business. Separations between services is a shared illusion, if something is against your values please tell your fellow human beings.
And I'm not talking as a prospective employee who "just wants to be a cog in the machine"; I'm talking as a founder and CTO who sets company goals. I'm worried about data leaks caused by poor implementation and short-sightedness, not those caused by company policy that I disagree with. If I disagree with company policy, I change it.
Spyware is spyware, and embedding it is rude no matter who it is from.
* Both of those specific services omit lots of specific identifiers in their data collection, and require you to go out of your way to send truly sensitive stuff to their servers
* By necessity, a technical analytics service gives you lots more control over where exactly it hooks in to your code.
I'm guessing if you're logged into facebook, now FB can correlate:
1) you use a backup service
2) all the metadata from the file names
Whoops.
This is why my network runs pi-hole / diversion with all tracking blocked network wide.
Names or identifiers are the kinds of things that very often end up in file names.
Rebecca_Range_will_and_final_testament.pdf
NadyaNayme - Resume - [Company Name].pdf
divorce_process.pdf
steps-to-take-after-a-car-accident.pdf
personal_injury_michael_perez_west_coast_trial_lawyers.pdf
embarassing_porno_filename.mp4
And this is a smalls subset of filenames that could not only provide PII but also potentially embarrassing or private information that isn't identifying but would be accompanied with files that are personally identifying.
I've also seen similar schemes used by doctor offices/hospitals, where you'll see the patient name and ailment. I once had to troubleshoot a backup problem for what turned out to obviously be an OB-GYN. Imagine my horror as I saw long filenames containing patient names and common STDs scroll by in the thousands.
That’s taking it to the extreme, but it’s technically correct.
If you say “Read the law strictly enough” please do...
I'm still wondering, because Backblaze will sign a BAA (their website says so), making them a business associate. I'm not talking about some private person uploading their own documents. My concern is that given that Backblaze will sign a BAA, then some companies must be using Backblaze and potentially storing PHI data there. Yes?
Backblaze then need to follow: "§ 164.312 Technical safeguards. (e)(1) Standard: Transmission security. Implement technical security measures to guard against unauthorized access to electronic protected health information that is being transmitted over an electronic communications network."
Facebook isn't authorized to access this data, but that might be more of a problem for Backblaze, even if Facebook could be required to delete the data.
Sexual Harassment Complaints/Persons Name.doc
...> "Hi Brett! The pixels we use are primarily for audience building when we advertise on other platforms like Facebook for example." [...]
The carefully calculated cutesy "Hi Brett!" with the exclamation point is the same reason big tech companies use infantile graphics [0]: by seeming playful, they create the illusion they are a Safe Friend you can Trust.
I do agree with your point about banality of evil.
Hi there!That's pretty much the definition of a conspiracy.
I believe that your analysis has flaws. It would be quite awkward on social media to use the a more formal way of addressing people. Twitter and other platforms have a "style" of conversation and trying to fit the square peg of formal writing into the round hole of internet conversation sounds stilted. I do not understand why you think corporations would decide to do that, no matter what their intentions are.
I humbly await your response.
Yours Truly, Mr. rPlayer
P.S. I hope your Grandmother is doing well. Please send my regards.
P.P.S. Please invest in my new cloud computing blockchain biotech startup where we sell NFTs.
Please update your signature to conform with the current standards, as outlined in the last month's circular.
Best Regards, TeMPOraL
--
TeMPOraL, Internet Compliance Officer (ICO)
ACME LLC - Synergizing Creative Accounting
ACME LLC, NaN NaN, Null Islands.
The content of this message is confidential and intended for the recipient specified in message only. It is strictly forbidden to share any part of this message with any third party, without a written consent of the sender. If you received this message by mistake, please reply to this message and follow with its deletion, so that we can ensure such a mistake does not occur in the future.
Please do not print this message unless it is necessary. Every unprinted message helps the environment. Think of the trees!
So for me it's not so much that the salutation not formal enough, it's more that it's odd that it exists at all.
But then again maybe usages differ in twitterworld.
https://knowyourmeme.com/memes/subcultures/corporate-art-sty...
It's just GPG-encrypted backups.
Look. Here's what the innerText says, with the `\n`s interpreted:
Upload
Download
New Folder
Delete
Snapshot
Selected: 0 Files: 0 bytes
Name
Size
Uploaded
dbbackup_06_23_2017.sql.gpg
10.2 GB
06/23/2017 22:24
dbbackup_06_24_2017.sql.gpg
10.2 GB
06/24/2017 22:24
dbbackup_07_02_2017.sql.gpg
10.2 GB
07/02/2017 22:25
dbbackup_07_03_2017.sql.gpg
10.2 GB
07/03/2017 22:26
dbbackup_07_04_2017.sql.gpg
10.2 GB
07/04/2017 22:25
dbbackup_07_05_2017.sql.gpg
10.2 GB
07/05/2017 22:25
key.asc
3.9 KB
02/24/2016 22:01
Once again, noise. Perfect.(For the curious: I couldn't be bothered to install https://projectnaptha.com/ (probably would've worked), I just resized a terminal to the same column width and blindly retyped the text into a shell printf statement.)
dRDDrOu44Rr84vzXJv2mcr2eg83zDN43mzUQ0N4
xzrI5Ha7HJ7gK3T8XfkGqNtvc7LMQPFjSwi
5Wb1XhHeR6LxQUC8XfyX9kvvooYvrp9fnxQVvH9C
jgAzjd56DGWFjcae0gw9A1LZxJEqVHW7UmkZ
XrpNNAZalPp6D4mnpLvVcCE3uWkDQzthSQwK9
(as generated by for example `head -c 30 /dev/urandom | base64 | tr -d '+/='`)are extremely suspicious-looking. To me that screams "obsessively paranoid". I reckon that if I were in law enforcement, the fact that there is absolutely nothing I can infer from these filenames, combined with the obvious complexity associated with correctly maintaining something like this, would actually make me that much more interested in decrypting this information simply to take a look at it and rule it out.
Which is exactly why this would be a scenario in which I _would_ want to "reuse someone else's password", if you will, and I'd theoretically go digging for common archival file patterns, and use the most common I came across.
If you want exactly 30 alphanumeric characters you could do `tr -dc '[:alnum:]' < /dev/urandom | head -c 30`
If I ever did something shady, I'd absolutely make a ton of honeypots like this.
https://twitter.com/backblaze/status/1373751015594356739
Not sure if they intended to send file names and sizes to FB, but in any case this doesn't look good. I'm currently looking for alternatives.
My guess the bulk of that price difference is due to economies of scale.
Given it's all personal data that I can stand to be without for days if needed, I could probably use Glacier (or even Glacier Deep Archive) and pay less than a third of the cost (or less than a twelfth!), but the absolute dollar amount isn't high enough for me to go through the trouble of changing up my backup scripts.
Sure, if you're running a business and are constantly generating lots of new data that needs to be backed up, that gets more expensive, but I also would expect a business with that much data could easily afford to spend several orders of magnitude more than I do on backups.
So far I'm happy with it, pricing is $5/month per 250gb + 1To egress. And, most importantly, none of the AWS complexity overhead. I have enough of that at work.
Technically you do have to pre-allocate, but it works like pay-on-demand because:
<paste>
We can increase the size of your filesystem at any time, with no downtime involved - so there is no need to overbuy in anticipation of future usage.
Further, you are granted a +10% grace at all times, so you always have some room to grow.
Finally, the system automatically emails you as you get close to your limit, so there is never an unexpected filling of space.
</paste>
OVH has two interesting products here:
OVH cloud archive works with rsync, costs roughly half as much for storage as blackblaze, roughly the same for egress, but charges for ingress (at the same rate as egress).
OVH object store is s3 compatible (like blackblaze), charges roughly the same for bandwidth, 2x for storage.
https://www.ovhcloud.com/en/public-cloud/prices/#439
Digitalocean has a blob store with the same pricing on bandwidth, 4x the pricing on storage, a minimum spend of $5/month on storage, but the first tb ($10) of bandwidth free.
I would be surprised if OVH had another fire in the future.
seems this the cheapest of all
However I do think there is value in Tarsnap if security is really that important. Colin is really switched on, and I trust both the service, and if anything happened, he would deal with it quickly and professionally. If I had the kind of profile that meant I needed to protect myself against determined attackers, then tarsnap would be a no-brainer for me.
> the original version, which can survive the loss of 2 datacenters, not the "reduced redundancy" version which can only survive the loss of a single datacenter
These days, there's not just reduced redundancy, but also infrequent access, which seems better for backups…
First, we only offer SSH / TCP22 so the transit is encrypted.
Second, we have installed, and maintain on the server side tools like 'borg' and 'rclone'. So while you might just 'rsync' or 'sftp' your data to us (in which cases it would not be encrypted on our end) you can also use sophisticated, encrypted backup tools like (borg, duplicity, git-annex, rclone, restic, etc.)
If you choose a tool like that, rsync.net does not hold the encryption keys. The data appears to be random from our viewpoint.
But why they need to submit that to Facebook for paying users I don't understand. The only thing I can think of is excluding active users from advertising... But is that worth the privacy intrusion?
This doesn’t look like the right data to be sending FB for that, though.
Good job, backblaze. So many years of work building goodwill with transparency in hard disk reports ruined completely with one stroke.
[1] https://www.facebook.com/business/help/164749007013531?id=40...
Having said that, this sounds like "Hey guess what? We are gonna snoop on you, and profile the hell out of you and leak all of your sensitive data all over the place (filenames can be) all because you are a paying customer"
That's about the worst way to disrespect a paying customer. Is there a way to easily identify companies that does this, I can avoid them?
> "easily identify companies that does this"
No. And this can happen for a lot of reasons as explained so even harder to check for.
– Cross-sell/Up-sell campaigns. Build an audience based on usage patterns, and create a create a campaign for a complimentary service or higher tier (say, for example, someone clicks the button for a gated feature they don't have access to).
– Suppression lists. If you don't want your campaigns to target existing users, you can build an audience from pixel data on your authenticated pages and suppress against that.
– Lookalike audiences. After you create an audience in Facebook, you can create a "lookalike audience" from that. So even if you aren't actively doing either of the above, you'd derive value from tracking your "best" customers and using it as a seed list for a lookalike audience.
You're also not limited to using the FB Pixel for any of the above. In addition to a browser-side pixel, FB allows you to upload hashed customer information and use those for conversion tracking and audience building. Which used to be completely transparent to end users, but now you're able to see a list of companies that have uploaded your info to FB in this manner (I can't recall where it's buried in the user settings, off the top of my head).
All of that said, it's entirely likely that Backblaze wasn't intentionally sending any of this data to FB to begin with. An insidious aspect of FB's Pixel is that it automatically attaches listeners to a bunch of stuff on the page such as buttons and sends back interactions and associated metadata[1]. The flag to disable this isn't mentioned in the implementation instructions that are generated upfront, and it's actually a fairly uncommon trait for ad pixels. So a typical implementation tends to leave it on out of ignorance rather than make a deliberate determination on whether to use or disable that functionality.
[1] https://developers.facebook.com/docs/facebook-pixel/advanced...
Um. Holy crap. Is this common knowledge? I don't get shocked easily these days but... wow
In the analytics space, you traditionally had to: – Initialize a tracking object on page load – Explicitly call methods on that tracking object when you wanted to actually send a hit
This is how it works for Adobe Analytics and Google Analytics historically. Many of the newer analytics providers instantiate auto-listeners, which gave them an edge on the out-of-the-box analytics features. And Google Analytics 4 (the newest release), also does this.
So it's not unheard of for site analytics. And a quick glance at a particular provider's website can usually make it obvious if this is occurring, based on the advertised features.
Ad pixels tend to be different though. You create a conversion event within the ad platform, and you're given a snippet of code to fire when that specific event occurs, which both instantiates the tracking object and calls the tracking method with the conversion event's configuration details.
Facebook's pixel works far more like a modern analytics library than an ad pixel. It vacuums up the hit data from the site and the marketer is able to sort it out after the fact in Facebook's interface and use what they want from it. Marketers working within Facebook can see this is happening because they set up the conversions and audiences against the hit data, but that's "just the way things work" in Facebook so they think nothing of it. Marketers coming from other channels will notice how different it is, but won't realize what's actually happening nor the implications behind it. Devs would realize pretty quickly what's happening after a few minutes exposure to the FB Pixel interface and it'd trigger a red flag for them, but that's marketing's territory and all devs see are the snippets provided for implementation. So the only time most people become aware of it is if marketing has someone technical working directly within FB's interface or if the person tasked with implementation has a reason to dig into Facebook's dev documentation rather than just plop the snippet on the page like they were told.
> What happens if you resolve http://facebook.com to 127.0.0.1 via hosts-file? (Or put it into your Pi-Hole DNS Ad-Blocker, or the like.) Does the Backblaze UI still work?
> Answer: Seems to work well. You need to add the main domain and sub domains (www. in this case), btw.
Also, huge warning for Wasabi newbies so you are not surprised. This might wreck your wallet.
Your egress is capped at your total amount stored. You cannot store 5TB and have 6TB of downloads against that account. The front page is covered in 'No Charges For Egress' 'no additional charges for egress or API requests'.. etc. This is not written anywhere on the main marketing pages, only behind a small link.
Another important point is file deletion. You are required to pay for multiple billing cycles (months) of files. Be very cautious about this. Using it as a temporary bucket will incur significant costs, if you upload a file, YOU WILL PAY FOR 90 DAYS OF THAT FILE. It is almost certainly not cost effective to store files in Wasabi any less than long term. Deleting a file that you just uploaded will incur a deletion fee equivalent to 3 months of storing that file.
The same goes for AWS IA and Glacier classes and Google's Nearline and Coldline. Unless you're storing files for a long time, always factor in the minimum retention period before estimating costs. It'll prevent any nasty billing surprises -- speaking from experience unfortunately.
There really isn't any competitor to B2 on price, or convenience (e.g. them sending you a NAS device for recovery).
As a provider of a paid service, it seems like they're in a better position to take the high road than a lot of tech companies are, but they have to decide that's who they are, and be clear they mean it.
I might very well start using backblaze next year or maybe even next month[1], but that is depending on the outcome of this event.
For a comparison: one good friend once called med to apologize that he had been laughing behind my back with some friends.
Guess who I definitely trust today? The one admitted his mistake. He always was a nice bloke and I guess he will never ever do anything like that ever again.
[1]: I won't start using it this week or the next however.
Who thought that Facebook tracking should be forced on paying customers?
If this passed review - it is likely even worse.
If they refrain from engaging, then I agree that you shouldn't judge them by that.
> In the unlikely event of a data breach, as defined in the GDPR, Backblaze will without undue delay send its affected customers a notification email, and provide at its discretion, updates through other communications channels. This notification will describe the nature of the data breach, including where possible, the categories and approximate number of data subjects concerned, the categories and approximate number of personal data records concerned, the contact point where more information can be obtained, the likely consequences of the personal data breach, and the measures taken or proposed to be taken by Backblaze to address the data breach, including, where appropriate, measures to mitigate its possible adverse effects.
Because that's the only valid fix here.
Also is this “bug” in breach of GDPR?
What you have right here and right now is a public relations disaster. Trust in your brand has been damaged. It cannot be repaired by you providing minimal information. Your standardised message is akin to "Don't worry your little heads over the details - trust us, everything is fine now", and to be honest I find it a bit insulting. As far as we know, "pushed out a fix" could mean that you have hidden the tracking, so it is harder to find. Your short message is making the public relations disaster worse, not better.
These are the steps that you need to take:
1. Provide an explanation of why tracking was being performed in the first place, including an analysis of how much of that was a mistake.
2. Make an apology for breaching your customers' trust. This is a really important step, and it should be repeated in each of your press releases.
3. Provide details on the steps you have taken to fix the problem, and what that means for tracking data.
4. Make a promise that strictly limits the level of tracking that you will be allowing yourself to make in the future. Ideally we would all want that to be zero, and if you intend to do business with certain jurisdictions then you are limited to what is legal, but you must in any case be clear about what tracking you will ever do.
Honesty and transparency are the keys at this point to restoring your brand. I do not think the community will accept anything less.
Yev, if you are not in the upper management chain of Backblaze, please show mnw21cam's message to someone who is.
The problem is not that there was a little bug which caused the Facebook tracker to get a few little pieces of information it shouldn't have.
The problem is that Backblaze failed to understand how to distinguish appropriate and inappropriate uses of third-party trackers for signed-in users on a security-critical application. The Facebook pixel should never have been there at all. It shouldn't even have been considered. It should've been an absolute no-brainer that Facebook has no business being on secure pages on a critical infrastructure service for paying customers.
The fact that the pixel even showed up at all on a logged in page represents a breach of trust for customers and casts doubt on Backblaze's competence in handling security issues. This warrants a serious reply from the CEO, not a copy-pasted meaningless reassurance.
He posts a lot on HN and Reddit in Backblaze posts.
Sort of like the Streissand Effect; if BB starts posting press releases about "breaching trust", filled with apologies, the audience that will become aware of this issue will be a lot bigger than it currently is.
So from a PR persepctive, a one line low key "we pushed a fix" reply makes total sense.
Right now, the issue is being downplayed, much to the chagrin of an increasingly knowledgeable set of computer users.
I'm not sure that's true.
There are a number of barely-technical subreddits, typically centered around Plex or some flavor of bittorrent, that consist of an all-day-every-day request for "as much cloud storage as possible for the lowest possible cost and please say it's free".
These are not HN readers. It's as if they attained consciousness two minutes ago and one minute ago they decided they needed cloud storage.
This is the audience these kind of analytical tools are geared toward and I don't think this changes their "engagement" or their "convertibility" or their "lifetime value".
The real information here is that in 2021, and in conjunction with a much more sophisticated product offering (B2), Backblaze is very aggressively pursuing flat-rate, loss-leaders who can be influenced and targeted by facebook.
Draw what conclusions from that you will.
I understand that is a lower price than your own product, but it doesn't necessarily follow that they're losing money.
You seem to be implying that their covering the losses on storage by selling customer data. Is that correct? In which case, would you like to be a bit more specific and explicit?
Aren't those the kind of users you can monetize through social media "engagement" ? I don't think it's the B2 users... which class of users had data leaks ?
But just to be clear: is your claim that the unlimited storage consumer backup plan is a loss leader, and they cover their costs by selling the backed up data (or metadata) to the likes of Facebook?
It seems like you're trying to waft a very serious claim in the direction of competitor, but you're being very careful to avoid saying anything explicit or specific.
Which is to say: the provider has a vested interest in minimizing your usage and incurred costs which runs directly counter to the consumers desire to use as many resources as possible. This antagonistic relationship leads to all manner of dysfunctions and bad patterns.
When I see serious businesses "enhancing engagement" with facebook pixels, I think that perhaps that is one more side effect of that antagonistic provider/customer relationship.
HOWEVER, it turns out that the tracking code was on the B2 side of things - the people-paying-money side of things - and not on the who-will-let-me-upload-movies-forever side of things.
So my sense was wrong.
I was suggesting that this might not be as brand damaging - and trust eroding - as my parent suggested. After all, both sides of that unlimited flat rate storage relationship are pretty dysfunctional. If this was on the B2 side of things then I take it back - it's probably quite damaging.
Regardless: I stand by my disdain - and continue to warn against - flat-rate service offerings. You want your provider to happily enable you to use more of their product.
[1] https://blog.kozubik.com/john_kozubik/2009/11/flat-rate-stor...
I think their best advertising for B2 is the quarterly hard disk stats, it's a shame they made this mistake with the fb tracker; I can't store data in my garage as cheaply as they do. ~100W continuous for a box with drives at $0.30/KWh (CA PG&E) is $20/month.
I'm jumping through the hoop of s3backer to turn B2 into a vdev for a zpool with encrypted datasets; I'll have to try something similar on rsync.net with sshfs.
As another data point, Backblaze has pretty much until this weekend to provide an update that includes "we have removed Facebook and are getting a 3rd party review for other security holes in our product". After this weekend, I won't care because I'll be on a different product.
This is embarrassingly bad, as in, I'm now embarrassed for recommending using Backblaze at my company.
Crap, I just realized I got Backblaze installed at two previous companies. Thanks!
I'd still use Backblaze but that's VERY dependent on how do you handle this. Just saying "we fixed it" doesn't answer the much more fundamental question of "what is the FB tracking pixel doing in a privacy-critical page in the first place?".
Please, do a thorough post-mortem analysis and publish it. Looking at the comments here, this could mean you get or lose the business of many.
I had to fight hard as an engineer to make sure that it does not happen. We had meetings after meetings, and it took a lot of effort for me to explain the risk of data leakage. I was questioned on my "insecurity" for not "trusting" people. It was not a nice experience. I had to inform them that tracking needs to be dealt with properly, not just lazily install google tag manager because it gives marketing 'flexibility'.
Google Tag Manager is an absolute cancer on web development.
Once it's in, you'll never get rid of it.
Apparently it lets people drop in random code from a bunch of different analytics platforms, so it's pretty much guaranteed to consist entirely of the sort of stuff I have NoScript enabled to block in the first place.
Some random guy in his basement assured someone in marketing they could handle our volume? Chuck their tag in and watch their website get DDoS'ed with millions of requests per minute, which takes out our website because marketing made it fail loudly.
IT can often be "we make someone else's bad idea happen," and that's because we simply lack veto power.
Hiring people is difficult as there is lack of supply of talented people. We hire InfoSec on contract basis not full-time and they don't join such meetings due to the nature of the contract. So all the responsibility fell on me to defend our technical decisions at that point in time.
I'm working on building out the engineering culture / awareness within management now, to ensure these things do not happen, and I don't have to be questioned as to why we cannot install "google tag manager" in our front-end.
It all comes down to creating awareness, and making people understand. Fortunately for me our CEO gets it, he ended up siding with me.
To me, this is an even more concerning issue. But then, I have no idea how the finance services world works, so maybe this is more common than I think?
Outsourced CISO/InfoSec is a valid and reasonable thing for some companies.
If I was running a financial services startup, those groups would be near the front of my list in terms of internal hiring.
But again, since it's not always easy to find the right people I end up having to fill in for everything we don't have a team member to execute on.
Usually I try to reason with management first. If it can be resolved internally we would not include outside consultants, however if it gets serious beyond something we can handle internally we would ask outside consultant to come in.
Personal attacks for doing your job? If you're still there, the stock options had better be enormous!
If you are the most knowledgeable person then you get blamed for their bullshit fantasies being impossible or unwise (or illegal)
You really need to present well and be careful about arguments like “millions of websites use GTM”. I did days of research and presented that while yes using GTM on Wordpress sites that hold no sensitive data might be fine however we are a financial services and we collect customer private data. So getting everyone on the same page and presenting alternative ways of solving the problem was critical.
"I trust people just fine. I trust people working for outside companies dependent on information gathering to gather information. Google has no fiduciary duty to our clients. Same with Facebook. We DO have a fiduciary duty to our clients and it includes not doing things that may send their confidential information to third parties because SOME of the information used may be useful to our marketing department."
It means there's either nobody reviewing the privacy implications of marketing decisions, or that somebody who knows better is reviewing these decisions and decided leaking data like this is acceptable.
Both those possibilities make backblaze a non-starter for me now.
> You can't expect marketing to understand these things.
I work in marketing ops and I know how difficult it is to get marketing folks to understand or care about how the tracking they use works. There’s a small but growing number of us trying to change behaviour and awareness from the inside out but if marketers fuck up on privacy an example should be made of them.
Unless you have really strong dev leadership, the trackers will end up in your product.
I'd love to hear from folks who have successfully blocked their business team on this. What tactics did you use?
* Announces this breach to the relevant data protection authorities. They have 48 hours from learning about it to giving an initial report to the UK's ICO.
* Makes a blog post apologizing, explaining how it happened, and what they've done to prevent it happening again.
I legit don't understand why a paid storage service would put a FB pixel on their dashboard which handles user files. It's a completely foreign concept to me. This seems like a screw up but also erodes a lot of trust which is unfortunate as I had been looking at them for past 2 months actually.
I even made a post just yesterday and another couple weeks ago on how BackBlaze's inability to set a specific file name, file size limit and expiry date on the pre-signed urls is preventing some of us from switching over from S3 to Backblaze for our storage of app data needs. And surprisingly, I wasn't the only one as I got a few people responding with same concern.
https://news.ycombinator.com/item?id=26430959
Basically:
> A limitation I ran across when using B2 was that their pre-defined url generation doesn't allow you to set file-size limits nor does it allow you to set the file name in the pre-defined url. It simply gives you a pod url to upload it to. So if you are using b2 for storage for lets say image uploads from browser, some malicious user has the ability to modify the network request with whatever file name or file size they want. Next thing you know, you have a 5gb sized image uploads happening.... This pretty much prevents me from using B2 for now.
> I ran into the same limitation! IIRC, there also wasn't a way to expire a signed upload URL sooner than whatever the default was, which was hours or maybe a day. I had the exact use case you mentioned, too - image uploads bypassing my backend server. I didn't want the generation of a signed url to, say, upload a profile photo, give carte blanche to create a hidden image host when combined with the limitation that you highlighted. All sorts of bad things could come of that. I ended up just going back to S3 - costs more, but still worth it.
Since this is for a site/app which lets users upload data, I am really trying to avoid S3 due to crazy costs. I might look into DigitalOcean's offerings. Anyone have any other recommendations?
Dumb suggestion: Run it yourself? Minio is easy to use, even in multi-server mode.
It's easy to attack Backblaze.
Before attacking them for this, please make sure the company that you are working for or building doesn't do the same thing. ( I know for a fact that a lot of startups make heavy use of Audience building).
Note : I have used their service in the past and moved on to OVH due to a different issue in the past . https://www.backblaze.com/blog/b2-503-500-server-error/
yes, they support rclone https://docs.ovh.com/gb/en/storage/sync-rclone-object-storag...
> Before attacking them for this, please make sure the company that you are working for or building doesn't do the same thing.
I don't get why these two are related at all. One can do both the second and the first. By their own admission in other HN threads, Backblaze earns several million dollars a year and is proud not to have VC backing. So it doesn't seem like anybody is attacking an underdog who's struggling to change the status quo and needs to be held to lower standards.
Have been a Backblaze customer for many years, mostly because of their state of HD here on HN. Lost all confidence in Backlaze.
Alternatives? Had been using rsyncnet in startups for many years, but was more expensive (now using Backblaze for storing GBs of raw images from DSLRs)
Even if you're paying for the product, you're probably still the product.
In this particular one you are not.
The data was harvested by the Facebook pixel as part of their audience building tech for acquisition of new customers. So you are being sold by Facebook, Backblaze just happens to have been quite careless here but they are not "selling your data".
The breach of trust is BB letting third party code on their platform and especially from a particularly untrustworthy third party. That's it, it's egregious and should be taken seriously, that also means discussing it seriously.
This kind of hyperbole is counter-productive, it only makes it easier to ignore your concerns as "crazy overreacting".
And: it only happens when rendering the filenames to a browser window right? So only when I browse folder y in bucket x are my filenames for that folder shared with FB?
Backblaze, I really enjoyed being a paying customer. Until now. Bunch of dorks.
Edit: goes to show you have to encrypt EVERYTHING at rest, even file names...
This is uncalled for and devalues the rest of your comment.
It was a heat of the moment thing. Mostly I'm being angry at invasive tracking being the norm when going down the 'growth hacking' path. I must lack perspective but it saddens me that contextual advertising and focus groups apparently aren't enough.
edit: to make up by adding something actually useful to the discussion... I checked and I don't see any DNS requests being made for any facebook domain when browsing my B2 buckets. Maybe by now they got rid of the tracking pixel?
I don't see any ad hominem there. Where is it?
Or just block all third-party javascript.
Facebook code shouldn't be running anywhere aside from facebook.com
As per the latest Backblaze blog post, [1] it has been going on since On March 8, 2021 at 8:39pm Pacific time.
Here's the HN submission on that blog post that hasn't gotten much attention. [2]
[1]: https://www.backblaze.com/blog/privacy-update-third-party-tr...
[1] https://github.com/restic/restic
[2] https://help.backblaze.com/hc/en-us/articles/115002880514-Ho...
Had a look at rsync.net - too pricey for my puny sub-500GB data.
FWIW, I agree with you on your previous posts/thoughts and I'm all for pay for use and against flat rate/unlimited plans, as long as it fits in a budget.
https://news.ycombinator.com/item?id=9186428
Just google for it and I'm sure you'll find more comments along those lines.
Never include the 3rd party marketing scripts on any pages where a user is authenticated.
But of course that would deprive many companies of mountains of valuable data so it ain't ever happening, right?
---
Also, am I missing something obvious here? If Backblaze -- or anyone else really -- wants analytics, what do they need the Facebook pixel for? There are so many good analytics services out there.
I'm paying for Backblaze, now I'm wondering if that was a mistake.
If they can have FB pixels in the admin area, they basically have no security processes working. If marketing drives their tech decisions, this is not a company to trust your data to.
HIPAA isn't some random cert you have to satisfy a single big customer. It needs to be a priority and all business decisions have to be made around it. GDPR is another one.
You need strong legislation and even stronger penalties. The only way to stop pervasive tracking is to financially ruin any company that employs it.
That's extra effort, but the big ones will do it
https://www.backblaze.com/maintenance.html > We're performing site maintenance. The page you're looking for will be back soon.
I guess somebody is having a rough Sunday night.
Ad tracking pixels in your object store dashboard is just clownshoes from a security engineering standpoint, over and above the fact that it's a slimy, dickhead move for a paid service.
It is just a dumb mistake.
If they don't pay attention to stuff like this, then why should I trust them with anything at all? This isn't some minor oopsie, this is failing to deliver on their core product[1]:
Top Backblaze B2 Use Case Solutions
Backup & Archive
Store securely to the cloud incl. safeguarding data on VMs, servers, NAS, and computers
It's not foolproof, but it would stop this.
https://github.com/truevault/hipaa-compliance-developers-gui... was on HN a week ago. It seemed to jive pretty well with our internal policies at the HIPAA compliant company I work for.
Walmart CEO already did couple of years ago https://www.wsj.com/articles/wal-mart-to-vendors-get-off-ama...
Want to see it, go to flickr.com in a private window, it should pop up something about cookies. Pick "Manage Settings". It's insane.
I can has reference? Sounds genuinely interesting, and that info didn't find its way through the rock I'm apparently under.
You mean other than storing them forever?
Are there audits to check compliance?
FB used 2FA phone numbers for spying, after repeatedly promising that they will be not used in this way.
See also FB promises about WhatsApp during merger.
Is there a truly independent confirmation?
Isn't there a lawsuit or similar going through the EU courts atm about Facebook not really being GDPR compliant? eg They claim they are, but the court case is about them not being.
First, we don't know that. Second, it doesn't matter even if they are not using it in any way. Backblaze shouldn't be sending this data to FB. And lastly, even if FB isn't using it in any way today, how do we know they won't use in future?
Nobody should be having access to any data that they don't absolutely need. It doesn't matter how mundane the data is, which company the data is going to, etc etc.
Like I said, it's definitely a problem regardless.
But this is like saying it that if you misplace your cell phone, it doesn't matter whether you left it in your friend's car or a taxi. Of course it matters whether or not data is being (mis)used.
https://www.reddit.com/r/backblaze/comments/madqug/backblaze...
I looked at my network requests and I no longer see the googletagmanager requests being made (which I think calls the FB pixel code).
- provide customers with an exact timeline of when the pixel was introduced so they can work out for themselves what data has been leaked
- report themselves to relevant regulators in each jurisdiction that they operate in for leaking customers' sensitive data – under the EU GDPR there is a time limit for this.
- get on the phone to Facebook and beg them to demonstrably delete the data
Filenames are not the worst possible thing to leak, but they can be sensitive and it's not good enough to just go "oh, oops, we implemented it wrong, we've fixed it now".
It would appear that the two have absolutely nothing to do with each other.
As a backup entity your first priority should be end-user trust, you can't just squander that in the name of growth.
I appreciate that their Twitter account is looking into it, but feel like this should be a pretty quick fix. I'm hesitant to even log into my Backblaze account again until it's sorted.
Might be worth looking into, if nothing else then for public perception reasons.
The major point to me is: If marketing at Backblaze drives decisions that influence security and tech has either no say in this or is not competent enough, it's not a company for me to trust my data with.
Not a single bit of decency were given to Backblaze. Instead HN demands them to be burned to oblivion.
So they should be good continuing to use "Backblaze"?
backblaze signs BAA agreements with companies storing personally identifiable medical information, I wouldn't believe anyone who told me that this facebook data leak was turned off for those customers; they should immediately be investigated and fined if any such breaches indeed happened.
Full agreement, any interest and goodwill towards the company is now completely gone.
Disclaimer: I am not a lawyer, this is not legal advice, YMMV.
> Believe that's the Facebook pixel we use for tracking, we've forwarded to our web team for review in case that is not intended behavior.
Just... "in case" it's unintended is not promising.
Especially a tech shop like backblaze who have engineers building amazing tech in general. But then you cheap out on implementing some basic metrics for the web UI. Do you even need all the bells and whistles Facebook offers?
This will rely on a user's FB account having the same email as used for BB, which could be unlikely in the case a company is paying for it. But it should work well enough for retail targeting.
https://www.facebook.com/business/help/341425252616329?id=24...
> They claim it is hashed before use so they never see the raw data
If they can match hash data with real data, they can know more than they did before. Depending on what algorithm they use for hashing (no mention of it), they could be using a similarity hash so that will minimize changes if there are minor differences in the dataset.
Let's say I find a profile through comparing hash of email to email in Facebook's database. I can then compare additional information to see if a customer has provided incorrect information to Facebook as a user. Facebook could check if my address is similar to online shopping sites I use and if not, flag the account.
Still, main point is you only need an identifier and none of the other data Facebook has. Pixels are not required for this as noted in the original comment, they probably have enough in the account details already.
they have the original data of their own audience. So they can send it explicitly to FB (instead of FB sucking it from the page's pixel so to speak) for FB to build the lookalike audience which FB would do using the wide FB owned data.
Did their web UI for backups have the same issue?
What meta-data and file name info is sent to Backblaze when using their end-to-end encryption option for backups?
Exactly what business purpose does it serve from a customer perspective? Why are you tracking ADVERTISING on a portal with private customer data?
Also, isn't this a violation of GDPR in EU, or is there a "you opted in by default when you logged in and so it's not our fault" argument in play here?
A resounding yes it seems...
https://en.wikipedia.org/wiki/Tarsnap
Recent talk with the creator:
Let this be a lesson.
If they fix this, goodwill will be burned, but it's not a dealbreaker for using them. If they don't fix it, then yes by all means leave the service.
They prevent quite a broad range annoyances, accidents, negligence, and malicious intent.
If file names are truly sensitive you have to actually be pragmatic in how you choose to render them.