A GitHub repository was public-viewable
blog.adafruit.com
blog.adafruit.com
Salesforce employee recently published source of one of their products. I’ve reported via email since I’ve been removed from their private Hackerone programme, presumably due inactivity. Sec team just said it was “test data”. Wish I’ve made a copy since it’s gone now and bullshit like this responses just wants me published everywhere.
But they won’t, will they?
Maybe someone else will fill that void
They screwed up by allowing that. The employee screwed up by committing it to git and then pushing to a public repo.
The employee wouldn't have been able to do that if they'd enforced using fake customer data for testing/training.
This isn't to excuse Adafruit; it's to remind everyone that the hot young startup you just signed up for is probably keeping your signup information in a mysql database that everyone in the company has access to right now with a plaintext password thumb-tacked to the one office wall they have.
Unless it is a HIPAA-compliant startup. Then scrubbing PII is priority number 1.
i know they exist. i just havent, personally, had that experience, thankfully.
When a stranger has their laptop stolen on a bus, who knows what data was on it. Fingers crossed most people have FileVault turned on these days.
You should not be using real data with PII for training exersizes.
Hopefully that's what they mean by:
> We are additionally putting in place more protocols and access controls to avoid any possible future data exposure and limiting access for employee training use.
It does however come with this risk, that they let data, or Jupyter notebook output or etc. get committed.
I'm not saying what happened here is excusable -- it's not, it's ridiculously bad. But I can see a few ways how it'd happen.
To me the equal or better question is why does this person have PII in the first place? It's not necessary to do analysis. Someone should have masked or removed it before this person ever got their hands on it. There's no way a PII dataset is needed for training.
... that's also a short path to dropping a user's PII in the clear.
Or at all. This is actually a fairly common problem among certain types of research scientists who think software engineering is "peasant work".
I remember a friend of mine working at a research organization where several researchers lost months of work due to an unscheduled server reboot. Turns out when you log into ephemeral containers and pretend they are VMs, things go poof.
As horrible as the robbery was, I was screaming inside "WHAT ABOUT VERSION CONTROL?!??!"
If he doesn’t come from a programming background version can seem foreign.
But like ya know… Dropbox. I’d be more concerned with my hard drive corrupting than anything else
For many people technology really is magic.
I think they might still be deadlocked to this day.
Docker offers persistent volumes, so my 30 second solution would be that
Is github the most accessible? It is the most popular and best for discovery but a discussion around if gitlabs or bitbucket being more accessible could be had.
You make that sound like it's generically a bad thing. Having the data you used and tested with is good in many ways. Having it in the same git branch and repo as everything else is suboptimal but usually a lot better than not safekeeping it at all.
But it's a bad idea to put PII on github which is apparently what happened.
As for the individual repos, here's the procedure:
cd <repo>
rm -rf .git/ # DON'T TRY THIS AT HOME unless you know what you're doing.
git add .
git commit -m "Brand new repo"
Do not do the above unless you completely understand the ramifications. But those ramifications might be precisely the ones you want for legal compliance.There are various tools to properly filter history, e.g. `git-filter-repo`, but the short answer is as the grandparent commenter said, things get hairy, you need to rewrite history and coordinate. It's not a situation you should hope to get yourself into...
Also, since that's a web store and the data comes from customer orders they would have no right to deletion as that information becomes protected for other legal requirements.
Addresses are sensitive information and whatever was happening sounds like multi-purpose-consent-necessary data processing and it was years old.
False. Transaction data must be kept for legal reasons and deletion requests do not apply to it.
But logs don't count as the transactional data that needs to be kept for legal reasons.
> We evaluated the risk and consulted with our privacy lawyers and legal experts, and took the approach that we thought appropriately mitigated any issues while being open and transparent and did not believe emailing directly was helpful in this case.
Seems like a pretty weak justification. "Our lawyers said we don't have to notify users."
I stopped trusting adafruit when they sent me broken boards and never replaced them. Their software side is very startupy, they love to break things, especially if it was working perfectly fine before. I have a USB to uart/SPI/i2c adafruit device that won't operate in any of those modes without changing windows drivers between them, and ontop of it requires local side circuit python and it used to just be C.
Circuit python was a huge mistake on their part. They didn't have the people to do C correctly (look at their nrf libraries), much less C and circuit python versions. Which ends up being a compatability nightmare.
So I just buy a 5 pack of esp32's or other dev boards off Amazon and be done with it. If they're broke, I get them for free. Sad to see that Amazon breaks less consumer laws than Adafruit in this case.
I don't know that I've ever seen rebranded Aliexpress merch on Adafruit. Boards and such are primarily their own design. Proof being ... it's all open source and the design files are on Github. There's occasionally other vendor's products there, like Pimoroni. Again, not rebranded Aliexpress. Besides boards, there's tools and such, but those have all been on brand from what I've seen. Adafruit is how I found out about the fantastic Japanese ENGINEER brand. They generally put a lot of effort into combing through suppliers and finding good quality stuff.
In general, they have reliable supply chains and manufacture everything in New York City. I have never heard of them selling counterfeit FTDI chips. I have two of their FT232H modules purchased directly through their store and they're legit.
They refunded me a $100 purchase, no return required, when I was sent defective product. I've always found their customer service to be friendly.
> Their software side is very startupy, they love to break things, especially if it was working perfectly fine before.
This is a very odd sentence to read in the context of embedded development. Adafruit is a hardware company with makers as a target market. I find Adafruit's software to be a good bit better than most vendor software I've worked with. And when there are bugs, most of the time the bug ends up being someone else's fault. cough Espressif cough.
That's what I think a lot of business owners and managemwnt don't quite understand about lawyers. If you listened to everything your lawyers are telling you, then you would never do anything interesting (i.e. anything risky). You don't necessarily have to listen to everything a lawyer says because their incentive is to limit your risk, not necessarily to ensure your long term success, which are different things. If it means doing wrong by your customers then lawyers may tell you to do it if it means a micron less liability.
It's bad because smaller companies can't compete. And many of these smaller companies do more for the community than Amazon does. Good luck calling Amazon for help with picking a sensor.
It's good because if you return stuff you'll save money with Amazon. Like the times I buy things outside of Amazon I'm amazed to see I can't return it, of the penalty for doing so is like 40% of the item cost.
I strongly considered purchasing my Raspberry pis from them, but I was put off by high shipping cost.
I mostly buy from DigiKey. They have a very wide selection of professional stuff, and they also resell adadruit and other similar hobby stuff (I think adadruit is a great company that makes great products personally).
You're not going to find a raspberry pi anywhere except from a scalper (~2x MSRP) for at least a few months, though. See https://rpilocator.com/?country=US .
I started off with rpi, but quickly realized I preferred not having a whole OS to manage. All my projects are now using adadruit boards, mostly feather nrf52840 ($20, Bluetooth low energy, great rust or circuitpython support).
Taking this advice, and weighting this against other factors (goodwill with customers, for example) is what a good CEO is able to navigate.
What is this criticism supposed to mean?
Putting PII into Github (even if the repo is private) is catastrophic. You just made Github into a third-party data processor by accident. Good luck explaining that from a CCPA/GDPR perspective.
If a diff is not useful, don't store it in git. Doing a code review on PII is not useful.
You also shouldn't add packages in a git repo if they can be downloaded immutably in CI/locally. It adds a lot of space. Every commit is a snapshot + new code. Commits are not diffs.
Similarly, I keep small pdf manuals in git. I add them in their own commit, which doesn't have a useful diff. In exchange, they're always there and I don't need to spin up some special one-off system.
I was replying to a comment asking why PII data shouldn't be stored in git vs a database. Just tried to give a general rundown of how to think about the tools.
Totally agreed that the data should have been synthetic/scrubbed and that a small sample set (large enough to validate) would be ok.
Sometimes a PDF or a few pictures are necessary (or just much simpler). I get that. I have seen some repos with an unreasonable amount of PDFs/Docs/pictures/etc.. and that's when a script that copies them into a gitignored directory from someplace (Dropbox/S3/etc..) is a better fit.
Shame on Adafruit. Hope they get sued out of existence.
They made a mistake. The response to the mistake is NOT GOOD. However, destroying a company that also does actually good things is not the answer here, either.
How should they fix this? I think they should start be writing another blog post apologizing for how bad the first one was, for starters. Emails to affected people next... But, let's be honest. How much better is _that_? Can't unleak the data. You can only be offered subscriptions to data protection services _sooo many times_. There's not really a monetary value as damages here...
"Adafruit team began the forensic process" - Checking audit logs is hardly a forensic process and the rest of the speel about "privacy lawyers and legal experts" feels like a poor effort to regain some trust.
Taking an extract including PII (names, usernames, IP addresses, or email addresses for example) from your customer prod db and dumping it in a csv in a private repo on GH is most likely a violation unless you have prior explicit consent for that very purpose and use.
This is true even if it's "pseudo-anonmized" in a way that the original PII can be deduced by combination with other datasets.
Finally, I wouldn't be surprised if many of those companies are operating illegally. Drinking and driving doesn't become legal just because a threshold of people start doing it. There is a lot of ignorance (willful or not) among EU businesses even today.
AIUI, it could swing either way depending on several factors, such as: the format and usage of the data; how and where the data is transferred, processed, stored and exposed; what access and role the colo provider has (are you purely renting a dedicated server in a DC with FDE that you unlock remotely with an HSM or is the data processed by one of their managed services?); how the consent you already acquired was formulated.
If the colo provider has an outsourced support engineer in Asia looking at logs/coredump or temporarily transferring a backup where the PII appears, that would constitute a transfer, for example, and full compliance needs to be guaranteed throughout.
It's years since I considered myself to have a clear and deep understanding of it and it's gotten a bit fuzzy since, so someone else might chime in with a more clear answer.
You got downvoted because storing encrypted secrets is fine.
Storing where?
Only until/unless the encryption is broken. I wouldn’t store long-lived secrets even encrypted in a public place.
"Update March 7, 2022: We appreciate the feedback from the community and our customers, and will be emailing users as part of this disclosure. We apologize for not doing that at the same time as the post/disclosure on Friday, March 4, 2022."
Why did this happen? Imagine it's this data analysts' first year out of school, something like a crappy statistics program that only teaches SAS or base-R with no VCS. This young padawan needs to practice analysis in preparation for a work project. They grab some data and stash it in a folder where Git is tracking. They do their analysis practice, it's Friday afternoon, they get lazy, they don't look at what they're committing, and they click buttons in the GUI without much thought.
It's incompetence and laziness, not malice. These tools that allow us to share widely, not just GH but also social media broadly: these tools have great power that should be used with greater training, responsibility, and care.
Sure it's easy to call it lazy for a data set to be in some local directory and accidentally get committed. Happens to all of us. The bigger problem is "why is that data sitting on your file system in a directory, when it should be in some data base, preferably not locally."
> these tools have great power that should be used with greater training, responsibility, and care.
This screams more and more that the tools are bad. Git is famously hard to use and even harder for non-plaintext data. Databases are annoying to initialize and get access to without a developer who's done it before. The tools suck, they can be better, and require less training. It's not wrong to be lazy - it's wrong to make the lazy path dangerous.
Um, yeah, it is - by most reasonable definitions of the word "database".
No doubt there are a few unusually-narrow definitions of "database" out there that would exclude Git, but I'm pretty certain they're in the minority.
This data wasn't "secrets" though. PD treated this loosely would still be an issue even if it was encrypted.
Much better to treat data, including secrets, as radioactive, and thus follow least priv + defense in depth. If it's never there, and better, they never got it, there's nothing to worry about
Ex: Before PII hits logs, encrypt it and don't give data scientists the key.
Ex: orgs we work with who version notebooks won't allow notebook output to be saved, and people know (+ software) to reject baked creds. Likewise, auth isn't baked generally, instead SSO.
If it's never there, for every stage, so much easier.
E.g. k8s sealed secrets https://github.com/bitnami-labs/sealed-secrets
E.g. Salt's GPG filters https://docs.saltproject.io/en/latest/ref/renderers/all/salt...
E.g. Ansible Vault https://docs.ansible.com/ansible/latest/user_guide/vault.htm...
E.g. Puppet hiera-yaml https://puppet.com/docs/puppet/6/securing-sensitive-data.htm...
PD/PII is a completely separate issue. First because even encrypting doesn't remove legal obligations concerning processing, and second because your DS/BI teams probably need access to the unencrypted data to like, do their actual jobs. You need completely orthogonal forms of access control for that (like SSO as you alude to).
I'd be curious if/how any data science teams (google/facebook/netflix/...) bake user credential & API secrets into data science notebooks. I've never seen it, but I don't get to see everyone's notebook environments. I have seen 1-2 projects attempting to do DB auth plugins/libs for jupyter notebooks, but not high-grade production ones. Instead, it gets baked into the deeper env (think Tableau, Databricks, ...), vs part of the ipynb.
Notebook security feels like 1990s/2000s browser security and the literal decades of unsafe web apis. DLP and all that is at the forefront in risks, yet most tools just do system auth and maybe a few special connector auth. The real threat model is outside of their system, so it's no surprise analysts & their data orgs fail in practice :(
Sealed secrets can be used in various k8s-aware notebooks.