In the meantime I'll be calling "private" repos "unlisted", seems more appropriate
In the meantime I'll be calling "private" repos "unlisted", seems more appropriate
The same for “deleted” repos.
Which is, indeed, what every modern database does.
I believe Kafka make deletion difficult, since it's an append-only log, but Kafka doesn't work well with laws that require deletion of data, so I don't believe it's a popular choice any longer (I.E. isn't modern).
^ (more likely they’ll just update the table to set a deleted flag)
To me, the idea that the deletion takes time to complete doesn't negate the idea that the data will be gone once the process completes.
WAL archive and backups are external systems. You could argue that nothing supports deletion because an external backup could exist, but that's not a useful conversation.
Can you point to an example of a modern database that "supports deletion" but keeps the data around forever? Maybe I've just used different tools than you. Knowing modern data retention concerns I'd be surprised if such a thing existed.
Pointing to which database you are talking about should clear this up quickly.
I don't think it's reasonable to talk about backups here. A backup is external to the database so it inherently cannot delete it. Similar to how a piece of paper cannot destroy a photograph of the paper, but burning the paper destroys it.
And sure, the DELETE FROM statement in postgres - or any other standards compliment sql db I know.
For Postgres you've got to consider vacuum. Auto vacuum is enabled by default. Deleted rows are removed unless you go out of your way to make it do something different.
- What was your "definition of delete" again?
- You mentioned some of the convenient technical defaults your frameworks and tools provide out-of-the-box, can you think of ways to improve the situation?
(You might re-run delete requests after restoring a backup; transaction should resolve in a timely fashion, failed deletes can be communicated to the user quickly etc.)
However, besides the technical aspect you talked about the "absolute best you could expect when asking for a delete in the UI^".
I think this where I, other posters in the thread, most people, and probably the GDPR and other legislature, would disagree. We expect significantly more effort to clean up deleted data.
This includes, for example, the ability to delete datasets from backups, as well as a general accountability of how often and where all the data is stored and if, and when a deletion process is complete.
Nope. GDPR allows deleted data to be retained in backups so long as there is an expiration process in place. Doesn’t matter how long it is. But certainly nobody has a right to forcing a company to pull all of their backups from cold storage and trove through them all any time any deletion request takes place. That’d be the quickest path to Distributed Denial of Bank Account Funds imaginable. Even the GDPR isn’t that bone-headed.
But yes, it is part of the law that the provider should tell you that your data isn’t actually being erased and instead it will be kept around until they get around to erasing everything as part of their standard timelines. But that knowledge doesn’t do anyone much good.
> CNIL confirmed that you’ll have one month to answer to a removal request, and that you don’t need to delete a backup set in order to remove an individual from it.
https://blog.quantum.com/2018/01/26/backup-administrators-th...
However I would like to avoid the impression that with the description of the technical status quo the topic is settled. To do so I would go back to my previous point: Imagine some truly illegal pictures are in that cold storage backup, and one day you might have to restore that data. (Since aparently the user's wish to delete data is not quite as respected as certain other hard legal requirements regarding content)
What solutions to mitigate the situation could a company, or backup tool/web framework etc. reasonably come up with? Maybe check the restored data against a list of hashes/IDs of to-be-deleted-data?
Or when its encryption key is overwritten.
But it probably is a good idea to stop returning deleted data from web APIs.
*obviously doesn't apply to soft deletes.
What GitHub is doing here is neither temporary nor inaccessible.
1. Create a private repo R
2. Create a private fork F of R
3. Push commits to the fork F
4. Make R public
The commits pushed to F prior to R being made public will become de facto public, even though F has always been a private fork. The post makes clear that commits pushed to F after R is made public are placed into a separate, private fork network.So basically, if you ever intend to open source anything, never do it to an existing private repo. Always start a from-scratch repo to be the root of your new public project.
However, if they "don't care" about such an issue, how can I trust them to care about other stuff?
- gives you control over gc-ing their hosted remote?
- does not to your knowledge have a third-party public reflog or an events API or brute-forceable short hashes?
if so, especially the second of those seems a fragile assumption, because this is "just" the way git works (I'm not saying the consequences aren't easy to mentally gloss over). Even if gitlab lacks those things curently (but I think for example it does support short hashes), it's easy to imagine them showing up somehow retroactively.
If you're just agreeing with the grandparent post that github's naming ("private") is misleading or that the fork feature encourages this mistake: agreed.
Curious to know if any git hosting service does support gc-ing under user control.
Or you know, self-host, preferrably on-prem.
Basic git hosting only needs a sshd running on the server. If you want collaborative features with a web UI then there are solutions for that available too.
Clone and mount, unmount and commit
i read headlines like the above with the implied "not just to the employees there anymore"
Order does not indicate any preference.
That might be a bit too strict. I'd still expect my private repos (no forks involved) to be private, unless we discover another footnote in GH's docs in a few years ¯\_(ツ)_/¯
But I'll forget about using forks except for publicly contributing to public repos.
> Users should never be expected to know these gotchas for a feature called "private".
Yes, the principle of least astonishment[0] should apply to security as well.
[0] https://en.wikipedia.org/wiki/Principle_of_least_astonishmen...
That doesn't fork, but does what you would expect, a fully private repo.
git is a "distributed" version control software afterall. It means a peer can't control everything.
And anyone in the world can pull what was pushed to a public git repo before you delete it. You should always assume that has happened.
"Anyone can access deleted and private repository data on GitHub"
Not everything needs to be designed for idiots and lazy people, it's ok for some tools and services, especially those aimed at technical people and engineers to require reading to use properly and to be surprising or unintuitive at first glance.
I agree generally that interfaces have been dumbing down too far, but "private is actually not private and it's on you for not knowing that, idiot B)" is a weird place to be planting that flag.