Claim: Private GitHub repos included in AI dataset
post.lurk.org
post.lurk.org
"I found two of my old Github repos in there. Both were deleted last year and both were private."
The Stack was constructed a while ago, so "deleted last year" wouldn't have an impact if it was constructed before then.
"Both were private" is the thing that needs to be unpacked here. Were these genuinely private repositories that had never been made public on GitHub?
https://huggingface.co/datasets/bigcode/the-stack-v2 talks about where the Stack comes from: "This dataset is derived from the Software Heritage archive, the largest public archive of software source code and accompanying development history"
You can search that here: https://archive.softwareheritage.org/browse/search/?q=simonw... - it would be interesting to know if the OP's "private" repos are included in that collection.
I've got a couple that are intended to be GPL, but you wouldn't know unless you go to the GitHub issues to find the issue I raised about the licence file not being in the repo. They are included.
The poster responded to that:
https://hachyderm.io/@emenel@post.lurk.org/11212861313743638...
> I’m pretty sure the repos of mine were private. If it were only me i could be misremembering something, but i have heard from a number of people that they also found repos in this dataset that were private.
Still, the post continues to be useful for those who want to opt-out regardless.
The other option is that the scraper got lucky with a tiny glitch (whatsoever).
On the one hand I bet github does everything to keep stuff secure, on the other hand I can't believe there wasn't a single glitch in the last few years.
And if there was a glitch then regular automated scrapers are a pretty likely siphon.
Heck I should check if any of my private repository show up.....
That's not "a tiny glitch". What you described is "GitHub may show your private repo to strangers randomly" which is a even more serious issue than they appearing in an archive some time later.
A lot of companies pay GitHub a LOT of money to securely host their code. A breach like this is a way bigger story than just another "AI is training on your data!" thing.
I continue to doubt that private repos being exposed like this actually happened here.
It used to be public by default, and enough people got confused by this that AWS & Github used to specifically scan repo's for accidentally public AWS credentials.
> I work for github. This post was shared in our slack, on a non work related channel. We don't think it's us.
> https://huggingface.co/datasets/bigcode/the-stack-v2
> They say they get data from SoftwareHeritage, a website that archives repos from github. If your repos were ever open, they might have been archived there even after you deleted from github.
> That's my best guess.
The definition of open source depends on a license existing in a repo. Without a license it’s not legal to copy and distribute.
Public vs Private repo is a platforms issue not the code maintainers.
If a public repo does not have a license, it does not mean it free to copy and distribute.
If a private repo has an open source license like MIT, then the crawler has a right to copy and distribute that repo. Regardless if it has authorization to access the repo or not.
e.x. if the repo has some sort of /used-licenses/ folder where the licenses for packages and the like are included, it could make a bad decision.
Yes it is. Due to both the terms you agree when you use GitHub and the general Implied License that covers everything public on the internet.
>Field had actual knowledge of the Googlebot. He also was aware of the ways to prevent Google from either listing his site at all or listing it but not providing a link to the cached version. Instead of opting out, however, he chose to allow Google to both index and provide a link to the cached version.
For the AI dataset, (A) did the person know their work was being collected by this group and for this purpose, and (B) did they know of a way to prevent that collection?
Is this true? When you post anything publicly, from sticking a poster on the street to making artwork like banksy, isn’t the default set to “it’s legal to copy, unless explicitly stated otherwise”?
There is also the practical issue that a lot of content is posted publicly without consent of the copyright owner. It's simply not true that just because someone else committed a copyright violation first, you can commit further violations without impunity based on that first violation.
Whether or not it is free to copy and distribute, it should be free to copy and distribute. (My opinion is that copyright is no good; if the file is public then you should be allowed to copy and distribute it.)
> If a private repo has an open source license like MIT, then the crawler has a right to copy and distribute that repo. Regardless if it has authorization to access the repo or not.
I should not think so. The license would only apply if you have a copy of it anyways. If you are not authorized to access it because it is private, then you would have to get a copy from somewhere else, and if nobody else is providing a copy, that shouldn't give you the right to unauthorized access. However, if it has been done, then it is done, so now there is a copy, and the license (if it is a license that allows copying it in this way) would authorize you to continue to use and distribute the copy that you have.
If repo is forked and the license is deleted the source code would need to be hashed to verify its the exact version of an open source repo. Mainly they don’t want copyleft or “malcious” license infecting their IP
If the hashes don’t match then it’s not technically the same code, so a company can’t safely use it without a license.
To quote the email they sent me:
"The mission of Software Heritage is to collect, preserve and share all the publicly available source code: https://www.softwareheritage.org"
So they're telling me their intent is to reproduce (share) my work, just because it was publicly available.
Upon my reply they did offer to cancel the request, but also told me they are facilitating the storage of my code for users private theft
" "Add forge now" requests are submitted by Software Heritage users who think that a forge is worth being archived.
After a careful examination of your arguments, we acknowledge that your forges may not be archived, so we won't process their ingestion, and close this add forge now request. However, we cannot prevent a publicly available from being archived by anyone using our "Save code now" feature
Surprisingly, the people generating real, novel, IP like to buy groceries now and then.
My own anecdotal experience is that it has no relation and I know many people who have violated the copyright on the IP they themselves produced.
I've had people "steal" IP by sending me a PDF of a paper they wrote whose copyright resides with their university. Not the typical image of a thief.
It's a shame because the GP had some valuable information to share about the emails software heritage was sending, but likely got downvoted because of this.
but frankly, it's not the first time. Github (after purchase by microsoft) did the same thing to my code. They reproduced it and put it on ice in some place in norway. That was the guise right? I mean I am sure they did, but you can bet your ass they're also using my code to train their AI.
Most of it was not licensed. It was publicly available as a portfolio so people would see I write serious code and would hire me.
So then I took down my entire github after being informed I was enrolled, couldn't opt out, etc, and then some OTHER foundation is now scanning my self hosted git repos? That made my blood BOIL!
Edit: I asked this before the parent commenter included the contents of the email. Thanks parent commenter!
The mission of Software Heritage is to collect, preserve and share all the publicly available source code: https://www.softwareheritage.org
We have received a request to add the forge hosted at the URL below to the list of software origins that are archived, and it is our understanding that you are or know the contact person for this forge.
In order to archive the forge contents, we will have to periodically pull the public repositories it contains and clone them into the Software Heritage archive. FAQs for our processes are available:
https://docs.softwareheritage.org/user/faq/#add-forge-now https://www.softwareheritage.org/faq/
Please let us know if there are any issues to consider before we launch the archival of the public repositories hosted on your infrastructure. Please use "Reply all" to ensure our system will process your answer properly.
In the absence of an answer to this message, we will start to archive your forge in the coming weeks. Only the publicly accessible repositories will be archived.
Thank you in advance for your help.
Kind regards, The Software Heritage team
-- bye, pabs
This is not true. They just had to remove about 500 public repos to comply with my copyright.
I was expecting to see at least one repo that shouldn't be there depending on when the dataset was put together. In 2015 I changed a repo from public to private, which I think might suggest that the dataset was built after 2015 since my now private repo isn't in the dataset?
If the repo was public even for a single commit, that was likely cloned and replicated elsewhere.
So it seems probable this is a case of repo owners misremembering.
> If you had code on GitHub at any point it looks like it might be included in a large dataset called “The Stack” — If you want your code removed from this massive “ai” training data go here:
> https://huggingface.co/spaces/bigcode/in-the-stack
> I found two of my old Github repos in there. Both were deleted last year and both were private. This is a serious breach of trust by Github and @huggingface.
> Remove all your code from Github.
> CONSENT IS NOT OPT-OUT.
There's no archived version on archive.org at least.
It is important to note that GitHub archive is not 100% accurate and there is over 319 missing hours.
Interestingly my public code with thousands of stars isn't in "The Stack".
Presumably github repo privacy state has an audit trail. This would allow GH to prove / disprove claims on any given repo easily. I hope a rep steps in to do so.
- https://github.com/emenel/dust
- https://github.com/emenel/portfolio
(based on https://archive.softwareheritage.org/browse/search/?q=emenel...)
Care to check them on gharchive? I bet they used to be public.
The data for The Stack’s dataset is sourced from the Software Heritage Archive, so checking that is redundant. We need different sources.
Now, this is only about it being a GitHub breach. Whether unlicensed (emenel/portfolio) or GPL (emenel/dust) code should be allowed in such datasets is a different matter.
[1] https://archive.softwareheritage.org/browse/origin/directory...
[2] https://archive.softwareheritage.org/browse/origin/directory...
[3] https://archive.softwareheritage.org/browse/origin/visits/?o...
The “different source” is supposed to be ghactions.
I'd really like to know if they were ever public at any point because that might explain it.
None of my repos that have always been private are included (apparently).
That's not to say I'm not concerned by this...
I have 66 repositories picked up and put in The Stack. I spot-checked the first 10. All 10 are Public on GitHub. 8 of the 10 do not have a license of any type, meaning they are covered by copyright, at least in the US, unless GitHub has some terms extending the license of public projects to 3rd parties.
One of mine is Private and was an extension I sold for a short time. I can't say if I ever made it public or not.
You can access that for your repos here: https://github.com/settings/security-log
Then search for repo:simonw/datasette or similar.
But... it looks like the audit log only goes back 6 months, so sadly it's not useful for reviewing this particular situation which involves repos that could have been 5 or more years old.
The ClickHouse copy of the GitHub Archive is useful for reviewing things and goes back a lot further. Try it here:
https://play.clickhouse.com/play?user=play
You can run this query to see relevant events for a specific username:
with public_events as (
select
created_at as timestamp,
'Private repo made public' as action,
repo_name
from github_events
where actor_login = 'simonw'
and event_type in ('PublicEvent')
),
most_recent_public_push as (
select
max(created_at) as timestamp,
'Most recent public push' as action,
repo_name
from github_events
where event_type = 'PushEvent'
and actor_login = 'simonw'
group by repo_name
),
combined as (
select * from public_events
union all select * from most_recent_public_push
)
select * from combined order by timestamp
The PublicEvent one is "When a private repository is made public" according to https://docs.github.com/en/rest/using-the-rest-api/github-ev...I just built a tool for running this query without having to type in the SQL: https://observablehq.com/@simonw/github-public-repo-history
Explained in this TIL: https://til.simonwillison.net/clickhouse/github-public-histo...
> I work for github. This post was shared in our slack, on a non work related channel. We don't think it's us.
> They say they get data from SoftwareHeritage, a website that archives repos from github. If your repos were ever open, they might have been archived there even after you deleted from github.
> That's my best guess.
https://hachyderm.io/@emenel@post.lurk.org/11212861313743638...
> thanks for the additional info. I’m pretty sure the repos of mine were private. If it were only me i could be misremembering something, but i have heard from a number of people that they also found repos in this dataset that were private. So i’m not sure what to think or how to explain it.
Using another platform/self-host would introduce a friction.
I am not opposed to issues per se. The problem is dealing with the amount, it's exhausting. It's just too easy to make issue or comment on GitHub because of network effect. This sums it up pretty well: https://nolanlawson.com/2017/03/05/what-it-feels-like-to-be-... (my workload is of course far smaller, but still it takes time and saps energy)
Yes, that issue has been opened for 4 years and I don't consider it important and neither did any of "me too" comments to do anything about it. At best I sometimes get drive by PR, which takes far more time to deal with than if I just did it myself.
My comment did not. The comment I replied to did. It said:
> especially with the number of actual FOSS alternatives available right now
I find GitLab quite lacking to GitHub and don't particularly see it as a compelling alternative to GitHub. Plus, there is no actual benefit to it being FOSS.
You are welcome to find GitLab lacking (there's currently at least 66500 people who agree with you), and "compelling" is up to you, but to say it's not a full featured competitor to GitHub is disingenuous
IANAL, but it seems like inclusion in the data set and subsequent distribution without the copyright notice would be a violation.
It’s the similar legal clauses used for decades on social and video hosting platforms.
---
You grant us and our legal successors the right to store, archive, parse, and display Your Content, and make incidental copies, as necessary to provide the Service, including improving the Service over time. This license includes the right to do things like copy it to our database and make backups; show it to you and other users; parse it into a search index or otherwise analyze it on our servers; share it with other users; and perform it, in case Your Content is something like music or video.
This license does not grant GitHub the right to sell Your Content. It also does not grant GitHub the right to otherwise distribute or use Your Content outside of our provision of the Service, except that as part of the right to archive Your Content, GitHub may permit our partners to store and archive Your Content in public repositories in connection with the GitHub Arctic Code Vault and GitHub Archive Program.
--- Any User-Generated Content you post publicly, including issues, comments, and contributions to other Users' repositories, may be viewed by others. By setting your repositories to be viewed publicly, you agree to allow others to view and "fork" your repositories (this means that others may make their own copies of Content from your repositories in repositories they control).
If you set your pages and repositories to be viewed publicly, you grant each User of GitHub a nonexclusive, worldwide license to use, display, and perform Your Content through the GitHub Service and to reproduce Your Content solely on GitHub as permitted through GitHub's functionality (for example, through forking).
---They don't even include training for copilot (though a dodgy lawyer will likely try to include that as "part of providing the service). 3rd parties only get a license to fork your repo, seemingly not even a license to do anything with that repo. (And hot take: Github should just let people disable the fork button already.)
I'd bet that this is enough to detect cases like GPL code, but I also bet that if this analysis fails instead of falling back to "unknown license, assume proprietary, don't copy" they fall back to "free lunch!". Because reasons.
Even though even permissive licenses like MIT and BSD require attribution and preservation of copyright notices. Maybe their AI just can't "reliably detect" licenses.
I may be misunderstanding your suggestion. If so, I'm curious to learn what I missed.
(edit: I think "v2" is already released, but you get my point)
Is there any other corroboration or proof?
If on the other hand, it's tracked to the latest commit, that is a different scenario.
If they are using private data in AI dataset (or other uses) then that is a serious issue; they are copying data which is meant to be private. (Public files are public and should be made copies that others can use too, though.)
Step 1) "Software Heritage" crawls the web (not only Github; they definitely crawl ANY gitlab instances they find online among practically everything else). They "store" ANY type of source code, irregardless of the license of that code. They admit as much in their own FAQ ( https://www.softwareheritage.org/faq/#24_What_is_the_policy_... ) :
> Software Heritage archives everything that is publicly available, without preliminary tests or checks [for LICENSE file or others]. You are responsible for checking whether the source code you find in the archive can be reused, and under which terms.
Step 2) "Hugging Face" uses the Software Heritage dataset to build some AI training dataset ("The Stack"), again, completely ignoring licensing. You apparently have to manually opt-out if you don't want your COPYRIGHTED source code to be included there. But as far as I can see, opting out is only considered AT ALL for Github repositories via https://huggingface.co/datasets/bigcode/the-stack-v2 . If you have your source code published outside Github, then your code appears to be used, period.
> The Stack v2 is a collection of source code from repositories with various licenses. Any use of all or part of the code gathered in The Stack v2 must abide by the terms of the original licenses [...]
Step 3) Eventually and inevitably someone releases some model trained on this dataset, ignoring licensing again.
Step 4) Someone uses code generated from such chatbot, unknowingly violating everything software license known to man.
Step 5) ???
Step 6) Profit!
Where do these "Software Heritage" guys mention which user-agent their crawler is using so that I can permanently ban them from my websites?
At least archive.org does that and they have a much nicer way to request exclusion from their archive.
To me, a learning path that's acceptable for humans is acceptable for AIs. All of the rationales for treating AI/human learning distinctly feel flimsy and arbitrary. "The difference is learning scale." So what? "LLMs just parrot derivatives of their training data." No, they don't. That's demonstrably false, and theoretically absurd, for all the common cited reasons. "AIs obviously aren't people. So their learning restrictions must be different." Why? "People are financially benefiting from the AI's knowledge!" Yep, that's generally how employment works. Etc, etc.
My take could be totally wrong. I get that. It'll be interesting to see where the courts fall on the issue. But personally, the critics' arguments feel profoundly crummy and counterproductive; applying their logic to pre-LLM systems usually results in horrendous alternative realities. But that's just, like, my opinion, man.
Anyway, this is strictly one level up above that debate, as here they are scraping source code which is NOT free, and whose licenses could explicitly forbid copy for use in training datasets, or by the military, or even by people whose eye color I don't like.
For example, example code for proprietary tools which usually allows you to only copy it strictly for purposes of extending the proprietary tool.
I eventually had to specifically block their IP addresses from accessing my Gitea and threaten legal action before they took the offending repos down.
I still have them blocked.
Consent is not Opt-out == Yes means Yes
I have "private" repos listed in the dataset but they were all at one time public. Searching the SoftwareHeritage site I can find those once-public repos with ancient commits.
My private repos that were always private are not listed in the dataset.
In all seriousness, without strong data privacy regulations (ie: GDPR), we will continue to see this sort of stuff, as the potential monetary rewards for using this sort of data far outweigh the potential liability. Cost of doing business type stuff, rather than an existential risk for abusing public trust.
My opinion is that data should be treated like a fissile element - very dangerous to hold and store, but extremely powerful when properly employed. However, it's only dangerous if the liability of storing it is significant and real, as of today, it's not (in the US).
Now the AI craze stole our code and collective knowledge to ultimately train their models, and hopefully replace SWEs and other knowledge based fields (ie, medicine).
At least with blockchain, some people got rich. But with the AI craze the only people getting rich are the rich themselves.
But plenty more lost money or went bankrupt. The one’s who got rich did so at the expense of other non-rich people.
[citation needed]
Where else would the money have come from?