Would be a terrible shame if lots of people opted out.
If you are a EU citizen you might also want to write a complaint to privacy@huggingface.co because they are collecting your personally identifiable information in machine-readable form which they are distributing to third parties.
Good thing I never finished a project!
They’re not able to scrape private repos, surely?
No, it seems to only be a subset of public repo’s.
I have like 65 repo’s and only 13 were scraped. I don’t get why they specifically scraped those though. They don’t have the most stars, they aren’t the oldest or newest, not the ones with the most forks, nor do I see a pattern based on programming language.Microsoft has access, and everything is for sale. So, maybe? Probably? 🤷
Embrace extend extinguish
When I hear this company’s name I can’t get the idea out of my head that it’s related to the face huggers from Alien.
Am I alone in not wanting to put my username into that field? If they don’t have it will they then just decide that it’s now a good time to scrape it? Or are they going to record that it was searched?
I just followed the opt out link and they’re making you open a public issue with a Markdown list with all your repositories you want removed.
Surely there has to be a better way to do this (probably the point).
Also why are they storing 4.71 TB in a Git repo? How are they going to deal with removal requests?
EDIT: It seems like they’re manually responding to the issues, wtf? There’s a perfectly fine GitHub auth system that they could use to verify everyone’s GitHub account / repository ownership
EDIT 2: This is apperently a collaboration between Hugging Face and ServiceNow, why is my companies IT ticketing system scraping all of GitHub?
with a Markdown list with all your repositories you want removed.
The repo readme linked FAQ says
You can choose to request either (1) all repos, or (2) you can specify select repos that you own to be removed.
so “all of them” should be acceptable
“Oh no, people are using information I put publicly available on the internet for everyone to see!”
Morons. The lot of you.
Edit: I rest my case.
I did it to invite collaboration and connect with other developers with similar interests. FOSS is more about building communities than building software, after all.
I did not anticipate that it could be (legally) used to dismantle the kinds of communities I wanted to build. (I did anticipate that it could be illegally used to that end, but historically that has tended to cause a Streisand Effect, so that risk seemed worth it.)
If people still want to be part of the community, they can.
You’re just coming up with reasons to fit in with the crowd.
Love you too, kitten. ♥️
Sure. Let’s see if they use it for endeavours in the same spirit.
Or are you fine with they using this data for-profit without benefitting the public by also making it open?
(not talking about hf, I know starcoder. just in general)
Maybe you should’ve released your code under a license with that stipulation.
In all honesty, you’re just coming up with reasons to fit in with the crowd.
Give me a crystal ball next time and I’ll do it genius, use your common sense.
Seems fitting the one thinking about “the crowd” the most is exactly the one trying to be the most distant from it. Projecting much are we?
Don’t hurt yourself with those mental gymnastics.
Hey, sure. I’m doing all kinds of gymnastics to think the ones profiting off our public data are in the wrong, I’m just in it to fit in and be in the crowd sure! Whatever makes you feel different. Hope your mirror works someday.
See? Even now you’re still trying to save face.
You need help.
Oh God why are all your replies in the same structure, am I talking to an LLM? I can’t decide whether that would be preferrable or not
We want to give developers agency over their source code by letting them decide whether or not it should be used to develop and evaluate machine learning models.
fuck them seriously; if you want to do that then don’t steal the repositories in the first place.
Unless I’m mistaken, this wasn’t written by the folks that scaped GitHub in the first place, someone just wrote a small tool to semi-automate the process of searching the scraped dats, and submitting a GitHub issue to have it removed.
it’s pretty clear from the language of the site, the fact that it’s on the hugging face domain, and the fact that the github organisation for the “opt out” makes it clear it’s hugging face.
don’t steal the repositories
How is this remotely stealing?

It’s not, but it may be violation of licenses. And also, if it has personal information on it, that’s probably illegal under the GDPR.
but it may be violation of licenses
They excluded code with non-permissive licenses apparently:
Each file is labelled permissive (at least one permissive license detected, no conflicting non-permissive license), no_license (no licenses detected, or only non-license legal texts such as CLAs), or non_permissive. The permissive allowlist follows the Blue Oak Council list plus licenses categorized as Permissive or Public Domain by ScanCode. Files classified as non_permissive are excluded from both released datasets.
also, if it has personal information on it, that’s probably illegal under the GDPR.
It’s all public so I would be extremely surprised if that were the case.
the method on github is to fork it not steal all the data in an external storage for external purposes without consent of the user.
edit: you can also report their whole account to github here: https://support.github.com/contact/report-abuse?category=report-abuse&report=bigcode-project&report_id=110470554&report_type=user
I’m pretty sure they just cloned the repos. That’s how GitHub is designed to work. Have you never cloned a repo from GitHub?
Yes, but I used it under the licence terms of the original developer.
So have they. They filter by license - see my other comment.
It doesn’t seem that way. My list contains several repositories that don’t have licenses, even some with GPL.
so if you are licenced under MIT and they use your code, they publish your copyright header?
Yes. If they train AI from your code? No, but the legality of that is yet to be settled and definitely leaning towards “it’s fine”.
not for private for profit use from shitty ai companies though…
Kinda hope it uses my code. It’s so terrible there’s no doubt it will make the resulting code from the model worse even if the impact is miniscule
I just opted out of all my repos except the God awful ones from middle and highschool. Those are basically prompt poison so fuckem
Modern tech companies love using the rapsist’s model of consent.
Meanwhile at work we just had a training course that specifically said doing “opt out” instead of “opt in” violates the principle of informed consent.
Notice how a shit load of these people keep turning up to be rapists and pedophiles… They have no shits to give about informed consent. In their mind you don’t even have the right to the same agency they do.
not surprising since the venn diagram of ai bros and rapists is a circle.
Is there a form to ask to be included in the next stack? they seem to have missed me this time
Scraping without consent steals developer personalIP.
they can do whatever they want to do with my code.
as long as it’s complaint to AGPL v3 :D
(doubt ai companies care about that)
Much like… -reads notes- Microsoft is via Github… huh…
Gothub.com is a thing now for OpenBSD’s Game of Trees project.
It’s gothub.org. You linked to some dumb AI startup.
It’s a sure sign of a bubble that a URL typo coincidentally brings one to a completely unrelated page for a completely unknown, in-beta AI startup with a dumb idea.
Since they stole my paper on ethics in computer science, maybe the model will learn to act better than its owners
tbh i’m thinking this alone isn’t that bad from an archiver/datahoarder perspective
This is what confuses me. The internet archive has been archive the entire Internet for years. Yet AI does the same to make a way for people to code easier and it is a problem all of the sudden?
IA is a nonprofit and archives to preserve human history. shitty AI startups do this to monetize the data, and their end goal is to “replace” the people who made that data in the first place.
Thanks to these AI mfs the IA now prevents access to many items because they can be used as training material which fucking sucks
As a developer, you hold the copyright to your code. When you make it open-source, you grant a license to use the code and the resulting program under certain terms.
This is a contract. If you copy my code without following these terms, then that’s theft.
The Internet Archive’s use complies with these terms for all open-source licenses. These AI companies do not. In particular, here’s a quote from the MIT license, which you will find in a similar wording in all open-source licenses:
The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.
In effect, what this means, is that when you copy my code, I demand that you also copy the license text along with it, so that anyone else looking at this code knows the permissions I grant and the terms I require.
And now guess what these AI companies are doing. They copy my code and reproduce substantial portions upon a user asking, yet they do not include my license terms. They violate the contract under which they obtained my source code.
I suspect you don’t realize how shit that is, because source code is so abstract.
It’s like spending hundreds of hours painting a great artwork and then deciding that everyone should be able to give a copy to everyone they know, under the simple condition that they inform those people that they have this right as well.
And then comes along a company and sells my artwork for money, without informing their customers that they can pass it on for free. That’s, plain and simple, a criminal operation.Deleted by author
If nobody owns the code, then then nobody can enforce the terms of the license it was released under, and free software under the FSF definition becomes impossible. All you have is public domain.
For example, a company could take the Linux kernel, modify it and distribute it with their gadgets. And they could simply not release the modifications they’ve made, as is required by the GNU Public License. But nobody would be able to do anything about it. Currently, copyright laws allow the people who wrote the Linux kernel to sue the company for breaking the license and violating the authors’ copyrights
In our current legal system, copyright is the basis for me to be able to set requirements on how my code can be shared. I do not care that I own it, I just care that it is shared under the conditions I set.
Without being able set these conditions, I would not open up my code.
No one. What people give a shit about is the license that is supposed to keep the code open, which is being removed for profit without consequence.
Internet Archive exists as a reference for your edification on a donation basis; AI companies intend to initiate a top down societal restructuring of jobs and thus access to benefits, paywall access to your own collective information, fund themselves through ouroboros leveraged deals and VC money (value that’s been stolen from the general populace over the years)
Deleted by author
As long as the datasets are open, it is our best hope. I know it doesn’t compensate the people whose work’s copyright and licenses have been violated, but I think it’s the only realistic hope we’ve got of getting out of this informational dystopia with a reasonably intact library of humanity’s knowledge that hasn’t been locked down and/or monetized. The AI scrapers and generators are in the process of burning down the great library of Alexandria that the Internet had become, and we are already starting to feel its loss. We cannot stop the wave of toxic pollution that is spreading through all our digital content now, but the archives from before this apocalypse started will become the most valuable thing humanity has ever produced. This is information war, and we are losing.
Wonder how many of those repos contain AI-generated code?
if I opt out, will the Roko’s basilisk come after me?
Reported for cognitohazard. Delete this immediately.
(I kid, but someone really did report your comment.)
lmao
As long as you haven’t been told what you’re opting out of, you’re good.
I think its hilarious that some people seem to actually take this science fiction variation on Pascal’s wager seriously.















