Would be a terrible shame if lots of people opted out.

If you are a EU citizen you might also want to write a complaint to privacy@huggingface.co because they are collecting your personally identifiable information in machine-readable form which they are distributing to third parties.

    • qaz@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      2 months ago

      No, it seems to only be a subset of public repo’s.
      I have like 65 repo’s and only 13 were scraped. I don’t get why they specifically scraped those though. They don’t have the most stars, they aren’t the oldest or newest, not the ones with the most forks, nor do I see a pattern based on programming language.

  • Prior_Industry@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    2 months ago

    When I hear this company’s name I can’t get the idea out of my head that it’s related to the face huggers from Alien.

  • purplemonkeymad@programming.dev
    link
    fedilink
    arrow-up
    0
    ·
    2 months ago

    Am I alone in not wanting to put my username into that field? If they don’t have it will they then just decide that it’s now a good time to scrape it? Or are they going to record that it was searched?

  • qaz@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    edit-2
    2 months ago

    I just followed the opt out link and they’re making you open a public issue with a Markdown list with all your repositories you want removed.

    Surely there has to be a better way to do this (probably the point).

    Also why are they storing 4.71 TB in a Git repo? How are they going to deal with removal requests?

    EDIT: It seems like they’re manually responding to the issues, wtf? There’s a perfectly fine GitHub auth system that they could use to verify everyone’s GitHub account / repository ownership

    EDIT 2: This is apperently a collaboration between Hugging Face and ServiceNow, why is my companies IT ticketing system scraping all of GitHub?

    • Kissaki@programming.dev
      link
      fedilink
      English
      arrow-up
      0
      ·
      2 months ago

      with a Markdown list with all your repositories you want removed.

      The repo readme linked FAQ says

      You can choose to request either (1) all repos, or (2) you can specify select repos that you own to be removed.

      so “all of them” should be acceptable

  • professor_prime@lemmychan.org
    link
    fedilink
    arrow-up
    0
    ·
    edit-2
    2 months ago

    “Oh no, people are using information I put publicly available on the internet for everyone to see!”

    Morons. The lot of you.

    Edit: I rest my case.

    • kibiz0r@midwest.social
      link
      fedilink
      English
      arrow-up
      0
      ·
      2 months ago

      I did it to invite collaboration and connect with other developers with similar interests. FOSS is more about building communities than building software, after all.

      I did not anticipate that it could be (legally) used to dismantle the kinds of communities I wanted to build. (I did anticipate that it could be illegally used to that end, but historically that has tended to cause a Streisand Effect, so that risk seemed worth it.)

    • lavember@programming.dev
      link
      fedilink
      arrow-up
      0
      ·
      2 months ago

      Sure. Let’s see if they use it for endeavours in the same spirit.

      Or are you fine with they using this data for-profit without benefitting the public by also making it open?

      (not talking about hf, I know starcoder. just in general)

      • professor_prime@lemmychan.org
        link
        fedilink
        arrow-up
        0
        ·
        2 months ago

        Maybe you should’ve released your code under a license with that stipulation.

        In all honesty, you’re just coming up with reasons to fit in with the crowd.

        • lavember@programming.dev
          link
          fedilink
          arrow-up
          0
          ·
          2 months ago

          Give me a crystal ball next time and I’ll do it genius, use your common sense.

          Seems fitting the one thinking about “the crowd” the most is exactly the one trying to be the most distant from it. Projecting much are we?

            • lavember@programming.dev
              link
              fedilink
              arrow-up
              0
              ·
              2 months ago

              Hey, sure. I’m doing all kinds of gymnastics to think the ones profiting off our public data are in the wrong, I’m just in it to fit in and be in the crowd sure! Whatever makes you feel different. Hope your mirror works someday.

  • katy ✨@piefed.blahaj.zone
    link
    fedilink
    English
    arrow-up
    0
    ·
    2 months ago

    We want to give developers agency over their source code by letting them decide whether or not it should be used to develop and evaluate machine learning models.

    fuck them seriously; if you want to do that then don’t steal the repositories in the first place.

    • JakenVeina@midwest.social
      link
      fedilink
      arrow-up
      0
      ·
      2 months ago

      Unless I’m mistaken, this wasn’t written by the folks that scaped GitHub in the first place, someone just wrote a small tool to semi-automate the process of searching the scraped dats, and submitting a GitHub issue to have it removed.

      • katy ✨@piefed.blahaj.zone
        link
        fedilink
        English
        arrow-up
        0
        ·
        2 months ago

        it’s pretty clear from the language of the site, the fact that it’s on the hugging face domain, and the fact that the github organisation for the “opt out” makes it clear it’s hugging face.

      • PlexSheep@infosec.pub
        link
        fedilink
        arrow-up
        0
        ·
        2 months ago

        It’s not, but it may be violation of licenses. And also, if it has personal information on it, that’s probably illegal under the GDPR.

        • FizzyOrange@programming.dev
          cake
          link
          fedilink
          arrow-up
          0
          ·
          2 months ago

          but it may be violation of licenses

          They excluded code with non-permissive licenses apparently:

          Each file is labelled permissive (at least one permissive license detected, no conflicting non-permissive license), no_license (no licenses detected, or only non-license legal texts such as CLAs), or non_permissive. The permissive allowlist follows the Blue Oak Council list plus licenses categorized as Permissive or Public Domain by ScanCode. Files classified as non_permissive are excluded from both released datasets.

          also, if it has personal information on it, that’s probably illegal under the GDPR.

          It’s all public so I would be extremely surprised if that were the case.

    • lps2@lemmy.ml
      link
      fedilink
      arrow-up
      0
      ·
      2 months ago

      Kinda hope it uses my code. It’s so terrible there’s no doubt it will make the resulting code from the model worse even if the impact is miniscule

      • kboy101222@sh.itjust.works
        link
        fedilink
        English
        arrow-up
        0
        ·
        2 months ago

        I just opted out of all my repos except the God awful ones from middle and highschool. Those are basically prompt poison so fuckem

    • kibiz0r@midwest.social
      link
      fedilink
      English
      arrow-up
      0
      ·
      2 months ago

      Meanwhile at work we just had a training course that specifically said doing “opt out” instead of “opt in” violates the principle of informed consent.

      • atomicbocks@sh.itjust.works
        link
        fedilink
        English
        arrow-up
        0
        ·
        2 months ago

        Notice how a shit load of these people keep turning up to be rapists and pedophiles… They have no shits to give about informed consent. In their mind you don’t even have the right to the same agency they do.

  • gnawmon@ttrpg.network
    link
    fedilink
    arrow-up
    0
    ·
    2 months ago

    they can do whatever they want to do with my code.

    as long as it’s complaint to AGPL v3 :D

    (doubt ai companies care about that)

  • goatbeard@beehaw.org
    link
    fedilink
    arrow-up
    0
    ·
    2 months ago

    Since they stole my paper on ethics in computer science, maybe the model will learn to act better than its owners

    • korendian@piefed.social
      link
      fedilink
      English
      arrow-up
      0
      ·
      2 months ago

      This is what confuses me. The internet archive has been archive the entire Internet for years. Yet AI does the same to make a way for people to code easier and it is a problem all of the sudden?

      • hexagonwin@lemmy.today
        link
        fedilink
        arrow-up
        0
        ·
        2 months ago

        IA is a nonprofit and archives to preserve human history. shitty AI startups do this to monetize the data, and their end goal is to “replace” the people who made that data in the first place.

        Thanks to these AI mfs the IA now prevents access to many items because they can be used as training material which fucking sucks

      • trem@lemmy.blahaj.zone
        link
        fedilink
        arrow-up
        0
        ·
        2 months ago

        As a developer, you hold the copyright to your code. When you make it open-source, you grant a license to use the code and the resulting program under certain terms.

        This is a contract. If you copy my code without following these terms, then that’s theft.

        The Internet Archive’s use complies with these terms for all open-source licenses. These AI companies do not. In particular, here’s a quote from the MIT license, which you will find in a similar wording in all open-source licenses:

        The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.

        https://mit-license.org/

        In effect, what this means, is that when you copy my code, I demand that you also copy the license text along with it, so that anyone else looking at this code knows the permissions I grant and the terms I require.

        And now guess what these AI companies are doing. They copy my code and reproduce substantial portions upon a user asking, yet they do not include my license terms. They violate the contract under which they obtained my source code.

        I suspect you don’t realize how shit that is, because source code is so abstract.
        It’s like spending hundreds of hours painting a great artwork and then deciding that everyone should be able to give a copy to everyone they know, under the simple condition that they inform those people that they have this right as well.
        And then comes along a company and sells my artwork for money, without informing their customers that they can pass it on for free. That’s, plain and simple, a criminal operation.

          • fruitcantfly@programming.dev
            link
            fedilink
            arrow-up
            0
            ·
            2 months ago

            If nobody owns the code, then then nobody can enforce the terms of the license it was released under, and free software under the FSF definition becomes impossible. All you have is public domain.

            For example, a company could take the Linux kernel, modify it and distribute it with their gadgets. And they could simply not release the modifications they’ve made, as is required by the GNU Public License. But nobody would be able to do anything about it. Currently, copyright laws allow the people who wrote the Linux kernel to sue the company for breaking the license and violating the authors’ copyrights

          • trem@lemmy.blahaj.zone
            link
            fedilink
            arrow-up
            0
            ·
            2 months ago

            In our current legal system, copyright is the basis for me to be able to set requirements on how my code can be shared. I do not care that I own it, I just care that it is shared under the conditions I set.

            Without being able set these conditions, I would not open up my code.

          • Squirrelanna@lemmy.blahaj.zone
            link
            fedilink
            arrow-up
            0
            ·
            2 months ago

            No one. What people give a shit about is the license that is supposed to keep the code open, which is being removed for profit without consequence.

      • Jtotheb@lemmy.world
        link
        fedilink
        arrow-up
        0
        ·
        2 months ago

        Internet Archive exists as a reference for your edification on a donation basis; AI companies intend to initiate a top down societal restructuring of jobs and thus access to benefits, paywall access to your own collective information, fund themselves through ouroboros leveraged deals and VC money (value that’s been stolen from the general populace over the years)

    • cecilkorik@lemmy.ca
      link
      fedilink
      English
      arrow-up
      0
      ·
      2 months ago

      As long as the datasets are open, it is our best hope. I know it doesn’t compensate the people whose work’s copyright and licenses have been violated, but I think it’s the only realistic hope we’ve got of getting out of this informational dystopia with a reasonably intact library of humanity’s knowledge that hasn’t been locked down and/or monetized. The AI scrapers and generators are in the process of burning down the great library of Alexandria that the Internet had become, and we are already starting to feel its loss. We cannot stop the wave of toxic pollution that is spreading through all our digital content now, but the archives from before this apocalypse started will become the most valuable thing humanity has ever produced. This is information war, and we are losing.