I know lemmy has a big hard on for anti-AI but a few things here:
- AI is over 60 years old and includes things like machine learning. Not all AI is generative or large language models.
- There is nothing fascist about running open AI (not “openAI”) models on your own hardware and that’s the future Cory Doctorow (mr “enshittification”) makes in his latest book The Reverse-Centaurs Guide to Life After AI. Basically he theorizes when the AI bubble collapasses it will be terrible but it will shift AI into the hands of consumers and allow people to use AI to do work for them, rather than their bosses using AI to replace workers.
- Does using facial recognition against ICE count as fascism?
There are many legitimate criticisms against AI, especially generative models, but this ain’t one
The problem is that the general public has adopted “AI” to mean the current popular machine learning thing, the general category of “generative AI”, and if you want a message about it to be received broadly, you need to use terms people will understand.
In my case, I’m not aware of any GenAI models that produce useful output that I’d consider to be ethical, so I support this message.
The problem is that the general public has adopted “AI” to mean the current popular machine learning thing,
Yes, that is the problem, because people don’t understand that it’s a short-hand, and don’t understand the difference between harmful uses, and non-harmful uses. The concept itself becomes a meaningless tribal indicator of hate. “This person is acting unethically and unlawfully,” has quickly become “I don’t understand what this technology is, but when I see ‘AI’ I automatically hate whatever it’s associated with.” That is counter-productive to humanity as a whole.
So stop saying that any given technology or method is intrinsically bad, and start focusing on WHO is doing the harm and HOW.
If I’m having a direct discussion with somebody, then I’ll make it clear what I’m talking about, and if they engage in a conversation about ethics, I’ll explain why those things are unethical.
But if I wanted to post a public message to voice my sentiment, I don’t think trying to post a long explanation every time would do any good, most people would probably just skip it when they saw it’s longer than a couple lines of text. So for those cases, I’d say you have to go down to the level of current discourse to engage with the masses on their terms if you want to be heard.
But if I wanted to post a public message to voice my sentiment, I don’t think trying to post a long explanation every time would do any good
That’s a non sequitur. The initial point was “Stop using this term incorrectly, because it spreads misunderstanding and misinformation.”
Nobody said you should write a thesis per reply, and my entire point was to focus on the harms, and the actors, not the tech.
Saying “AI is unethical,” is a truly nonsensical statement. Even if we grant that AI is shorthand for LLM, is still isn’t justified as all I need to do is find an instance of an open source model being trained on public data to refute it, and there are MANY.
Unless you think that using public data to make public software is inherently unethical, which is a hard stance to justify.
If you’re upset with Sam Altman and ChatGPT, then you can simply say “Sam Altman is a thief, and ChatGPT uses stolen data.” Which is infinitely more defendable than “AI is trained on stolen data and is unethical.”
If you’re upset with Sam Altman and ChatGPT, then you can simply say “Sam Altman is a thief, and ChatGPT uses stolen data.” Which is infinitely more defendable than “AI is trained on stolen data and is unethical.”
The issue is, I’d say the same applies to every model that produces useful outputs. LLaMa, Anthropic, Grok, DeepSeek, whatever else is out there. If you tell somebody that the LLM they’re using is unethical, they’ll nust go use another convenient corporate model.
And “public data” is not enough for me, because that typically means scraping copyrighted content from public websites, I’m not aware of a model that uses only data with permission (either explicit or granted by the license) that’s useful, and the people who need to hear about ethical problems with GenAI especially don’t know about that.
The issue is, I’d say the same applies to every model that produces useful outputs.
That you know of. Your lack of awareness is not an indication on the stance of the technology itself. If you have a problem with Grok, as I do, then condemn Grok, Twitter, and that pathetic man child that owns them.
If you tell somebody that the LLM they’re using is unethical, they’ll nust go use another convenient corporate model.
Yeah, so actively fight against these toxic corporate activities. Fuck OpenAI. Fuck Google. Fuck Twitter. But the instant negative reaction to their shared tool is getting quite a few additional people caught in the blast.
And “public data” is not enough for me, because that typically means scraping copyrighted content from public websites, I’m not aware of a model that uses only data with permission
Again, this is an expression of your ignorance, not reality.
OLMo 2 was trained on Wikipedia and other fully public forums. Its training sources and data is fully accessible and open.
GPT-NeoX is an untrained model that you can train yourself. It’s literally just the foundations anybody could use.
Pythia is fully open source, and one of its training sources was GitHub. Which does genuinely bring in to question whether or not you can use GitHub to train AI. Personally, I don’t see how branching a repository is any different than using the code to train a model. But the current anti-AI trend has people EXTREMELY sensitive to this concept.
Again, this is an expression of your ignorance, not reality.
OLMo 2 was trained on Wikipedia and other fully public forums. Its training sources and data is fully accessible and open.
Decided to check out the first example quickly. It’s hard to dig through the information, but following the chain of sources:
In the first stage, which covers over 90% of the total pretraining budget, we use the OLMo-Mix-1124, a collection of approximately 3.9 trillion tokens sourced from DCLM, Dolma, Starcoder, and Proof Pile II.
As part of DCLM, we provide a standardized corpus of 240T tokens extracted from Common Crawl
So, it’s using data from Common Crawl. What data, exactly? That’d be harder to dig up. DCLM has a repository, but they don’t make an effort to point out how, or if, they’re filtering the data.
What I can quickly find is information from Common Crawl itself, which is the ultimate source of data. On that I can immediately see only two things:
- They have an opt-out list, and from what I understand, they’re scraping websites by default, providing mechanisms for you to block them or opt-out… If you’re even aware they exist.
- According to their stats page, they have 313423 pages scraped from
github.com. I doubt github has that many pages of non-user-generated content, so the question is… What data is going in there, who wrote it, and did anybody agree to that? What about other websites from those top domains, such asblogspot.com,wordpress.org,readthedocs.io? I doubt they got permission to use the 17175161 pages of content from blogspot they’ve scraped.
So yeah, maybe there’s a “good” model out there I’d actually accept, but I’ve seen “open” models being released, and I can’t possibly check all of them, just checking the websites for one of the sources for one of the models took 20 minutes here, if I wanted to verify this properly I’d have to setup the tooling to query the terabytes of data for this source, and all others, and even there I’m not sure if I’d find answers.
Largely agree. AI isn’t going away; the left needs to take a page out of Sanders and advocate instead for nationalizing and regulating instead.
There is no such thing as “fascist technology”. Technology is just a tool. It can be used for good or evil.
Except fascists are more likely to use it because they are artistic nincompoops
I see. So using it to cure cancer has to be stopped in its tracks now cuz you had a shallow political opinion.
No one is using LLMs to cure cancer.
Medical AI and LLM’s are not the same thing
Read it again. Post says AI. all AI not just LLM.
like i said: shallow political opinions. just so you idiots can rage bait. this is pathetic.
I think there’s a pretty clear implication of referring to generative AI specifically when lay people use the term at this point, especially when creative labor is mentioned.
Yes, this is good. Keep doing everything you can to make yourselves weaker.
to make yourselves weaker.
Considering that reliance on AI weakens people’s critical thinking skills, I’m curious what your reasoning is here.
Aren’t they taking a loss on every AI operation? Like a big loss. So use enough, especially free versions, and you will drive them to bankruptcy. Lol
They can just make more money to keep the bubble from popping. There is a limit to that but how much damage to the environment will they do before that point?
How? They’re running at a huge loss.
Anthropic’s last quarter posted a profit.
The machine gun was an imperialist and fascist technology of oppression. But call it a Kalashnikov and it becomes a tool for liberation.
Do not listen to neoliberal fuckai morons. Think for yourself what tools you can use or not. If machine generated AI slop is indistinguishable from real human “art”, then it was never art but human slop to begin with. What generations of rom-coms and participation trophies does to a people’s ego.
Push for meaningful regulation of AI instead of endless whining and villifying people like e.g. non-capitalist china does: https://regulations.ai/regulations/china-summary
Guns and Muscles are fascist too! Fighting is fascist.
Guns were invented literally thousands of years before fascism
Firearms are barely a 1000 year old technology.
have you ever heard of the hyperbowl






