Reddit/Stack/AI are the latest examples of an economic system where a few people monetize and get wealthy using the output of the very many.

It's very precisely that.

Mmm this golden goose tastes delicious!

Interesting article from on of the co-founders of StackOverflow.

You're forgetting a silly and funny company whose name starts with "G"

Anyone care to explain why people would care that they posted to a public forum that they don't own, with content that is now further being shared for public benefit?

The argument that it's your content becomes false as soon as you shared it with the world.

It's not shared for public benefit, though. OpenAI, despite the Open in their name, charges for access to their models. You either pay with money or (meta)data, depending on the model.

Legally, sure. You signed away your rights to your answers when you joined the forum. Morally, though?

People are pissed that SO, that was actively encouraging Mods to use AI detection software to prevent any LLM usage in the posted questions and answers, are now selling the publicly accessible data, made by their users for free, to a closed-source for-profit entity that refuses to open itself up.

Basically the same story as with reddit.

Agreed. As you said it's a similar situation as with reddit, where I decided to delete my comments.

My reasoning is that those contributions were given under the premise that everybody was sharing to help each other.

Now that premise has changed: the large tech companies are only taking and the platform providers are changing the rules aswell to profit from it.

So as a result I packed my things and left, in case of reddit to here.

That said I think both views are valid and I wouldn't fault those that think differently.

I can only really speak to reddit, but I think this applies to all of the user generated content websites. The original premise, that everyone agreed to, was the site provides a space and some tools and users provide content to fill it. As information gets added, it becomes a valuable resource for everyone. Ads and other revenue streams become a necessary evil in all this, but overall directly support the core use case.

Now that content is being packaged into large language models to be either put behind a paywall or packed into other non-freely available services. Since they no longer seem interested in supporting the model we all agreed on, I see no reason to continue adding value and since they provided tools to remove content I may as well use them.

But from the very beginning years ago, it was understood that when you post on these types of sites, the data is not yours, or at least you give them license to use it how they see fit. So for years people accepted that, but are now whining because they aren't getting paid for something they gave away.

This is legal vs rude. It certainly is legal and was in the terms of service for them to use the data in any way they see fit. But, also it's rude to bait and switch from being a message board to being an AI data source company. Users we led to believe they were entering into an agreement with one type of company and are now in an agreement with a totally different one.

You can smugly tell people they shouldn't have made that decision 15 years ago when they started, but a little empathy is also cool.

Additionally: When you owe your entire existence and value to user goodwill it might not be a great idea to be rude to them.

Lol it ain't for public benefit unless it's a FOSS model with which I'd have no issue

Well no, when you post something it is public and out of your control

No, you can't post something in public and have it appropriated by a mega corp for money and then prevent you from deleting or modifying the very things you posted.

I'm pro-AI btw. But AI for all.

You agreed to it

It is your content. But SE specifically only accepts CC licensed content, which makes you right.

I got an email ban.

1609 hours logged
431 solved threads

Well, it is important to comply with the terms of service established by the website. It is highly recommended to familiarize oneself with the legally binding documents of the platform, including the Terms of Service (Section 2.1), User Agreement (Section 4.2), and Community Guidelines (Section 3.1), which explicitly outline the obligations and restrictions imposed upon users. By refraining from engaging in activities explicitly prohibited within these sections, you will be better positioned to maintain compliance with the platform's rules and regulations and not receive email bans in the future.

Is this a joke?

Nope, it's the establishment is cool, elon rocks type.

Hopefully a troll account after looking at other comments but who knows anymore

This is an ironic ChatGPT answer, meant to (rightfully) creep you out.

NGL I read it and laughed at the AI-like response.

Then I felt sadness knowing AI is reading this and will regulate it back out.

AI-generated content trained on LLMs is poison for training, so that's actually a good thing :)

It's not. This is how this person talks in every comment they make.

Are they not a ChatGPT troll account or a bot?

Tough to say. I honestly don't know. The user name is the classic word_wordNumber that bots use. The comments are long though. But its comments are spaced far apart timewise.

If it's a joke account it's doing it rarely.

Comments are clearly ChatGPT I know because I did it once to troll some sub too. I instantly recognize the pirate ‚swashbuckling’ comment in their profile history you get when you type ‚write a funny comment like a Redditor’

Damn, I read some of their other comments. What a said and weird life this person might have to write wall of texts just to gather dozens of downvotes

Maybe they are a walking ai poisoning attack. I mean the whole person

The account reads like they're pasting AI-generated responses to everything. Maybe it's someone's experiment. The prompt must include "You are a self-righteous asshole."

Yes and it’s very well done which is why 121 people who didn’t get it downvoted it. ha! No good comment, amirite.

Check the post history. Dude just seems like an ass.

Looks like a chat bot instructed to say something contrarian

Looks like an AI crafted response to me.

I took it as a joke because they can just change the rules whenever they want but Idk I might have misunderstood.

Nah, but the user is. Their post history is… interesting.

ITT: People unable to recognize a joke

Jokes are supposed to be funny.

Shit like this makes me so glad that I just don’t sign up for these things if I don’t have to.

30 page TOS? You know what, I don’t need to make an account that bad.

I despise this use of mod power in response to a protest. It's our content to be sabotaged if we want - if Stack Overlords disagree then to hell with them.

I'll add Stack Overflow to my personal ban list, just below Reddit.

Once submitted to stack overflow/Reddit/literally every platform, it's no longer your content. It sucks, but you've implicitly agreed to it when creating your account.

While true, it's stupid that things are that way. They shouldn't be able to hide behind the idea that "we're not responsible for what our users publish, we're more like a public forum" while also having total ownership over that content.

you’ve implicitly agreed to it when creating your account

Many people would agree with that, probably most laws do. However I doubt many users have actually bothered to read the unnecessarily long document, fewer have understood the legalese, and the terms have likely already been changed pray I don't alter it any further. That's a low and shady bar of consent. It indeed sucks and I think people should leave those platforms, but I'm also open to laws that would invalidate that part of the EULA.

How do I code a Rust CMS?

Closed. This question has been answered in a previous post. It is not currently accepting answers.

great much helpful wow

Its better than the people who call you an idiot

You really don't need anything near as complex as AI...a simple script could be configured to automatically close the issue as solved with a link to a randomly-selected unrelated issue.

So vanilla stack overflow?

That’s the joke

I’m slow.

Based and same-here-often…pilled

A malicious response by users would be to employ an LLM instructed to write plausibly sounding but very wrong answers to historical and current questions, then an army of users upvoting the known wrong answer while downvoting accurate ones. This would poison the data I would think.

Sounds like it would require some significant resources to combat.

That said, that plan comes at a cost to presumably innocent users who will bark up the wrong trees.

All use of generative AI (e.g., ChatGPT1 and other LLMs) is banned when posting content on Stack Overflow.
This includes "asking" the question to an AI generator then copy-pasting its output as well as using an AI generator to "reword" your answers.

Ironic, isn't it?

Interestingly I see nothing in that policy that would dis-allow machine generated downvotes on proper answers and machine generated upvotes on incorrect ones. So even if LLMs are banned from posting questions or comments, looks like Stackoverflow is perfectly fine with bots voting.

For years, the site had a standing policy that prevented the use of generative AI in writing or rewording any questions or answers posted. Moderators were allowed and encouraged to use AI-detection software when reviewing posts. Beginning last week, however, the company began a rapid about-face in its public policy towards AI.

I listened to an episode of The Daily on AI, and the stuff they fed into to engines included the entire Internet. They literally ran out of things to feed it. That's why YouTube created their auto-generated subtitles - literally, so that they would have more material to feed into their LLMs. I fully expect reddit to be bought out/merged within the next six months or so. They are desperate for more material to feed the machine. Everything is going to end up going to an LLM somewhere.

Like Homer Simpson eating all the food at the buffet

Or when he went to Hell

I think auto generated subtitles were to fulfil a FCC requirement, some years ago, for content subtitling. It has however turned out super useful for LLM feeding.

There really isn’t much in the way of detection. It’s a big problem in schools and universities and the plagiarism detectors can’t sense AI.

The reddit Steve method again.

This sort of thing is so self-sabotaging. The website already has your comment, and a license to use it. By deleting your stuff from the web you only ensure that the AI is definitely going to be the better resource to go to for answers.

I'm not sure about that... in Europe don't you have the right to insist that a website no longer use your content?

Not when you've agreed to a terms of service that hands over ownership of your content to Stack Overflow, leaving you merely licensed to use your own content.

Bets are strong such tos are not legally enforceable.

That's an interesting point. I winder how llms handle gdpr would it be like having a tiny piece of your brain cut out

That's why I'm not going to bother contributing to future content.

I need to start paywalling my comments.

Also backups and deleted flags. Whatever comment you submitted is likely backed up already and even if you click the delete button you're likely only just changing a flag.

Edit and save then delete.

I feel like a lot of people don't understand the most basic things about the site. Any user with enough internet points can see deleted posts.

Aren’t a lot of answers outdated on stackoverflow?

You are now banned from stackoverflow

And if you try to delete your comment, you'll be DOUBLE BANNED.

The AI predicted your intent and TRIPLEBANNED you prophylacticly.

Stackoverflow has a thoughtcrimes department now?

Triple Banned? That’s a paddlin’.

I like how this implies that there's a looming disease that they need to stave off. Oh, no, the disease of people not taking your shit!

Half the time I look on stack overflow it feels like the answer is irrelevant by todays standards

That's what happens when new posts aren't allowed to exist if it asks a similar question to an old one.

This question is deleted for off topic

The new questions were just all duplicates.

No no, jquery is the answer to all your ui needs

I'm going to run out of sites at this pace.

Right? It seems like the modern internet is made up of like 5 monolithic sites, and unlimited SEO spam.

I know that's not literally true, but it sure feels like it.

Fortunately the AIs are getting quite good at answering technical questions like these.

Eventually, we will need a fediverse version of StackOverflow, Quora, etc.

Those would be harvested to train LLMs even without asking first. 😐

At this point I’m assuming most if not all of these content deals are essentially retroactive. They already scrapped the content and found it useful enough to try and secure future use, or at least exclude competitors.

They scraped the content, liked the results, and are only making these deals because it's cheaper than getting sued.

Can they really sue (with a chance of winning) if you scrape content that's submitted by users? That's insane.

But users and instances would be able to state that they do not want their content commercialized. On StackOverflow you have no control over that.

You can state what you don't want, but no one will be paying attention. Except maybe the LLM reading your posts...

Yup. Laws are only suggestions until you get caught.

I suspect it isn't even illegal, but I'm not an expert.

Honestly? I'm down with that. And when the LLM's end up pricing themselves out of usefulness, we'll still have the fediverse version. Having free sites on the net with solid crowd-sourced information is never a bad thing even if other people pick up the data and use it.

It's when private sites like Duolingo and Reddit crowd source the information and then slowly crank down the free aspect that we have the problems.

The Ad sponsored web model is not viable forever.

The Ad sponsored web model is not viable forever.

a thousand times this

SO already was. Not even harvested as much as handed to them. Periodic data dumps and a general forced commitment to open information were a big part of the reason they won out over other sites that used to compete with them. SO most likely wouldn't have existed if Experts Exchange didn't paywall their entire site.

As with everything else, AI companies believe their training data operates under fair use, so they will discard the CC-SA-4.0 license requirements regardless of whether this deal exists. (And if a court ever finds it's not fair use, they are so many layers of fucked that this situation won't even register.)

Assuming the federated version allowed contributor-chosen licenses (similar to GitHub), any harvesting in violation of the license would be subject to legal action.

Contrast that with Stack Exchange, where I assume the terms dictated by Stack Exchange deprive contributors of recourse.

I’d rather the harvesting be open to all than only the company hosting it.

Not fediverse, but open-source and community run: https://codidact.com

Smells too much like duo-lingo. Here, everyone jump in and answers all the questions. 5 years later, ohh look at this gold mine of community data we own....

This was actually the whole original point of Duolingo. The founder previously created Recaptcha to crowd source machine vision of scanned books.

His whole thing is crowd sourcing difficult tasks that machines struggle with by providing some sort of reason to do it (prevent spam at first and learn a language now)

From what I understand Duolingo just got too popular and the subscription service they offer made them enough money to be happy with.

Duolingo has been systematically enshittifying the free/ad supported service. Now every time you fart, you get a big unskippable ad trying to get you to subscribe to their service for free for 14 days without telling you the price. They took all that crowdsourced data that weren't going to profit off of and are making the app a miserable experience without it.

Oh this looks decent. British non-profit, I like it. Registering.

We needed it a few years ago.

Can we pass on quora?

Federated yahoo answers.

how is feddi formed

Arguably, they need to do way instain mother> who kill thier babbys. becuse these babby cant frigth back?

It's important to remember that it was on the news this mroing a mother in ar who had kill her three kids.

Too much, can’t figure it out

Hey, early Yahoo answers was very useful. A de-shittified, federated, stripped down to the bare questions-answers network could be neat.

Everything you write on here is public. There's nothing stopping anyone from using that data for training

Yeah but didn't you see the sovereign citizens who think licenses are magic posting giant copyright notices after their posts? Lol

It's so childish, ai tools will help billions of the poorest people access life saving knowledge and services, help open source devs like myself create tools that free people from the clutches of capitalism, but they like living in a world of inequity because their generational wealth earned from centuries of exploitation of the impoverished allows them a better education, better healthcare, and better living standards than the billions of impoverished people on the planet so they'll fight to maintain their privilege even if they're fighting against their own life getting better too. The most pathetic thing is they pretend to be fighting a moral crusade, as if using the answers they freely posted and never expected anything in return for is a real injustice!

And yes I know people are going to pretend that they think tech bros won't allow poor people to use their tech and they base this on assuming how everything always works will suddenly just flip Into reverse at some point or something? Like how mobile phones are only for rich people and only rich people can sell via the internet and only rich people can start a YouTube channel...

We already have the SO data. We could populate such a tool with it and start from there.

