Much of modern AI has been built using enormous quantities of online material. Creators, publishers and AI developers disagree over whether publicly accessible content should be available for training.
If we ban training on public data, we don't stop AI; we just ensure that AI only represents the wealthy and the powerful (https://www.technologyreview.com/2024/01/10/licensed-data-bias-ai/). If companies can only train on data they explicitly buy, they will buy the archives of the New York Times, Getty Images, and academic journals. They won't buy the blogs, forums, and public posts of average people, minorities, or marginalised groups. The resulting AI models will be incredibly biased, represen...
R
ruth29Jul 27, 2026
0
We need to consider the 'Right to be Forgotten.' If I post something stupid on a public forum when I'm 16, I can delete it when I'm 25. But if an LLM scraped that post in 2023, those words are permanently baked into the weights of the model. You cannot 'delete' data from a trained neural network without retraining the entire multi-million dollar model from scratch (which companies won't do) (https://arxiv.org/abs/2209.00939). Public scraping destroys the human right to evolve and erase our past....
D
diana12Aug 25, 2026
0
The US Copyright Office's Part 3 report on generative AI training is the most authoritative analysis of the legal landscape (https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf). Their conclusion is nuanced: some training uses may be fair use, others may not be, and the determination is fact-specific. The policy question is whether we want to wait for courts to resolve this case by case over a decade, or whether we w...
K
kyle20Sep 3, 2026
0
The internet was built on the premise of machine readability. Search engines have been scraping the entire internet for 25 years to build their indexes, and we accepted it because they sent traffic back to the creators. The problem with LLMs isn't the scraping; it's the lack of attribution and traffic routing. If AI companies figure out a micro-transaction model to pay creators a fraction of a cent every time their data influences an output, the entire copyright debate disappears (https://a16z.c...
A
amy102Sep 5, 2026
0
The practical reality is that the models that exist were trained on public internet content and that can't be undone (https://www.youtube.com/watch?v=MvnN70at75I). The policy question is forward-looking: what rules should govern future training? I'd argue for an opt-out system with clear mechanisms, creators who want to exclude their work from training can do so, and AI companies must honour those requests. That's not perfect but it's implementable. Would an opt-out system satisfy you as a creat...
Z
zane107Sep 7, 2026
0
As someone whose work has almost certainly been used to train AI models without my consent or compensation, I have a direct stake in this. The MIT Sloan 'learnrights' proposal is the most interesting solution I've seen, it would create a licensing framework where creators are compensated for their contribution to training data (https://mitsloan.mit.edu/ideas-made-to-matter/how-learnrights-would-compensate-creators-ai-model-training). The current situation, where my decade of creative work is fre...
Z
zane107Jul 19, 2026
0
Also, 'representing my worldview' doesn't pay my rent. If an AI company uses my portfolio to train an image generator that puts me out of work, I don't care how 'diverse' the model is. I care that I was robbed.
L
LaborUnionRepSep 11, 2026
0
@CognitiveScientist This is a false dichotomy. You are arguing that the only way to achieve 'diversity' is through corporate theft. We could build public, state-funded data trusts where citizens voluntarily donate their data for the public good, rather than letting private corporations strip-mine the internet for profit.
Join the Conversation
Share your AI tool experiences and help others make informed decisions.
This forum is actively moderated. All posts and replies can be reported by community members using the Report button. Our team reviews flagged content to keep discussions constructive and safe. Read our Community Guidelines for more details.