Human-LLM Language Collapse, or why AI May Lead to Stagnation and Homogenisation of Culture and Technology
Aug 2, 2026
There are scenarios that could play out with AI that I worry about, and this is one of them.
Model collapse is the idea that models trained on the output of others models are inferior, with less diversity and less accuracy. It has been publicised recently (July 2026) that AI companies are buying up used —and often rare— books in order to scan and shred them. They shred them as a kind of copyright defence: scan a copy, shred a copy, only one version exists. Of course this ignores that they are then training models on the one version and incorporating it into their output weights. Whether or not that is a breach of copyright remains an open question. Regardless, apparently the companies are specifically seeking out pre-2022 books, i.e. books not yet tainted by the output of the earlier models. This suggests that the AI companies themselves are worried about model collapse.
LLMs lower the barrier to entry to writing a book, or an article, or a piece of advertising copy. This has lead to a large quantity of “AI Slop” content, particularly noticable in Amazon EBook listings and auto-generated web articles. Writers are then under pressure to compete against a far larger cohort which can output content far quicker than they can. For an author it must be incredible disheartening and overwhelming to see the ubiquity of LLM content online. I would expect under this pressure writers will either give up or turn to LLM use themselves. Perhaps existing writers won’t “give up”, but prospective writers, people who would have been writers, will be less likely to bother, because what’s the point? The knowledge that anything they write will be scooped up by an LLM and gargled back out creates additional negative pressure on non-LLM writing. We have already seen otherwise respected authors caught out using AI and this is sure to continue.
Thinking more on prospective writers, what is more concerning to me is that they will grow up in a world surrounded by LLMs. They will be watching LLM written content on YouTube, reading LLM written books, and no doubt “talking to” LLMs as part of daily life. They will get used to the “LLM voice” and sublanguage. Anyone, but especially children, will learn to parrot the style and language of those around them. My fear then is that children growing up will talk and write like an LLM. At some critical mass of affected children, this will act like herd immunity in a kind of perverse opposite way and infect the non-LLM-raised children in the same way. Put another way, if say 10% of children are raised sans-LLM, the LLM-language of the remaining 90% will cause the 10% to pick up that same voice. I don’t know what a critical mass for this would be but presumably it’s above 50%. Given how many “iPad parents” we see in the wild, 50% of children raised on LLMs is quite believable.
Sam Altman, parent, on parenting
This is a homonogenisation, a “monoculture”, of language. It can happen with LLMs like with nothing before them because of the sheer overwhelming quantity of content that they produce, and how they have become the easiest way to produce the forms of content that they do produce. This is a new and different problem.
To come back to model collapse. If we allow this hypothesis of a homogenised language and writing going forward, and we are at all worried about model collapse, then it follows that at some point the LLMs will have been trained on all reasonably useful datasets. I’m not sure training data from before the 20th century is particularly useful for writing a modern works in the modern voice, so the training corpus is somewhat bounded and well-defined. If the AI companies are already looking for rare copies of used books, and frankly the books are not necessarily even any good — there is probably a reason they are out of print, then it seems likey that the training corpus is mostly complete. This leads to the second point, which is actually the first in the title, stagnation. If we ingest all the training data that is useful, and all the human output going forward is LLM produced or LLM affected (the 10%), then the content it outputs stagnates. Writing and language becomes same, and it does not change.
The pre-2022 content, the “low background steel”, becomes older over time, and that’s the truth. In fairness if everything is stagnant and homogenous as I suggest then perhaps that does not matter. But presumably some things still happen in the world, and we feed those things, the news of the world, into the LLMs. Does is matter if the news fed into the LLM training post-2022 is all LLM generated? Model collapse, if true, decreases model accuracy. So if we’ve got news happening in 2035, for example the USA going to war in the middle east for some as-yet-unknown reason, and the news reports are all LLM generated, then does the LLM summary of those events suffer in accuracy versus a news report of one of the other times the USA went to war in the middle east before 2022?
I’ve taken language and writing as an example, but I believe this extends to culture more broadly (language is culture), and to technology. And I think it would be a bad thing.