The Internet Ouroboros
AI ate the human internet. Now it is coming back for the fossils.
Somewhere, some way, somebody out there is buying an awful lot of old books.
Not Hemingway first editions or illuminated, gold-leaf manuscripts hidden in ancient monasteries. No, much stranger things you probably don’t ever think about.
A 1995 guide to metropolitan Denver. A manual for WordPerfect harking back to 1991. Forgotten technical books, local histories from villages you’ve never heard of, academic texts, obscure nonfiction from decades ago, drier than the Gobi Desert. The kind of books that survive for decades because nobody wants them badly enough to buy them and nobody hates them enough to throw them away.
Well, in case you’ve not been paying attention, somebody does indeed want them now.
At one used bookstore in Houston, the manager told The Atlantic that 95 of the previous 100 books sold through one online platform had gone to the same buyer. Other booksellers have reported similarly strange bursts of demand, including orders where shipping cost more than the books themselves.
Over in Europe, a Dutch bookseller thought a request for 3,000 copies was "spam or phishing."
Naturally, the internet has gone and decided that AI is eating the libraries.
That isn’t exactly paranoia speaking. AI companies like Suno and Udio are ravenously dining on the musical output of civilization as we speak.
We already know that at least one AI company has bought physical books at enormous scale for AI training purposes.
For some reason, Anthropic, the safety-minded frontier AI lab behind Claude, called its operation “Project Panama.”

The company bought millions of new and used books, removed their bindings, cut the pages apart, scanned them, and shredded the originals. Internal material later exposed through litigation contained a sentence that seems scientifically engineered in a laboratory to infuriate bibliophiles:
“Project Panama is our effort to destructively scan all the books in the world.”

Anthropic has reportedly already spent tens of millions of dollars doing it.
But I think most of the outrage is misplaced. Looking at the wrong part entirely…
The Wrong Thing to Be Angry About
There is something genuinely unpleasant about watching a book being burned or slashed or shredded or torn apart.
Kids forced to tear through a summer reading list and prepare book reports for the first day of school may disagree.
That visceral reaction to maiming books is not entirely sentimental nonsense, though. Books are obviously more than mere containers for some text. People write in them, spill things on them, gift them to each other, drag them through moves and messy relationships and entire periods of their lives. Sometimes a specific copy matters even though the words inside it are identical to a million other books out in the wild.
There are also books where destroying the object destroys something that a scan cannot preserve.
An exquisitely rare edition, a heavily annotated copy, a local history with almost no surviving copies, some obscurely strange technical document that was printed once in 1976 and then lost to the forgotten depths of time.
If those disappear into a scanner and a private dataset, something has actually been lost.
But this is where, in my humble opinion, the conversation around Anthropic starts getting ever so slightly confused.
Most books are not unique cultural artifacts with outsized historical significance. Libraries remove books from their collections all the time. Publishers pulp inventory. People throw badly-battered books away during cross-country moves. Entire sets of outdated encyclopedias have been sent migrating toward recycling centers since Wikipedia arrived, like dusty, musky lambs to slaughter.
The destruction for sure matters. That being said, the fact that a book was physically destroyed does not automatically make it an act of cultural vandalism.
Anthropic has now publicly stated that it was not deliberately buying rare or antiquarian books for Project Panama. The more recent buying spree is much harder to explain.
Some sellers have reported unusually obscure material disappearing in bulk, but nobody has shown that Anthropic, or even AI companies generally, are behind all of those purchases.
That distinction does indeed matter because there is a tendency with stories like this to let the most disturbing, conspiratorial, tin-foil hat explanation become the accepted one before anybody has actually proven it.
And you know what?
It is exceedingly easy to see how we arrived there, given how fashionable hating everything remotely AI-related has become in 2026 A.D.
Still, we already know enough to ask the more interesting question. But there are more curious things lurking in the tech shadows. More interesting and uncomfortable questions to be asked, and more precarious truths to dislodge…
Why would anyone building frontier AI want a thirty-year-old computer manual in the first place?
The outrage tends to stop at the part where books get violently dismembered. That’s where I think the story really starts to get fascinating.
The Paper Guillotine
Part of the answer is embarrassingly mundane.
Copyright law has made the physical book unusually useful.
Anthropic had previously downloaded millions of books from pirate libraries, which is the polite legal way of describing an operation that looks an awful lot like stealing millions of books.
Judge William Alsup distinguished between using lawfully acquired books for training and keeping pirated copies in a general-purpose internal library, and Anthropic eventually agreed to a $1.5 billion “slap on the wrist” settlement over the piracy claims. Not too shabby for a company valued at $965 billion as of May 2026.
Project Panama followed a different sort of logic altogether. Buy the physical book legally, scan it, destroy the original during the process, and keep the digital copy in an ever-growing digital archive.
There is something wonderfully stupid and unsettlingly human about the fact that this can somehow make more legal sense than downloading the same damn text from the internet.
That helps explain the bizarre mechanics of Project Panama, yet it still doesn’t explain why the books were valuable enough to bother acquiring in the first place.
For a brief period, the bleeding edge of artificial intelligence looked less like an Asimov novel and more like a slaughterhouse for literature, where books go to die.
The broader argument about AI training data is obviously much bigger than Anthropic alone. OpenAI’s GPT and other frontier models were built from staggering amounts of human-created material, and the circumstances under which that material entered training datasets vary wildly.
Some was licensed legally in accordance with the law. Some was scraped from public websites. Some creators had no idea their work was being used at all. In other cases, pirated copies were absolutely involved.
You are free to choose your preferred moral vocabulary for all of that. Theft, scraping, learning, stealing, training, gathering inspiration, appropriation, fair use, industrial-scale borrowing without consent.
Courts and regulators are still duking it out, trying to differentiate and settle on which descriptions have legal force.
I am not especially interested in trying to settle that fight here. At least not here in this particular piece.
What interests me more at this moment is what happened after the models had eaten so much of the human internet.
Because eventually, somewhere along the line, they started contributing to it too.
The Internet Ouroboros
Books were attractive training material long before the current obsession with synthetic data, and long before most people had any reason to think about AI training corpora at all. Before people even knew what AI was or why they should be afraid of it.
They are methodically edited, generally more coherent than random web pages, and full of hidden wisdom that seemingly never made the transition cleanly onto the open internet.
A manual written for a specialized industry in 1989 might contain information that barely exists anywhere in any fashion online. A forgotten academic book might cover some obscure subject with more depth than twenty years of search-engine-optimized summaries ever managed.
I can say from experience that a history degree leaves you with at least one durable skill: finding books nobody else knew were there. If you have ever done any archival research, you know how much of human history never made it online. Sometimes the thing you need is still sitting in a gloomy room inside a book nobody has opened since 1986.
Of course, AI labs are more than capable of combing through old websites, newspaper archives, and already-digitized books too.
Paper matters because plenty of useful material never made it online, because published editions come with dates and provenance, because books tend to offer denser and more coherent writing than the average webpage, and, thanks to the strange legal incentives above, sometimes buying the physical object is simply the cleanest way to acquire and extract the text.
Older books now have another advantage that would have sounded ridiculous a few years ago.
Their age tells you something about who, or what, forged them into existence.
If a book was printed in 1981, Claude did not help write it. ChatGPT did not smooth out the prose. Nobody was able to ask Gemini to turn chapter three into something “more engaging and digestible.” Whatever its flaws, the thing came from an overwhelmingly human-authored, pre-generative-AI information environment.
Once upon a time, that distinction used to be so obvious that nobody even needed a word for it.
The early frontier models arrived at an unusually convenient point in history.
Humans had spent decades producing enormous quantities of material online before machines became capable of producing convincing language at comparable scale. Blogs, newspapers, forums, code repositories, research papers, recipes, technical documentation, arguments, reviews, obscure hobbyist sites, bad poetry, good poetry, and millions of people patiently explaining the same problems to strangers.
It was a gigantic accidental dataset, overwhelmingly created by humans. By people like you and me.
Then the frontier models came and absorbed it all, which in turn made something rather funky happen.
That overwhelmingly human-only information environment began changing quickly and dramatically.
Generative AI started producing articles, comments, product descriptions, summaries, emails, code, forum posts, books, and just about everything and anything else people could persuade it to make.
Some of that material is surprisingly useful in a pinch. Some of it is outright sloppy garbage. Most importantly, more and more of it now sits neatly nestled beside human-created material without a reliable or trustworthy way of distinguishing between them.
I won’t pretend to be smart enough to suggest that synthetic data is inherently bad for AI. Far more brilliant machine learning experts are chipping away at that problem already.
That is one of the places where the public version of this argument becomes complicated.
You might not know it, but researchers deliberately use synthetic data all the time. It can be filtered, verified, combined with human data, and used to train models very effectively.
The problem appears when that loop is uncontrolled.
Research on recursive training has shown that when models repeatedly learn from generated material without preserving enough of the original distribution, the output can gradually narrow and distort.
There is no evidence that Anthropic launched Project Panama because someone over in San Francisco was panicking about model collapse. The documented attraction of books is much, much broader than that: quality, depth, access, and huge quantities of useful text.
But… The provenance problem is where I think the story gets stranger.
Rare information starts falling away. Certain mistakes become more common. The model begins learning from an increasingly degraded reflection of what earlier models learned from humans. Almost like a photocopy of a photocopy of a photocopy ad infinitum.

The phrase that became popular for this was ‘model collapse’, although I personally think the provenance problem is more interesting than the dramatic name.
If a future training corpus contains millions of pages that were generated by previous models, how do you reliably know what originally came from us supposedly learned, upright simians?
And once models are rewriting human work, summarizing previous model output, generating articles from generated press releases, and filling search results with content derived from other content, the ancestry gets chaotic, very, very quickly.
That is the loop I’ve become obsessed with recently. A rabbit hole I haven’t gotten the slightest hope of crawling out of anytime soon.
AI learned from the human internet, then slowly became part of the very same internet itself. The environment that trained the first generation no longer exists in the same form for the next one. And it never will, ever again.
And the older material also starts looking different in this post-machine-learning world.
Libraries, archived websites, old magazines, dead forums, dusty books. None of these became better sources of information overnight. What changed is the confidence we can have about where they came from.
The snake has begun eating its own tail!
I Found a Fossil in Google Docs
While researching this, I started looking through some of my old files.
I found a Google Doc called The Lizard, dated May 19, 2018.

I wrote it while I was living in Vietnam. It is about something I experienced there, written in whatever style and state of mind I happened to have eight years ago. I had largely forgotten it existed. I’ll even go so far as to say that it’s perhaps a painful reminder that I hate my own writing and haven’t gotten much better in a decade.
There is literally nothing especially important about this particular short story. It is not some lost masterpiece from my youth. Mostly it is an old piece of writing sitting in an archive.
But reading it now, it struck me that the document has acquired a characteristic I had absolutely no reason to think about in 2018.
I know beyond a shadow of a doubt who wrote it.
It would hardly satisfy a forensic data auditor, but I don't need one to tell me where this thing came from. I remember writing it.
There is no ambiguity about whether a model generated the first draft, fixed the grammar, rewrote half the paragraphs, or quietly inserted itself into the process somewhere along the way. I wrote it years before any of those tools were available to the world.
The document may not have changed, but the world around it certainly has gotten weirder and wilder.
That makes it something like a fossil or a relic of time.
Not in the sense that it is precious or worth hanging in the halls of some magnificent museum, but because it comes from a layer of the information environment formed under conditions that won’t ever exist again.
There must be trillions of things like it too…
Old blogs that have not been updated in fifteen years. Abandoned message boards. University essays. Email archives. Newspaper databases. Flickr accounts. Personal websites maintained by people on some sort of spectrum who knew far too much about orchids or model trains or tropical fish. Photographs sitting on discontinued hard drives. Old company documents nobody cared to delete.
Most of this material is individually worthless.
Collectively, though, it records a period when human authorship was still that default assumption.
That is yet another thing that makes those old physical books interesting too. They belong to the same historical layer.
The Oil Under the Library
The comparison I keep coming back to is oil, although this stretch of a conceit breaks if you take it too literally.
Information is very obviously not a consumable resource in the same way. Training a model on a book does not burn the words. The text can be copied endlessly, and humans can keep producing new material every day.
What is finite, however, is the historical period itself.
There is unfortunately a closed corpus of information created before generative AI became an ordinary participant in our ecosystem.
We can preserve that material and duplicate it as much as we damn well like, but we cannot produce another twenty years of internet magically out of thin air from the pre-GPT times.
That is where my silly, little geological analogy becomes useful.
Oil fields matter because enormous deposits formed under conditions that cannot simply be recreated whenever we need more. We simply don’t have the time.
The individual organisms that eventually became part of the deposit are borderline irrelevant. But the layer matters.
The pre-generative internet is beginning to look like an informational version of that.
Researchers at Epoch AI have tried to estimate how much high-quality public human text exists for AI training. The number is necessarily rough, but their work points toward something critical. Frontier AI has been consuming the easily accessible stock of high-quality human-generated text much faster than humans can create equivalent material, and the deficit grows.
There are really two problems colliding here. AI labs want enormous amounts of useful human material, and the easiest high-quality supply does not expand at anything like the rate compute has. At the same time, the newer internet is becoming more complicated to trace back to a human source. Old material happens to be unusually attractive on both fronts.
That does not mean models are about to run out of things to learn from, though. There is private data, video, audio, scientific measurements, interactions with the physical world, newly created human material, synthetic data, and training techniques that extract more value from existing datasets.
Still, the first frontier models received a strange one-time inheritance from their forebears.
They arrived after decades of public human writing had accumulated and before generative AI had become a major contributor to that same environment.
No future model gets that same cushy deal.
The archive itself is not going anywhere. We can duplicate it endlessly forever. But the historical window in which it was produced has closed. Every new year of internet now arrives under different conditions from 2004, 1997, or 1986.
A geological layer can be drilled again. It cannot be freshly deposited on command.
This is why the awkward, old material suddenly becomes intriguing. A forgotten forum from 2004 might be full of nonsense, but the nonsense has unusually clear provenance. The same is true of a printed book from 1978 or some embarrassing short story I wrote in Vietnam when I was trying to be artistic and profound.
Provenance is becoming the harder part to pin down. We have more information than we know what to do with, but we are less certain than ever about who made it, when it was created, and how many machines may have touched that data along its journey.
Certified Organic Humans
We are already starting to build products around that problem.
The Authors Guild now offers a Human Authored certification for books. Writers can register qualifying works and display a mark indicating that a human indeed wrote the text, with certain allowances for research, grammar, and spelling, and allowances for limited editing tools.

The startlingly strange part of all this is how quickly it already started feeling ordinary.
For almost the entire history of publishing, putting “written by a human” on the cover of a book would have sounded like certifying that it contained words.
One would have been relentlessly mocked.
It’s not like there was anybody else around to write it.
Now the assumption has become the opposite. Something publishers and authors may want to prove.
It reminds me a little of food labeling. The underlying items did not suddenly change, but the environment around them did, and eventually the absence of something becomes a selling point.
We are already drifting toward the language of food labels: no preservatives, no artificial colors, human-authored.
That sounds ridiculous until you scroll through Amazon and realize why those absurd labels exist in the first place.
The Machine Writes Back
A recent study looked at more than 14,000 self-published genre-fiction titles sold on Amazon between 2023 and 2026.
Those researchers found a growing presence of books containing substantial amounts of AI-generated text in them.
Those books were not necessarily dominating sales, which does matter. This is not a story about readers suddenly deciding AI fiction is better than human fiction. Maybe I’ll be writing that essay in a year’s time.
The more peculiar effect was volume.
Generative AI made it dramatically easier to pump out something resembling a book at scale.
The researchers found a market where the number of titles competing for attention grew much faster than the amount of money readers were spending.
That changes a creative market even if the average AI-generated book is mediocre (sorry, but they usually are).
It also completes the loop in a way that would have sounded absurd only a few years ago.
Human books helped train the models. The models became capable of writing books themselves. Those books entered the same digital markets as human books, which forced human authors to start proving that their books were certifiably human-made.
At roughly the same time, old physical books became attractive to AI companies partly because they come from a simpler era when none of this ambiguity even existed.
This is why I find it difficult to make Anthropic the villain of this specific story. And trust me, there are plenty of other things to be outraged about.
There are legitimate questions about how AI companies acquired training material and whether creators should have been compensated. There are also very reasonable concerns about obscure or rare works disappearing into private corporate datasets, and vanishing altogether.
But the image of Anthropic chopping up books is almost too convenient. It gives the story a physical act that looks obviously wrong, which makes it easy to stop thinking there.
The more consequential change is harder to take a snapshot of.
For decades, digitization solved an obvious problem. Physical information was difficult to search, harder to copy, and often impossible to access. We spent years converting books, photographs, newspapers, and almost everything else we could find into raw, scannable data.
Then the digital world became capable of generating information on its own.
Once the digital world could manufacture material at essentially unlimited scale, provenance began carrying a hell of a lot more weight.
There is more text than anyone could ever read. More images than anyone could ever see. More material can be generated in an afternoon than a person could consume in a thousand bajillion lifetimes.
The tricky part now is sleuthing around to figure out who produced that content. Where it may have come from and for which purpose. Oh, and how many generations of machines sit between you and whatever human experience originally existed at its core.
Which brings us back to the books nobody wanted.
Somewhere there is probably a computer manual from the early 1990s sitting on a shelf because nobody ever bothered to throw it away. For years, the internet made that book look increasingly obsolete.
Now its obsolescence may be exactly what makes it useful. It turns out sitting still for thirty years has its advantages.