Ask most people who follow self-publishing what generative AI has done to Amazon’s Kindle store, and you’ll get some version of the same answer: the platform is being flooded with machine-written novels, and whoever is producing them is quietly cashing in while human authors watch their sales evaporate. It’s a tidy story, repeated across writer forums and publishing newsletters for the past two years. Cheap words in, real money out.
That version of events isn’t hard to believe. AI detection has gotten good enough to flag machine-written prose at scale, and sightings of obviously synthetic listings, ones with nonsensical author bios or covers that look assembled from the same three templates, have been easy to find in genres like romance and thriller. The assumption that follows feels almost like common sense: more AI books on the virtual shelf should mean more AI dollars landing in someone’s pocket.
A 2026 analysis by Tuhin Chakrabarty and three co-authors at Stony Brook University, Columbia Law School, and the University of Michigan tested that assumption directly instead of taking it on faith. The team ran full-text AI detection across 14,419 self-published genre-fiction ebooks sold on Amazon between January 2023 and March 2026, matching each title to its own daily sales record through the end of June 2026.
They used a detector called Pangram v3.3, which its developers report catches AI-generated text with a false-positive rate of just 0.04 percent, and sorted every book into one of three buckets based on how much of its text the tool flagged as machine-generated: none, light exposure at 25 percent or under, and substantial exposure above that threshold. That threshold matters. A book that used AI for a light copyedit or a few transitional paragraphs lands in a different category than one built mostly from generated text, and the study treats them separately rather than lumping every book that touched a chatbot into one pile.
The results complicate the flooding story rather than confirming it. According to the study, books with substantial AI text made up 20 percent of the catalog analyzed, but captured only 12.1 percent of unit sales and 11.3 percent of revenue. Books with no detected AI text, by contrast, made up 62.9 percent of the catalog and pulled in 72.5 percent of the revenue. As the researchers put it, these books “make up a large share of the catalog but a smaller share of sales.”
That gap is real. It doesn’t cancel out what comes next, though. The same data show AI-heavy books gaining ground over time, winning a growing share of sales and taking more of the scarce top-rank positions once held by books with no detected AI text. A fifth of the catalog earning an eighth of the money looks, at first glance, like proof the flooding strategy is failing. Followed across three years, it starts to look more like a foothold.
The bigger shift might not be about AI-generated books specifically at all. Across the period studied, the researchers found that the cumulative catalog of released titles grew 38.3-fold and the number of books selling in a given quarter grew 19.2-fold, while quarterly revenue grew only 8.9-fold. Thousands of new titles, AI-assisted and human-written alike, poured onto virtual shelves far faster than buyers showed up to pay for them. Revenue per selling book fell across most genres, and books with no AI text lost the most ground specifically in genres where AI-generated titles had spread the furthest, especially where Kindle Unlimited availability was high.
Kindle Unlimited matters here because it changes how a book earns money in the first place. Instead of a single purchase price, KU pays authors per page read from a shared pool, which rewards volume and frequent new releases more than it rewards any single title selling well. A catalog strategy built around publishing often, even with heavily AI-assisted books, can do reasonably well under those incentives even while any individual title underperforms on its own. That helps explain why the erosion for human-written books wasn’t evenly spread. It concentrated in exactly the genres and pricing models most exposed to volume publishing.
A second finding buried in the same dataset speaks less to sales and more to originality. Among top-selling titles, the researchers found that books with substantial AI text drew on more distinctive language already present in existing books than human-written top sellers did, and that this overlap climbed alongside revenue. Books with no detected AI text showed no such pattern. The authors note this bears on the market-effect question at the center of fair-use arguments over copyright, since one legal test for whether a use is fair asks whether it substitutes for the original work in the marketplace.
A browsing reader couldn’t have worked any of this out on their own. Amazon listings don’t disclose whether a book contains AI-generated text, and the researchers needed full-text detection software run against the entire manuscript, not just a skim of the sample pages Amazon shows before purchase, to sort the catalog this precisely. A cover, a blurb, and a few free pages simply don’t carry enough signal either way. That’s part of why the flooding narrative took hold in the first place: the visible evidence, a handful of obviously synthetic listings, was easy to spot, while the actual sales math required a proprietary sales panel and machine detection run across full manuscripts before anyone could see it clearly.
It’s worth sitting with what this data can and can’t tell you. This is one dataset, covering self-published genre fiction specifically rather than all of Amazon’s catalog, and it’s a preprint that hasn’t yet gone through peer review. The findings describe an association between detected AI text and sales performance across a large sample, not a controlled experiment into why any individual reader bought or skipped a given book. Detectors also aren’t flawless. Even at a reported false-positive rate of 0.04 percent, a sample of this size will misclassify some titles in both directions, and genre-fiction ebooks are not the same market as, say, self-published nonfiction or hardcover releases.
What the numbers do support is a narrower claim than the one usually repeated. AI-generated books aren’t quietly draining the same dollars away from human authors, title for title, the way the flooding story implies. They’re doing something slower and, in some ways, harder to counter.
Related Stories from The Blog Herald
- Is blogging still profitable? The bloggers who reach a full-time income wait an average of four years to get there, and the wait is getting longer, not shorter
- Virginia’s mandatory one-hour daily limit on teen social media use was struck down by a federal judge less than two months after it took effect, and NetChoice’s First Amendment win over Virginia’s usage caps could stall similar bills in other statehouses that would have limited how publishers reach younger audiences.
- Podcast advertising rates are climbing and YouTube just doubled its monetization bar – here’s why that might not be a coincidence
By adding volume the market didn’t ask for, they’re making the whole shelf more crowded and slightly less profitable to sell into, whether or not the book sitting next to them was touched by a machine at all.
Related Stories from The Blog Herald
- Is blogging still profitable? The bloggers who reach a full-time income wait an average of four years to get there, and the wait is getting longer, not shorter
- Virginia’s mandatory one-hour daily limit on teen social media use was struck down by a federal judge less than two months after it took effect, and NetChoice’s First Amendment win over Virginia’s usage caps could stall similar bills in other statehouses that would have limited how publishers reach younger audiences.
- Podcast advertising rates are climbing and YouTube just doubled its monetization bar – here’s why that might not be a coincidence
