I honestly do not know whether anything I have ever published has ended up training a chatbot somewhere, and until this week I had never actually looked into how a person would even find out.
That is not a comfortable thing to admit as someone who writes for a living. It turns out I am not alone in not knowing, and a lawsuit filed this summer is trying to force an answer for a much bigger group of publishers than just me.
What the lawsuit actually claims
On June 24, 2026, a coalition of 34 newspaper companies, representing nearly 400 local publications across 33 states, filed suit against OpenAI and Microsoft in federal court in New York. The Arkansas Democrat-Gazette and its parent company, WEHCO, are named plaintiffs. The complaint alleges that the two companies systematically scraped publisher websites, copied articles without permission, stripped out copyright management information, and fed the material into the datasets that trained ChatGPT and Copilot. The legal claims rest on the federal Copyright Act and the Digital Millennium Copyright Act, and the newspapers are seeking statutory damages along with an injunction that would force the removal of their work from the underlying models and training data, not just a promise to stop scraping going forward.
Part of the newspapers’ evidence is OpenAI’s own words. In written evidence the company submitted to the UK House of Lords in late 2023, OpenAI stated plainly that “it would be impossible to train today’s leading AI models without using copyrighted materials.” The complaint frames that admission as the whole case in miniature: a company that argues copyright is essential to its own product while treating other people’s copyrighted work as free raw material. As the filing itself puts it, the publishers are trying “to hold Defendants to the same standard they insist upon for themselves.”
Why this matters even if your blog is not one of the 400
Most independent bloggers and small digital publishers are never going to be plaintiffs in a case like this one. Joining a coalition lawsuit takes lawyers, resources, and a body of published work large enough to be worth the fight. But the legal theory being tested here does not only apply to newspapers with decades of archives. If a court agrees that scraping a publisher’s site to train a commercial model without permission or payment is copyright infringement, that finding does not stay contained to the plaintiffs named in this one complaint. It becomes a data point every smaller publisher can point to.
A quieter reason this case matters sits beyond whatever legal precedent it eventually sets. Most independent writers have never had the option WEHCO and its co-plaintiffs are exercising, the option to say no and be taken seriously. A newspaper chain with decades of archives and in-house counsel can credibly threaten litigation. A single blogger, or even a small network of them, generally cannot, which means the practical leverage over how AI companies treat smaller publishers’ work has always sat with whoever could afford the lawsuit, not with whoever actually wrote the content.
A quick check on where you actually stand
1. Look at your own robots.txt file
Most AI crawlers, including OpenAI’s, are supposed to respect a robots.txt file that blocks them, at least when a site owner has actually configured one to do so. Check whether your site’s file currently blocks known AI crawler user agents, or whether it was never updated past whatever your CMS shipped with by default. A lot of sites are wide open simply because nobody ever touched the setting.
2. Search for your own content inside publicly available training-data trackers
Several independent researchers and outlets have built lookup tools that let a site owner check whether their domain appears in commonly used AI training datasets, including the kind of web-scrape datasets named in this lawsuit’s complaint. Running your own domain through one of these is a five-minute task that at least tells you whether you are asking a hypothetical question or a concrete one.
3. Decide whether blocking or licensing is actually the goal
Blocking every AI crawler outright is not automatically the right call for every publisher. Some are actively negotiating licensing deals instead, trading access for payment rather than trying to keep every crawler out. Figure out which outcome you actually want before you spend time implementing either one. A blogger who depends on search traffic to survive has a very different calculation than one whose income comes mostly from a newsletter or direct sales, since blocking a crawler can carry search-visibility tradeoffs of its own that have nothing to do with AI training specifically.
4. Keep records now, even if you take no other action
Whatever you decide to do about crawling going forward, save dated screenshots or archived copies of your own published work and your site’s historical robots.txt settings. If this legal theory keeps winning, a record of what you published and when, and what you did or did not authorize, is the kind of thing that turns a vague sense of being wronged into something a lawyer could actually work with later.
Final thoughts
I am not a lawyer, and nothing above is a legal strategy, just the plain, practical version of what I would want to know about my own site before deciding whether any of this is worth acting on.
Whatever happens to this particular case, the underlying question, whether the internet’s smaller publishers get paid or just get scraped, is not going away because the newspapers with the resources to sue are the only ones asking it out loud. I would rather know the answer for my own site now than find out the hard way later that I never bothered to look.
Related Stories from The Blog Herald
- New York’s SAFE Act now blocks algorithmic feeds and overnight notifications for any user a platform believes is under 18, and the rules Governor Hochul’s office finalized this summer could quietly reshape how blogs and digital publishers reach teen readers who discover content through social recommendations rather than search.
- People who learned to write in longhand often think differently on paper — and many quietly miss the slowness of it
- New York’s new AI advertising law took effect June 9 and requires a “conspicuous” disclosure any time a synthetic AI performer appears in an ad, with fines of $1,000 for a first violation and $5,000 for each one after that, and while the law spares publishers who merely host the ad, any blog or brand actually producing AI-generated spokespeople for a New York audience is squarely on the hook.
