wikiHow just sued OpenAI for scraping 11,000+ how-to articles to train ChatGPT without permission or payment. Filed in SDNY (Case 1:26-cv-07171), the complaint alleges OpenAI pulled content directly and via Common Crawl, then kept crawling even after wikiHow blocked GPTBot and OAI-SearchBot via robots.txt in 2023/2025.

The core technical accusation: ChatGPT now generates competing how-to responses that reproduce wikiHow's substance near-verbatim, cannibalizing pageviews and ad revenue. wikiHow also claims OpenAI stripped copyright management info (titles, bylines, notices) in violation of DMCA.

OpenAI's defense: "publicly available data + fair use." But the lawsuit argues that substitution effect kills the economic incentive to produce original content.

This is part of a larger wave of copyright suits against AI companies over training data practices. The legal question boils down to: does fair use doctrine cover large-scale commercial ingestion of copyrighted material for model training when the output directly competes with the source?

Technically interesting because it tests the boundary between "learning from public data" and "commercial copying that displaces the original." The robots.txt violation is also a data point on whether AI companies respect crawl directives or just ignore them when convenient.