OpenAI insists training AI models on copyrighted data is fair use.
OpenAI has publicly responded to a copyright lawsuit by The New York Times, calling the case “without merit” and saying it still hoped for a partnership with the media outlet.
In a blog post, OpenAI said the Times “is not telling the full story.” It took particular issue with claims that its ChatGPT AI tool reproduced Times stories verbatim, arguing that the Times had manipulated prompts to include regurgitated excerpts of articles. “Even when using such prompts, our models don’t typically behave the way The New York Times insinuates, which suggests they either instructed the model to regurgitate or cherry-picked their examples from many attempts,” OpenAI said.
OpenAI claims it’s attempted to reduce regurgitation from its large language models and that the Times refused to share examples of this reproduction before filing the lawsuit. It said the verbatim examples “appear to be from year-old articles that have proliferated on multiple third-party websites.” The company did admit that it took down a ChatGPT feature, called Browse, that unintentionally reproduced content.
One thing that seems dumb about the NYT case that I haven't seen much talk about is that they argue that ChatGPT is a competitor and it's use of copyrighted work will take away NYTs business. This is one of the elements they need on their side to counter OpenAIs fiar use defense. But it just strikes me as dumb on its face. You go to the NYT to find out what's happening right now, in the present. You don't go to the NYT to find general information about the past or fixed concepts. You use ChatGPT the opposite way, it can tell you about the past (accuracy aside) and it can tell you about general concepts, but it can't tell you about what's going on in the present (except by doing a web search, which my understanding is not a part of this lawsuit). I feel pretty confident in saying that there's not one human on earth that was a regular new York times reader who said "well i don't need this anymore since now I have ChatGPT". The use cases just do not overlap at all.
it can’t tell you about what’s going on in the present (except by doing a web search, which my understanding is not a part of this lawsuit)
It's absolutely part of the lawsuit. NYT just isn't emphasising it because they know OpenAI is perfectly within their rights to do web searches and bringing it up would weaken NYT's case.
ChatGPT with web search is really good at telling you what's on right now. It won't summarise NYT articles, because NYT has blocked it with robots.txt, but it will summarise other news organisations that cover the same facts.
The fundamental issue is news and facts are not protected by copyright... and organisations like the NYT take advantage of that all the time by immediately plagiarising and re-writing/publishing stories broken by thousands of other news organisations. This really is the pot calling the kettle black.
When NYT loses this case, and I think they probably will, there's a good chance OpenAI will stop checking robots.txt files.