Authors Are Furious After Finding Their Works on List of Books Used To Train AI

stopthatgirl7@kbin.social · 1 year ago

Authors Are Furious After Finding Their Works on List of Books Used To Train AI

👁️👄👁️@lemm.ee · 1 year ago

It literally shares passages verbatim

BetaDoggo_@lemmy.world · 1 year ago

So does any site that quotes the book. Just being trained on a work doesn’t give the model the ability to cite it word for word. For most of the books in this set you wouldn’t even be able to get a single accurate quote out of most models. The models gain the ability to cite passages from training on other sources citing these same passages.

lloram239@feddit.de · edit-2 1 year ago

It shares popular quotes from books, it can’t reproduce arbitrary content from a book. The content needs to be heavily duplicated in the training data to stick around (e.g. from book reviews), and even than half of it might still end up being made up on the spot.

Also request for copyrighted content will be blocked by ChatGPT and just receive the stock “I can’d do that” response anyway.

If you have some damning examples that show the opposite, show them.

BURN@lemmy.world · 1 year ago

Being blocked by ChatGPT just means that the interaction layer you see doesn’t show the output, not that the output wasn’t generated.

Everything you see that’s public facing and interfacing with an AI is an extreme filtering layer for what is output. There’s tons of checks that happen to ensure that they don’t output illegal content or any of a million other undesirable things.

👁️👄👁️@lemm.ee · 1 year ago

I’m too lazy and care too little but you can basically get it to roleplay as a book expert or something and to “remind” you of certain passages. It gets around the filter pretty easily, that’s how jailbreaks work.

Piecemakers@lemmy.world · 1 year ago

That claim is disingenuous at best, and misinformed otherwise.

PsychedSy@sh.itjust.works · 1 year ago

That’s maybe an issue. I mirror speech a lot, though. How large are the passages?