I see the punditry commenting on the NYT lawsuit is is pretending that current era LLMs aren’t specifically designed to memorise and regurgitate answers in standardised tests as a marketing ploy.

(Memorisation and verbatim copying is pretty much required to pass, say, a bar exam)

You can design an LLM to mimimise memorisation and verbatim copying. OpenAI or Google have not done this. Systems who minimise memorisation by design, perform well in other benchmarks, tend to flop hard at human-oriented exams which are designed to test that a human has spent a lot of time reading the right books

Instead, because “our AI can be a lawyer, it passes the bar exam” is such a boon for marketing, they’ve optimised their systems to do the oppisite, despite the inherent plagiarism risk

So, IANAL, so I don’t know enough to assess the overall legal risk this exposes you to (although see all the lawsuits) which is going to vary depending on whether you’re doing it en masse to replace a search engine, using it for code completion, or to write text for publication

What I do know is the copying it does would’ve been enough to get you fired if you’d have done it yourself before the era of LLMs

But all of a sudden its fine now that it’s done by a chatbot that won’t ever join a union

Follow

@baldur At least for code completion, that's simply not true. There's an old adage in software engineering: Novice programmers take code from StackOverflow. Experienced programmers take the *right* code from StackOverflow.

@LouisIngenthron Copying unattributed GPLed code into a proprietary project will absolutely get you fired.

Sign in to participate in the conversation
CleverLibre Social

CleverLibre Social is an inclusive social instance for open discussion, learning, and community.
All cultures welcome.
Hate speech and harassment strictly forbidden.