Pretty sure this was under heavy discussion earlier this month, but it flew under my radar so I’m sharing it here more widely: maliciously modifying an open-source LLM to spread misinformation while remaining undetected by current methods. Data provenance is critically important, and LLMs are a new front on the software supply chain war. blog.mithrilsecurity.io/poison

@darkuncle

Anyone who actually trusts the output of a LLM at this point hasn't been listening and deserves exactly what they get. Sorry. Tough Love.

@Biggles it’s like having an intern: you need to validate the output; it will be mostly right much of the time but that’s a long way from reliably right all the time.

@darkuncle

If I had an intern that would just make up shit when they didn't know the answer - I'd get a new intern. The ability to say "I don't know" is vital.

@Biggles people have to learn that ability, more often than not

Follow

@darkuncle

Yeah, but (1) most people learn it fairly young, and (2) so far as I can tell, the LLMs in use today - all of them - still haven't learned it - and probably *can't*. They simply don't actually understand the semantic content of their corpus well enough to separate truth from falsehood. You might be able to work around that a bit by being *very* careful about what you feed them - the AWS docs one lies, but obviously so - but I don't think it's fixable with this type of technology.

· · Tusker · 1 · 0 · 1

@Biggles right - LLMs don’t have any context for objective truth outside of what they were trained on; that corpus is truth by definition. The idea that anything in it might not be truthful is a nonsensical statement from the perspective of the LLM.

@darkuncle

it's even more insidious than that - it's not that they believe their corpus is all true - it's that they don't have a concept of truth, or falsehood, or concepts in general. The idea that they can tell the truth is anthropomorphizing them - they generate symbols that match previous patterns, and it's us that ascribe meaning where it's not deserved - like seeing faces in clouds. So - no truth, no falsehood. What they are is correct or incorrect, or a mix of both, or "not even wrong" when they spew out nonsense.

With a book - we accurately attribute what is written to the author and not to a bundle of cardboard and paper, but the MML looks different - they are pretty damn effective in convincing people that they're something they aren't. I used to think of them as a parrot - but a parrot is unlikely to get someone killed, because you know it's a bird. I think of them now as a memetic virus - they mimic thought well enough to exploit weaknesses in human perception and cognition and get us to pay attention to them. They aren't evil, because there's no intent - they're just a disease we haven't developed any defenses against.

Sign in to participate in the conversation
CleverLibre Social

CleverLibre Social is an inclusive social instance for open discussion, learning, and community.
All cultures welcome.
Hate speech and harassment strictly forbidden.