Pretty sure this was under heavy discussion earlier this month, but it flew under my radar so I’m sharing it here more widely: maliciously modifying an open-source LLM to spread misinformation while remaining undetected by current methods. Data provenance is critically important, and LLMs are a new front on the software supply chain war. https://blog.mithrilsecurity.io/poisongpt-how-we-hid-a-lobotomized-llm-on-hugging-face-to-spread-fake-news/
Anyone who actually trusts the output of a LLM at this point hasn't been listening and deserves exactly what they get. Sorry. Tough Love.
@Biggles it’s like having an intern: you need to validate the output; it will be mostly right much of the time but that’s a long way from reliably right all the time.
If I had an intern that would just make up shit when they didn't know the answer - I'd get a new intern. The ability to say "I don't know" is vital.
Yeah, but (1) most people learn it fairly young, and (2) so far as I can tell, the LLMs in use today - all of them - still haven't learned it - and probably *can't*. They simply don't actually understand the semantic content of their corpus well enough to separate truth from falsehood. You might be able to work around that a bit by being *very* careful about what you feed them - the AWS docs one lies, but obviously so - but I don't think it's fixable with this type of technology.
it's even more insidious than that - it's not that they believe their corpus is all true - it's that they don't have a concept of truth, or falsehood, or concepts in general. The idea that they can tell the truth is anthropomorphizing them - they generate symbols that match previous patterns, and it's us that ascribe meaning where it's not deserved - like seeing faces in clouds. So - no truth, no falsehood. What they are is correct or incorrect, or a mix of both, or "not even wrong" when they spew out nonsense.
With a book - we accurately attribute what is written to the author and not to a bundle of cardboard and paper, but the MML looks different - they are pretty damn effective in convincing people that they're something they aren't. I used to think of them as a parrot - but a parrot is unlikely to get someone killed, because you know it's a bird. I think of them now as a memetic virus - they mimic thought well enough to exploit weaknesses in human perception and cognition and get us to pay attention to them. They aren't evil, because there's no intent - they're just a disease we haven't developed any defenses against.
@Biggles right - LLMs don’t have any context for objective truth outside of what they were trained on; that corpus is truth by definition. The idea that anything in it might not be truthful is a nonsensical statement from the perspective of the LLM.