“AI researchers found that widely used safety training techniques failed to remove malicious behavior from large language models — and one technique even backfired, teaching the AI to recognize its triggers and better hide its bad behavior from the researchers.”

Poisoned AI went rogue during training and couldn't be taught to behave again in 'legitimately scary' study
livescience.com/technology/art

Research pre-print:
arxiv.org/abs/2401.05566

@kegill Now that's interesting! If regular folks were able to generate and host "AI poison pills" on their websites, the AI companies would have far greater incentive to obey their robots.txt. Right now, it's just a polite request, but if there could be a poison pill behind it, it's more like a minefield sign.

Follow

@kegill Right, but as it says in the post you quoted, that's "not about breaking models", it's about protecting the work in question.

By "poison pill", I'm referring to tech or content that would intentionally break models and cause more anomalous output.

@LouisIngenthron

Louis, there are two tools in that article. One is protective. The other is designed to BREAK any content derived from the stolen work.

Sign in to participate in the conversation
CleverLibre Social

CleverLibre Social is an inclusive social instance for open discussion, learning, and community.
All cultures welcome.
Hate speech and harassment strictly forbidden.