“AI researchers found that widely used safety training techniques failed to remove malicious behavior from large language models — and one technique even backfired, teaching the AI to recognize its triggers and better hide its bad behavior from the researchers.”
Poisoned AI went rogue during training and couldn't be taught to behave again in 'legitimately scary' study
https://www.livescience.com/technology/artificial-intelligence/legitimately-scary-anthropic-ai-poisoned-rogue-evil-couldnt-be-taught-how-to-behave-again
Research pre-print:
https://arxiv.org/abs/2401.05566
@kegill Now that's interesting! If regular folks were able to generate and host "AI poison pills" on their websites, the AI companies would have far greater incentive to obey their robots.txt. Right now, it's just a polite request, but if there could be a poison pill behind it, it's more like a minefield sign.
@LouisIngenthron
Louis, there are two tools in that article. One is protective. The other is designed to BREAK any content derived from the stolen work.