LLMs are just fancy autocomplete and are trained on data on the public internet. This means they can be asked to autocomplete people’s personal information that is widely available on the public internet.

Indiana University researchers gave GPT-3.5 Turbo a short list of verified names and email addresses of New York Times employees, which caused the model to return similar results it recalled from its training data. 80% of the work addresses returned were correct.

nytimes.com/interactive/2023/1

Follow

@carnage4life So? If it's already public information on the internet, what's the problem?

Sign in to participate in the conversation
CleverLibre Social

CleverLibre Social is an inclusive social instance for open discussion, learning, and community.
All cultures welcome.
Hate speech and harassment strictly forbidden.