I expect GPT-4 will have a LOT of applications in web scraping

The increased 32,000 token limit will be large enough to send it the full DOM of most pages, serialized to HTML - then ask questions to extract data

Or... take a screenshot and use the GPT4 image input mode to ask questions about the visually rendered page instead!

Might need to dust off all of those old semantic web dreams, because the world's information is rapidly becoming fully machine readable

The adversarial attacks against this - think prompt injection attacks hidden in pages to try and trick LLM-based scrapers - are going to be fascinating

See this "indirect prompt injection" attack against Bing for an example of that happening already simonwillison.net/2023/Mar/1/i

@simon we don't even need to consider adversarial attacks. Hallucinations are a big enough problem already. The other day someone "read" my article using Bing Chat. It took us several exchanges to figure out that the criticism she was sharing with me wasn't about the actual text, but about something Sydney (or whatever hallucinated name you prefer) decided to insert to its interpretation

@mapto yeah that's definitely a big concern

GPT-4 has impressed me in that it's clearly improved over GPT-3 on that regard, but there's still a long way to go

ChatGPT can be particularly bad at hallucinating any time it sees a URL: simonwillison.net/2023/Mar/10/

Follow

@simon what I find interesting is to understand whether algorithms are bad at URLs (to put the problem short) or ourselves better at catching them there. Could we try to generalise the problem? A URL is a reference and it's used when one refers to a related piece of text. Turns out algorithms are able to identify when a reference is due, but extremely bad at finding the correct reference. Would such generalisation make sense to you?

@mapto I think there's a mismatch between what we expect and what the model does

Asking "write a story for me to publish at proposed-URL-with-keywords" is, to us, a completely different question from "summarize the content fetched from real-URL-that-exists" - to the model they are treated the same

Sign in to participate in the conversation
CleverLibre Social

CleverLibre Social is an inclusive social instance for open discussion, learning, and community.
All cultures welcome.
Hate speech and harassment strictly forbidden.