How bad are the thousands of new stochastically-generated websites?
Last night I wanted to roast some hazelnuts, and I could not remember the temperature I used last time. So I searched on DuckDuckGo. Every website that I could find was machine-generated with different temps listed. One site had three separate methods listed that were essentially differently worded versions of the same thing. With different temperatures.

So I pulled my copy of Rodale’s Basic Natural Foods Cookbook off the shelf and looked it up there.

I think it may be time to download an archive copy of the 2022 Wikipedia before we lose all of our reference material. It was nice having all the world’s knowledge at my fingertips for a couple of decades, but that time seems to be past.

@grumpasaurus @bhawthorne Oh man, there is a website that I can't seem to find right now, but it's essentially a bunch of recipes in a sort of "raw" format; no ads, no stories, no ratings, just lots of diverse recipes.

I'm letting you know even though I haven't found it, in case anyone else sees this and goes "Oh, yeah, it's this".

Follow

@b4ux1t3 @grumpasaurus @bhawthorne A while back, I actually trained an LLM to crawl recipe sites, detect pages with recipes, and extract them from their paragraphs of pointless stories, while also rating the page for clarity and brevity. Had 98% accuracy.
Was thinking about creating a search engine that prioritized things people actually wanted instead of just wordiness. But I'm not friends with any rich VCs, so 🤷‍♂️

@LouisIngenthron @grumpasaurus @b4ux1t3 @bhawthorne psst, try Kagi...

Also incidentally they also have a LLM summarizer in their web extension...

I genuinely like it better than using Google, not just because getting away from the monopoly or the excessive inserts... but because I definitely am finding better results faster (and when I find a site like this, I can tell it to never show it to me in a search result ever again)

@shiri @LouisIngenthron @grumpasaurus @b4ux1t3 The bad news is that Kagi delivered the same top 3 results as Google and DDG. This confirms to me the problem is not the search engine. It is the content. It’s all SEO, all the time, and now being contaminated by generated content. On the other hand, I may be able to set a default date range to always cut off any results after 2022-11-30.

@bhawthorne @shiri @LouisIngenthron @b4ux1t3 2022-11-30 has a rift in time and space that prevents time travel before that date.

@bhawthorne @grumpasaurus @b4ux1t3 @LouisIngenthron fair, though I find the ability to tune your results to be helpful in adjusting for this (means you'll only see that junk site once).

You can even take advantage of their stats page to pre-tune: kagi.com/stats?stat=leaderboar…

Given that SEO takes time and domains are limited resource that cost money, tuning out crap results as they come up should pretty easily keep up.

Sign in to participate in the conversation
CleverLibre Social

CleverLibre Social is an inclusive social instance for open discussion, learning, and community.
All cultures welcome.
Hate speech and harassment strictly forbidden.