Curation Algorithms are Important, We Should Control It
The curation algorithms, from search result ranking to timeline ranking to recommendations, may be the most impactful and useful algorithms humans have ever derived. They cut through vast amount of noise to deliver signal relevant to each user. The lack of them on Mastodon and other places is sometimes touted as a feature -- it's not, it's simply lack of an important feature.
That's not to say the algorithms deployed by Twitter et al are perfect. Much has been written about how they are not tuned for the benefit of the users, but the corporations; not tuned for the improvement of the users, but the addiction of them; etc. But a competititor must provide a fix for those problems to measure up to the existing networks.
In a sense, the curation algorithms are bubble creators, because curation is the process of creating bubbles. I hypothesize that there is a trade off among three desirable qualities of these bubbles: effort spent to curate, granularity of the boundary, coverage of the potentials. For the same effort, one must trade off granularity and coverage, as more granularity demands more effort for examining each item.
The importance of the curation algorithms then is to scale up the curation effort available beyond human limitations, leading to improvements in both granularity and coverage. Improvement in coverage is straightforward: Google Books enables searching in unprecedentedly many books in one query. Improvement in granularity is more subtle: recommendations are relevant to complex queries, not simple combination of source and recency. Except, instead of bothering the user to input the complex queries, the corporations choose to infer inputs from user tracking data.
Simply opting-out of tracking is not really taking back control of the curation algorithms from the corporations, as many have suggested. The control of the curation algorithms is in the complex queries. Currently, the corporations sell that control to their customers: advertisers. They can use a set of complex queries to control in what kind of rankings or timelines they want their ads to appear. But the users don't get to control what kind of rankings or timelines they want to view.
Curation algorithms are too important and beneficial to give up. To really take control of a machine learning based curation algorithm, there needs to be a language for both describing its learned parameters to users and for expressing complex queries by users to it.