I watched a neat if somewhat synthetic ML demo based on this paper - https://arxiv.org/pdf/2005.11401.pdf - and on the one hand, the technique is interesting because it lets you isolate cited sources for the generated text fairly easily. But on the other hand... I remember when people thought using Wikipedia for research was cheating. And it still - sort of - is, in the sense that it's not your primary research, it's other people's distillates of other people's research. And that's what this demo is doing.
And to take that logic a step further: the reason Wikipedia was considered cheating or in some way adjacent to plagiarism wasn't just its presumed unreliability - some people just need to see gatekeepers, I get it even if I don't agree - but the fact that all this supplementary labor went not just uncited or uncredited but unacknowledged.
That is to say, laborwashing.
Yes, it's novel that ML demos can distill coherent sentences out of the products of a tectonic amount of human effort. But.
So far I haven't seen one of these tools that isn't, in one way or another, glossing over the step where a lot of un- or under-paid humans are doing all the heavy lifting that makes any of this work at all. Do you presume to criticize the great and powerful Oz^H AI model? Pay no attention to the hundreds or thousand of people behind the curtain.
https://www.youtube.com/watch?v=YWyCCJ6B2WE
Put differently, every working AI model is an abstraction layer that gives us license to ignore the details of exploitation.
I'm naive enough to believe humane computing is possible, that computational literacy matters, that access to computation should be a human right, but for any of those wild-eyed techno-hippie ideals to matter, computation can't have exploitation at its core. And all the talk about "AI safety" and "human-in-the-loop decision making" is a distraction from the fact that it does, hiding questions about where humans already are in these systems with chin-stroking questions about where they should be.
To some extent, I think our language is failing us here, in the way it always fails us in the presence of real novelty; we seek out analogy to tether the novelty to our existing frames, and the shortcomings of those analogies are where the charlatans of the world do their best work. "We are encoding knowledge" no sir you're encoding text. The knowledge is represented by the text but it is not knowledge. "Signs of AGI" sir you put googly-eyes on a very complicated spreadsheet, we are not fooled.
@internic Who throws away backups though?