RT @mapto@twitter.com
Not long ago I presented my experiments on illustrating famous fairy tales with #Midjourney v4 at the #IRCDL23 (http://lacam.di.uniba.it/IRCDL23/) This experimentation is part of the VAST project https://www.vast-project.eu/ that studies values present in different… https://qoto.org/@mapto/110111252142814403
Now Midjourney v5 is out and it is all about realism - it gets 5 fingers correct. Also, in my experiments I limited myself to not using image prompts and chaos coefficients. So this experimentation definitely could be continued, but certainly I'm done with with commercial closed generators, and certainly we need a more scalable way to evaluate image correctness
Another encountered issue was that it was extremely difficult to force the generator to draw 7 dwarfs. Notoriously, 5 fingers of a hand are also difficult, but the larger the number the smaller hope to get it by chance. A third difficulty I found was the impossibility to generate impossible scenes, typical for fairy tales, such as #RedCap popping out happily from the wolf's belly or #Gretel shoving the witch into the oven. A quick comparison with #dallE showed that @openai's model handles somewhat better two of the above challenges (other than quantities). Yet, it struggles with unwanted #artefacts
The 4-step process is not to be seen as something strict. Rather, it is a suggestion that it would bring little benefit to proceed to a subsequent step, without getting the one before at least approximately right. Moving backwards in the process is not ruled out, yet again, before moving back, try consolidating your progress so far. Although it allowed me to generate the intended number of illustrations, while doing this I also confirmed that specific scenes are close to impossible to generate. For example, trying to illustrate how Cinderella plants a tree at her mother's grave, I couldn't find a way to make #Midjourney generate the scene without a grown tree already present - something that makes the scene absurd
Once composition is at least roughly right, the third stage to choose a style that helps the efficiency. One that possibly reduces hallucinations, yet eases interpretability by the viewer. For fairytales, "book illustration" is a possibility
In the second stage, considering the outcome of the first step, I aimed to isolate parts of the prompt to be removed, added or replaced to improve the composition of the image, aiming to add important elements and to remove unwanted ones.
The first stage of the process is converting the intended text to a prompt, without deviating from the original vocabulary. This means condensing content into a single phrase, removing words that are not meant to be visualised, e.g. this, here, he, and substituting them with what they refer to.
Other models followed, in November came Midjourney v4 and in December Structured Diffusion Guidance (https://arxiv.org/abs/2212.05032). Better composition was notably easier to achieve. This allowed me to proceed with (and complete) my tasks of illustration of fairy tales. All resulting images can be seen in the paper, but the other important outcome is the definition of a preliminary process for the generation of images aligned to the original story text.
Unlike the widespread #AIart, this task had the constraint to adhere to the original text, thus keep in check some peculiarities of generator models, such as #hallucinations. When working on my task, I quickly discovered the obvious: that #DescriptiveText-s are somewhat easier to generate for. (example RedCap). However, fairy tales contain much less descriptions than one might recall from child memories. So the challenge of the task was to generate illustrations also for #narrative. While it might be too ambitious to try to illustrate for a sequence of events (what would make a narrative), even describing a single event requires an interaction or a scene composition. However, interactions or compositions were notoriously difficult to get right by generative models. Until in mid-2022 Google's Parti (https://arxiv.org/abs/2206.10789) made a notable breakthrough by linking the image generation to (text) transformer models
To initiate such a process, I engaged with an investigation following the principles of action research (https://sonyaterborg.com/2016/02/17/action-research/) - an iterative methodology where research and practice go hand in hand. While a task is being completed, in parallel a reproducible process is being developed, lessons learned are being collected.
What does it take to be more rigorous about the composition of prompts? Well, first of all, not only talking about what words are used, but also about how they are chosen, why one choice of words might be better than another? Then, it is important to be careful how features are tried, putting an effort to separate different effects as much as possible. This takes time, but allows for a better understanding of the interplays (https://doi.org/10.1207/s15327809jls1502_2). And this is only for a start of a long reflective practice that allows for the collection of insights that can be valid also across models, hopefully also for models that are yet to come.
With AI text-to-image generator models gaining popularity, there's a lot of talk about #promptEngineering, the process of refining the textual input used to obtain a desired result. However, calling the current practices "engineering" is an exaggeration considering that 1. it commonly limits itself to everyone using their ad-hoc techniques, 2. that are barely validated, and 3. and transferability across models is not even considered. For an overview of practices among #aiartists, consider Jonas Oppenlaender's ethnographic study (https://arxiv.org/abs/2204.13988) He gives a good idea of some of the employed witchcraft, such as quality boosters, repetitions and magic words
Not long ago I presented my experiments on illustrating famous fairy tales with #Midjourney v4 at the #IRCDL23 (http://lacam.di.uniba.it/IRCDL23/) This experimentation is part of the VAST project https://www.vast-project.eu/ that studies values present in different texts, including some fairy tales recorded by the Grimm brothers. My generations were an aside activity, and I attempted to generate 5 illustrations for each of 5 fairy tales: #LittleRedRidingHood, #Cinderella, #LittleSnowWhite, #HanselAndGretel, and #FaithfulJohannes
On the "sparks" paper:
https://twitter.com/emilymbender/status/1638891855718002691?s=20
On the GPT-4 ad copy:
https://twitter.com/emilymbender/status/1635697381244272640?s=20
On "general" tasks:
https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/084b6fbb10729ed4da8c3d3f5a3ae7c9-Abstract-round2.html
>>
And could the creators "reliably control" #ChatGPT et al. Yes, they could --- by simply not setting them up as easily accessible sources of non-information poisoning our information ecosystem.
And could folks "understand" these systems? There are plenty of open questions about how deep neural nets map inputs to outputs, but we'd be much better positioned to study them if the AI labs provided transparency about training data, model architecture, and training regimes.
>>
Okay, so that AI letter signed by lots of AI researchers calling for a "Pause [on] Giant AI Experiments"? It's just dripping with AI hype. Here's a quick rundown.
First, for context, note that URL? The Future of Life Institute is a longtermist operation. You know, the people who are focused on maximizing the happiness of billions of future beings who live in computer simulations.
https://futureoflife.org/open-letter/pause-giant-ai-experiments/
>>
Join the rally in #SanFrancisco! Big Media is suing to cut off #libraries’ ownership and control of digital books, opening new paths for #censorship & #surveillance.
On 4/8, we rally at 🐦InternetArchive to demand #DigitalRightsForLibraries ✊🏻📚
RSVP: https://actionnetwork.org/events/dont-delete-our-books-rally-in-san-francisco?source=twitter&
An enquiry into photograph smiles and how Midjourney distorts these
RT @hackylawyER@twitter.com
"It was as if the AI had cast 21st century Americans to put on different costumes and play the various cultures of the world. Which, of course, it had."
Great article from @babiejenks@twitter.com https://medium.com/@socialcreature/ai-and-the-american-smile-76d23a0fbfaf
🐦🔗: https://twitter.com/hackylawyER/status/1641157473833762833
At @IslabUnimi@twitter.com we're using bespoke (quantitative) word embeddings as a tool for comparative corpus analysis. I see it as a first step towards better informed critical analysis. Here's an example by @eliroc98@twitter.com and Tommaso Locatelli
http://tales.islab.di.unimi.it/2023/03/27/john2vec-or-embedding-deweys-philosophy/
#DigitalHumanities
Studying how people interact, in the past (#CulturalAnalytics) and today (#EdTech #Crowdsourcing). Researcher at @IslabUnimi, University of Milan. Bulgarian activist for legal reform with @pravosadiezv. I use dedicated accounts for different languages.
My profile is searchable with https://www.tootfinder.ch/