Show newer

RT @mapto@twitter.com

Not long ago I presented my experiments on illustrating famous fairy tales with v4 at the (lacam.di.uniba.it/IRCDL23/) This experimentation is part of the VAST project vast-project.eu/ that studies values present in different… qoto.org/@mapto/11011125214281

🐦🔗: twitter.com/mapto/status/16413

Now Midjourney v5 is out and it is all about realism - it gets 5 fingers correct. Also, in my experiments I limited myself to not using image prompts and chaos coefficients. So this experimentation definitely could be continued, but certainly I'm done with with commercial closed generators, and certainly we need a more scalable way to evaluate image correctness

Show thread

Another encountered issue was that it was extremely difficult to force the generator to draw 7 dwarfs. Notoriously, 5 fingers of a hand are also difficult, but the larger the number the smaller hope to get it by chance. A third difficulty I found was the impossibility to generate impossible scenes, typical for fairy tales, such as popping out happily from the wolf's belly or shoving the witch into the oven. A quick comparison with showed that @openai's model handles somewhat better two of the above challenges (other than quantities). Yet, it struggles with unwanted

The 4-step process is not to be seen as something strict. Rather, it is a suggestion that it would bring little benefit to proceed to a subsequent step, without getting the one before at least approximately right. Moving backwards in the process is not ruled out, yet again, before moving back, try consolidating your progress so far. Although it allowed me to generate the intended number of illustrations, while doing this I also confirmed that specific scenes are close to impossible to generate. For example, trying to illustrate how Cinderella plants a tree at her mother's grave, I couldn't find a way to make generate the scene without a grown tree already present - something that makes the scene absurd

Show thread

The final fourth step, once we have a roughly satisfactory result, is to try variations that might just fix the remaining issues through randomisation

Once composition is at least roughly right, the third stage to choose a style that helps the efficiency. One that possibly reduces hallucinations, yet eases interpretability by the viewer. For fairytales, "book illustration" is a possibility

Show thread

In the second stage, considering the outcome of the first step, I aimed to isolate parts of the prompt to be removed, added or replaced to improve the composition of the image, aiming to add important elements and to remove unwanted ones.

Show thread

The first stage of the process is converting the intended text to a prompt, without deviating from the original vocabulary. This means condensing content into a single phrase, removing words that are not meant to be visualised, e.g. this, here, he, and substituting them with what they refer to.

Show thread

Other models followed, in November came Midjourney v4 and in December Structured Diffusion Guidance (arxiv.org/abs/2212.05032). Better composition was notably easier to achieve. This allowed me to proceed with (and complete) my tasks of illustration of fairy tales. All resulting images can be seen in the paper, but the other important outcome is the definition of a preliminary process for the generation of images aligned to the original story text.

Show thread

Unlike the widespread , this task had the constraint to adhere to the original text, thus keep in check some peculiarities of generator models, such as . When working on my task, I quickly discovered the obvious: that -s are somewhat easier to generate for. (example RedCap). However, fairy tales contain much less descriptions than one might recall from child memories. So the challenge of the task was to generate illustrations also for . While it might be too ambitious to try to illustrate for a sequence of events (what would make a narrative), even describing a single event requires an interaction or a scene composition. However, interactions or compositions were notoriously difficult to get right by generative models. Until in mid-2022 Google's Parti (arxiv.org/abs/2206.10789) made a notable breakthrough by linking the image generation to (text) transformer models

To initiate such a process, I engaged with an investigation following the principles of action research (sonyaterborg.com/2016/02/17/ac) - an iterative methodology where research and practice go hand in hand. While a task is being completed, in parallel a reproducible process is being developed, lessons learned are being collected.

Show thread

What does it take to be more rigorous about the composition of prompts? Well, first of all, not only talking about what words are used, but also about how they are chosen, why one choice of words might be better than another? Then, it is important to be careful how features are tried, putting an effort to separate different effects as much as possible. This takes time, but allows for a better understanding of the interplays (doi.org/10.1207/s15327809jls15). And this is only for a start of a long reflective practice that allows for the collection of insights that can be valid also across models, hopefully also for models that are yet to come.

Show thread

With AI text-to-image generator models gaining popularity, there's a lot of talk about , the process of refining the textual input used to obtain a desired result. However, calling the current practices "engineering" is an exaggeration considering that 1. it commonly limits itself to everyone using their ad-hoc techniques, 2. that are barely validated, and 3. and transferability across models is not even considered. For an overview of practices among , consider Jonas Oppenlaender's ethnographic study (arxiv.org/abs/2204.13988) He gives a good idea of some of the employed witchcraft, such as quality boosters, repetitions and magic words

Show thread

Not long ago I presented my experiments on illustrating famous fairy tales with v4 at the (lacam.di.uniba.it/IRCDL23/) This experimentation is part of the VAST project vast-project.eu/ that studies values present in different texts, including some fairy tales recorded by the Grimm brothers. My generations were an aside activity, and I attempted to generate 5 illustrations for each of 5 fairy tales: , , , , and

And could the creators "reliably control" #ChatGPT et al. Yes, they could --- by simply not setting them up as easily accessible sources of non-information poisoning our information ecosystem.

And could folks "understand" these systems? There are plenty of open questions about how deep neural nets map inputs to outputs, but we'd be much better positioned to study them if the AI labs provided transparency about training data, model architecture, and training regimes.

>>

Show thread

Okay, so that AI letter signed by lots of AI researchers calling for a "Pause [on] Giant AI Experiments"? It's just dripping with AI hype. Here's a quick rundown.

First, for context, note that URL? The Future of Life Institute is a longtermist operation. You know, the people who are focused on maximizing the happiness of billions of future beings who live in computer simulations.

futureoflife.org/open-letter/p

#AIhype

>>

Join the rally in #SanFrancisco! Big Media is suing to cut off #libraries’ ownership and control of digital books, opening new paths for #censorship & #surveillance.

On 4/8, we rally at 🐦InternetArchive to demand #DigitalRightsForLibraries ✊🏻📚

RSVP: actionnetwork.org/events/dont-

An enquiry into photograph smiles and how Midjourney distorts these

RT @hackylawyER@twitter.com

"It was as if the AI had cast 21st century Americans to put on different costumes and play the various cultures of the world. Which, of course, it had."

Great article from @babiejenks@twitter.com medium.com/@socialcreature/ai-

🐦🔗: twitter.com/hackylawyER/status

At @IslabUnimi@twitter.com we're using bespoke (quantitative) word embeddings as a tool for comparative corpus analysis. I see it as a first step towards better informed critical analysis. Here's an example by @eliroc98@twitter.com and Tommaso Locatelli
tales.islab.di.unimi.it/2023/0

Show older
CleverLibre Social

CleverLibre Social is an inclusive social instance for open discussion, learning, and community.
All cultures welcome.
Hate speech and harassment strictly forbidden.