Saltar al contenido
Let's talk
Back to the blog
Artificial Intelligence

Understanding Deep Learning through image generation.

Images 1 and 6 are synthetic: generated by AI from a time series of images.

6 min read

Images 1 and 6 are synthetic: generated by AI from a time series of images.

2, 3, 4 and 5 are human creations. Mine, to be more precise. I draw them in an app called "Paper" (originally from FiftyThree, now from WeTransfer) with my fingers on the screen. I generated image 2 first and so on. Human, and time to unwind.

I shared with the AI a set of pieces with a similar style and aesthetic, and let OpenAI generate what would hypothetically have been the first of the series and the next one.

We humans would read into it a compositional logic, a chromatic logic, a spatial logic. For the AI it's a game of attributes, iterations and probabilistic distance. And even so, the result is simply wonderful.

๐™๐™š๐™˜๐™๐™ฃ๐™ž๐™˜๐™–๐™ก๐™ก๐™ฎ: ๐š๐š‘๐šŽ ๐š–๐š˜๐š๐šŽ๐š• ๐šŽ๐š—๐šŒ๐š˜๐š๐šŽ๐šœ ๐šŽ๐šŠ๐šŒ๐š‘ ๐š’๐š–๐šŠ๐š๐šŽ ๐š’๐š— ๐šŠ ๐š‘๐š’๐š๐š‘-๐š๐š’๐š–๐šŽ๐š—๐šœ๐š’๐š˜๐š—๐šŠ๐š• ๐š•๐šŠ๐š๐šŽ๐š—๐š ๐šœ๐š™๐šŠ๐šŒ๐šŽ, ๐š ๐š‘๐šŽ๐š›๐šŽ "๐šœ๐š๐šข๐š•๐šŽ" ๐š’๐šœ ๐š›๐šŽ๐š™๐š›๐šŽ๐šœ๐šŽ๐š—๐š๐šŽ๐š ๐šŠ๐šœ ๐šŠ ๐šŸ๐šŽ๐šŒ๐š๐š˜๐š›. ๐™ธ๐š ๐šŠ๐šŸ๐šŽ๐š›๐šŠ๐š๐šŽ๐šœ ๐š๐š‘๐š˜๐šœ๐šŽ ๐šŸ๐šŽ๐šŒ๐š๐š˜๐š›๐šœ ๐š๐š˜ ๐š’๐š—๐š๐šŽ๐š› ๐š๐š‘๐šŽ ๐š๐š’๐š›๐šŽ๐šŒ๐š๐š’๐š˜๐š— ๐š˜๐š ๐š๐š‘๐šŽ ๐šœ๐šŽ๐š›๐š’๐šŽ๐šœ ๐šŠ๐š—๐š, ๐š๐š‘๐š›๐š˜๐šž๐š๐š‘ ๐š›๐šŽ๐šŸ๐šŽ๐š›๐šœ๐šŽ ๐š๐š’๐š๐š๐šž๐šœ๐š’๐š˜๐š—, ๐š’๐š ๐šœ๐š๐šŠ๐š›๐š๐šœ ๐š๐š›๐š˜๐š– ๐š—๐š˜๐š’๐šœ๐šŽ ๐šŠ๐š—๐š ๐š™๐š›๐š˜๐š๐š›๐šŽ๐šœ๐šœ๐š’๐šŸ๐šŽ๐š•๐šข ๐š๐šŽ๐š—๐š˜๐š’๐šœ๐šŽ๐šœ ๐š’๐š (๐š›๐šŽ๐š๐šž๐šŒ๐š’๐š—๐š ๐š๐š‘๐šŽ ๐š—๐š˜๐š’๐šœ๐šŽ) ๐šž๐š—๐š๐š’๐š• ๐š’๐š ๐š–๐šŠ๐š๐šŽ๐š›๐š’๐šŠ๐š•๐š’๐šฃ๐šŽ๐šœ ๐šŠ ๐š™๐š’๐šŽ๐šŒ๐šŽ ๐šŒ๐š˜๐š‘๐šŽ๐š›๐šŽ๐š—๐š ๐š ๐š’๐š๐š‘ ๐š๐š‘๐šŠ๐š ๐šœ๐š๐šข๐š•๐š’๐šœ๐š๐š’๐šŒ ๐š๐š›๐šŠ๐š“๐šŽ๐šŒ๐š๐š˜๐š›๐šข. ๐™ธ๐š ๐š๐š˜๐šŽ๐šœ ๐š—๐š˜๐š ๐šŒ๐š˜๐š™๐šข: ๐š’๐š ๐š’๐š—๐š๐šŽ๐š›๐š™๐š˜๐š•๐šŠ๐š๐šŽ๐šœ.

In other words, the model encodes the images in a latent space, or a mathematical map of thousands of dimensions where visual features are represented as semantic vectors. By averaging these vectors, the AI identifies the stylistic trajectory of the series. It doesn't copy images: it interpolates values.

And meanwhile, I don't lose my capacity for wonder at the nuances, and at how an AI is able to generate the first and the sixth of a human series of four images.

Images 1, 2, 3 and 4. Human creations

Latent Space: Mapping creativity mathematically

Imagine you had to describe each of your drawings with, let's say, 512 sliders. One controls "how circular vs. organic it is", another "chromatic saturation", another "density of concentric layers", another "presence of navy blue"โ€ฆ and so on up to 512. No slider has a name a human would recognize โ€”they are dimensions the algorithmic model has discovered on its own (deep learning) during trainingโ€”, but between them all they capture the essence of any image.

That 512-dimensional space is the latent space. The number, actually, is quite a bit larger: in models like Stable Diffusion, the space where generation happens has around 16,000 dimensions, and the associated semantic spaces โ€”the ones that encode "style" or "concept"โ€” are around 768 or 1,024.

The observable universe contains around 10โธโฐ atoms. With just 512 binary sliders, the latent space already has more possible combinations than there are atoms in the universe, raised to almost twice that power.

In larger closed models, the figures go up to several thousand. The idea of "512 sliders" is a pedagogical simplification. What matters is to retain that we are talking about thousands of axes, not hundreds, and that each one is a dimension learned automatically, with no human name. Inferences.

Every image in the world is a point in there. Similar images live close by; radically different images, far away.

What is powerful about the latent space is that inside it you can do arithmetic with meaning. The classic example from word embeddings: king โˆ’ man + woman โ‰ˆ queen. With images it works the same way: smiling portrait โˆ’ neutral portrait = "smile vector", which you can then add to any other portrait to make it smile.

The six drawings are six points in that space. They form a small cloud. The AI can calculate the center of that cloud (the "average style"), its dominant direction (where the series is heading), or a point interpolated between any two of them. That is what it does to infer "the first one" (image 1) and "the next one" (image 6): it locates a plausible position inside your stylistic cloud. It doesn't look for the exact point; it samples a plausible position. That is why, if we asked it to generate "the seventh image" ten times, it would give us ten different images, all coherent with my style. It is a stochastic process, not a deterministic one. Does the idea of non-deterministic models click now?

The latent space is a compressed map of the visual world where geometry means something. Moving in one direction amounts to transforming the image in a coherent way.

Image 1 and Image 6 in the series

Diffusion: Stabilizing creation out of noise

Diffusion is the generation mechanism. And its logic is counterintuitive but beautiful.

During training, the AI does the opposite of generating. It takes a real image and adds Gaussian noise to it little by little, step by step, until it turns it into pure static โ€”that white noise of a badly tuned televisionโ€”. And at each step, it learns to predict what noise has just been added. It repeats this millions of times with millions of images. It learns to destroy images in a controlled way. It learns to tell what's superfluous for an image to make sense again.

When the moment to generate arrives, the AI reverses the process. It starts from pure noise โ€”literally random values drawn from a Gaussian bell curveโ€” and, step by step, subtracts the noise that "it" itself predicts. Little by little it cleans up that noise. At each step it removes a part of the excess until a coherent image starts to appear. With each iteration the image is a little less dirty, a little more defined. After fifty or a hundred steps, a coherent image emerges from the chaos. And all of this happens in seconds, maybe minutes.

A technical nuance: the noise it starts from doesn't live in the semantic space โ€”the one of "style, concept, composition"โ€”, but in the autoencoder's latent space, a compressed representation of the image. The semantic space acts differently: as a conditioning vector that guides the process. At each of the ~50 steps, the model receives two inputs โ€”the current noise and that conditionโ€” and predicts which portion of the noise is excess in order to move closer to an image coherent with the style requested. The random noise is the raw material; the embedding of your series is the force field that sculpts it.

The most precise metaphor we can use is that of the sculptor Michelangelo. Michelangelo di Lodovico Buonarroti Simoni used to say that the David was already inside the block of marble, and that he only removed the excess.

Diffusion is exactly that. The image is already latent in the noise.

The model only removes, step by step, what does not belong. With one addition that Michelangelo did not have: the "David" that emerges depends on which stylistic condition is whispered to the chisel at each blow.

These days the chisel or the brush can be physical or semantic. Starting from words or from images or from audio or from text, etc. The chisel or the brush are the variable. And they can also be a constant. Creativity still exists. And what we call art is an emotional consideration.

Are images 1 and 6 art?

The author

Bernardo Crespo

C-suite advisor in AI, data and strategy. CEO of Quantum Markethink and Academic Director at IE. He helps leadership teams make sound decisions in the age of AI.

LinkedIn