Section 1: The dataset and the problem
The hidden distribution, the dataset I actually get, and why generative modeling starts there.
I am working through Mathematical Foundations of Generative AI, Prof. Prathosh AP’s public lectures. The playlist is the spine. These notes are my deep dive on each section: a visual when the picture is the point, a detour when a prerequisite is doing real work, and the formula in my own words.
Section 1: The dataset and the problem
Everything later in the playlist, GANs, VAEs, diffusion, transformers, state-space models, and preference tuning, starts from the same setup. I want this part solid before I touch a model.
The setup I am carrying:
Data (unknown), with , where is the dimensionality of the data. Example: is an image, , so . Goal: estimate and learn to sample from it.
Intuition
Somewhere there is a hidden recipe, a probability distribution, that produced every cat photo, every English sentence, every speech clip. I never get the recipe. I get a pile of examples it produced. Generative modeling is using that pile to reconstruct the recipe well enough to cook up new examples from the same kitchen.
An analogy I can check
Think of a production web service whose request process I cannot inspect. All I have is an access log of 1,000 requests. I want a load-testing tool that produces synthetic traffic indistinguishable from real users: same mix of endpoints, same payload sizes, same timing.
| Load-testing analogy | Notation |
|---|---|
| The real, unknown user behaviour | , the true data distribution |
| One logged request | , one data point |
| The fields in a log line | the coordinates of |
| The access log | dataset |
| The traffic generator | the generative model |
| Synthetic requests it emits | samples from |
Two consequences. I do not need to write the user-behaviour formula down. I need a generator whose output is statistically indistinguishable. That is why a GAN never writes down at all. And replaying the log is not the goal. I want new requests, not copies.
Visual
A hidden 2-D distribution plays the role of . At first I only see the dots, the dataset. Reveal to see what is normally invisible, and drag to see how more data reveals the shape.
Each “Draw a new dataset” gives a different and the same green underneath. The dataset is random. The distribution is the fixed thing I am after. With the shape is a guess. With the two clumps are obvious.
Detours I needed before the formula
Random variable, and random vector. A random variable is a quantity whose value is decided by chance, like a die roll . A random vector is the same idea with several numbers at once, such as height and weight of a randomly chosen person, . Capital means the random thing in general. Lowercase means one specific value that actually came out.
Distribution and density. tells me how likely each possible value is. For a die it is a table: each face has probability 1/6. For continuous values, such as pixel intensities, it is a density , a heat map over space. High density means values land there often. In the chart above, the green shading is that density.
Two facts I will need constantly:
The second says all probability mass integrates to 1. I will use loosely for both the distribution and its density. That is standard, and I will be explicit when the difference matters.
iid, independent and identically distributed.
- Identically distributed: every comes from the same . Every log line came from the same user population, not some from production and some from a test environment.
- Independent: knowing tells me nothing extra about . The first roll of a die does not influence the second.
In notation, and .
The formula
Read aloud: the dataset is a collection of points. Each point is an independent draw from the same unknown distribution . Each point is a vector of real numbers.
| Symbol | What it is | Type / shape | Role |
|---|---|---|---|
| the dataset | a set of vectors | the only thing I actually observe | |
| number of samples | integer | more data, a better picture of | |
| one data point | vector in | an image, a sentence embedding, a signal | |
| dimensionality | integer | how many numbers describe one point | |
| ”is drawn from” | relation | links data to its source | |
| iid | independent and identically distributed | assumption | makes learning from samples legitimate |
| true data distribution | distribution on | the hidden target, unknown |
Why “unknown” is the whole problem. If I knew , I would sample from it and stop. Every model in the playlist is a different trick for working with through samples alone.
Why iid matters. It licenses replacing an expectation, an average over , with an average over the dataset. That is the law of large numbers:
That line is the engine behind every training loss later. The chart shows a small version: the sample mean settles as grows.
Why is so large. A 400×400 RGB image is one point in , one coordinate per pixel per colour channel. Real images occupy a thin region of that space. Random pixel vectors look like television static. This is the manifold hypothesis, and it comes back when naive GAN training falls over.
Two jobs hiding in one goal
Given , estimate and learn to sample from it. Those are different jobs, and different models emphasise different ones.
| Job | Meaning | Who emphasises it |
|---|---|---|
| Density estimation | answer “how likely is this ?“ | autoregressive models, VAEs through the ELBO |
| Sampling | produce a new | GANs, which never write a density down, and diffusion |
A GAN can sample well and still be unable to tell me for a given image. I want to remember that split.
A worked example small enough to do by hand
Let and . Each point is hours slept and coffees drunk for a random person:
Each , so . Here means 6 hours of sleep and 2 coffees. iid means three different people, surveyed separately, from the same population.
I want under the unknown . I cannot integrate against a distribution I do not have. iid lets me estimate it from the sample: .
The generative goal is a machine that outputs new pairs such as : plausible, not copied from , and following the same pattern. More sleep tends to go with less coffee. That dependency between coordinates is what captures, and a per-column average does not.
The same picture with and in the millions is image generation.
Where the rest of the playlist sits, as I understand it now
| Stretch | Model | How it attacks “learn from “ |
|---|---|---|
| Early | GAN / divergence minimisation | push through and shrink an f-divergence to |
| Next | VAE | latent-variable , maximise the ELBO |
| Then | WGAN | the same game, with the Wasserstein distance |
| Then | DDPM | add noise to , learn to reverse it |
| Then | Autoregressive / transformer | factorise |
| Then | SSM / Mamba | another sequence model |
| Late | RLHF / DPO | move the learned distribution toward preferences |
The next section turns the goal into a recipe: pick a parametric family , define a divergence, and minimise it using only samples.
Questions I want to be able to answer before I move on:
- If a dataset mixes 5,000 phone photos and 5,000 medical scans, which part of iid breaks, and why does that hurt?
- What is for a 28×28 grayscale digit? For a 64×64 RGB image?
- Once I have , do I know exactly? What changes as ?
Next: Section 2: The general principle of generative models.