◈Training Images
Drop images here
or click to browse (2-12 images)
0/12 uploaded
Images are downscaled to 32×32 for training. We're running a neural net in JS, not a data center 😅
◈Training Controls
Network Visualization
◈Generation
Train the model first, then watch noise become art
(Art is a strong word for 32×32 pixel blobs, but we believe in this model)
1. The Forward Process (Adding Noise)
We take each training image and gradually add Gaussian noise to it at different timesteps. At t=0, the image is clean. At t=max, it's pure noise. The model needs to learn what noise was added at each step.
x_noisy = α(t) × x_clean + σ(t) × ε, where ε ~ N(0, I)
2. Training the Denoiser
A small MLP (multi-layer perceptron) is trained to predict the noise ε that was added. It takes the noisy image + timestep embedding as input and outputs the predicted noise. We minimize MSE loss between predicted and actual noise.
3. The Reverse Process (Generating)
Start from pure random noise. At each step, the network predicts the noise component, and we subtract a fraction of it. Step by step, structure emerges from chaos. Like sculpting, but with math.
x_{t-1} = (x_t - β × predicted_noise) / α + small_noise
4. Why This Is Simplified
Real diffusion models (Stable Diffusion, DALL-E) use U-Nets with millions/billions of parameters, attention mechanisms, trained on millions of images with hundreds of GPUs. This toy uses a 3-layer MLP with ~800K parameters running in your browser's single JS thread. It learns patterns from your tiny dataset — don't expect Mona Lisa! 🎨