News

Currently, no news are available

Visual Generative Models

This lecture introduces continuous generative modeling, diffusion and flow matching, as the dominant paradigm for visual synthesis, and traces its extension from static images to dynamic, embodied domains. We start with the core mathematical picture: forward noising processes, score/velocity-field learning, and the shared ODE/SDE view that unifies diffusion models and flow matching as a generative modeling framework that learns a velocity field to transform a source distribution into a target distribution through a continuous flow. From there, we move to video diffusion models and their emerging role as implicit world models, before turning to robotics: Diffusion Policy as a case study in using denoising processes to represent multimodal action distributions, and video generation as a substrate for planning and simulation in physical AI. The lecture closes with open questions around sampling efficiency, controllability, and what it means for a generative model to also be a world model.

Website: https://vip.mpi-inf.mpg.de/vgm.html

CMS Registration will open after the introductory lecture on Oct 12!


Prerequisites

  • Linear algebra and multivariate calculus
  • Basic probability
  • Deep learning fundamentals
  • Python and PyTorch

Dates and Locations

  • Lecture: Mondays, 4pm–6pm (starts at 4 pm, from 2nd lecture), Room E1 5 (MPI-SWS building) R.002
  • Tutorial: Wednesdays, 12pm–2pm (starts at 12:15), Room E1 4 (MPI-INF building) R.024
  • The first, introductory lecture will be on October 12 and will start at 4:15 pm.
  • Tutorials will start on October 21.

Preliminary Schedule of Content

The following is a preliminary, high-level outline of the topics covered and may still change.

Foundations
  • Generative modeling as transport: probability paths, flows, and velocity fields
  • Flow matching: the conditional flow matching objective
  • Diffusion models as a special case of flow matching
  • Samplers, discretization error, and the cost of sampling
Practice, Steering and Judgment
  • Latents and architectures: autoencoders, tokenizers, diffusion transformers
  • Conditioning, guidance and control
  • Evaluation: measuring fidelity and coverage, reward models as learned evaluators
Efficiency and Post-Training
  • Few-step sampling: distillation, consistency models, flow maps
  • Reasoning and RL post-training of flow models
Video
  • Video generation: spatiotemporal flows, generation regimes, long-horizon consistency
Embodied Generation and World Models
  • Robot learning
  • From denoising to velocity-field policies: Diffusion Policy and flow-based policies
  • World models and simulators: action conditioning, planning, synthetic data

 

Material

Material will appear here.

Privacy Policy | Legal Notice
If you encounter technical problems, please contact the administrators.