Overview
How does a text prompt become a photorealistic image? This session demystifies diffusion models — the technology behind AI image generators — plus speech-to-text, text-to-speech, and voice cloning. We finish with the uncomfortable flip side: deepfakes, and how to spot them.
Learning objectives
- Describe how diffusion models turn noise into images, in plain language.
- Generate images with well-crafted prompts and iterate toward a goal.
- Use speech-to-text to caption a video and text-to-speech to voice one.
- Identify common signs of AI-generated and deepfaked media.
Session agenda
- 6:00 – 6:10Warm-up
Study-guide homework debrief: where did the AI shine and where did it fail?
- 6:10 – 7:15Lecture: Diffusion models, speech, and deepfakes
Noise-to-image intuition for diffusion; speech-to-text and text-to-speech; a live voice-cloning demo and the deepfake problem.
- 7:15 – 7:25Break
- 7:25 – 9:00Lab: Create and caption
Generate images from prompts and refine them; auto-caption a short video; experiment with text-to-speech voices.
- 9:00 – 9:10Wrap-up
Gallery walk of generated images; preview AI and data.
Materials
- Free image generation access (e.g. Bing Image Creator, Gemini, or similar).
- A short video clip to caption (yours or provided in class).
- Headphones for the audio portions of the lab.
Homework
Generate one image you are proud of and save the full prompt history that got you there. Find one real-world deepfake news story and bring a one-sentence summary.
Detailed slides, lab worksheets, and demos for this session are in progress and will be posted here before class.