Multimodal AI: Vision & Images

Show a photo of a fridge to a modern model and ask "what can I cook tonight?" It reads the ketchup, the half-onion, the carton of eggs, and gives you an omelette recipe. The model never "sees" pixels the way you do. It turns the image into the same kind of num

8 lessons, each with runnable code in the browser.

  1. Models That Can See
  2. Sending an Image to a Model
  3. OCR & Document Understanding
  4. Image Generation Basics
  5. Costs & Limits of Vision
  6. Prompting With Images
  7. Multi-Image and Video Frames
  8. Editing and Variations

Compilearn home