Caption & Alt-Text Generator

A vision model never opens a photo file. It reads bytes that have been base64-encoded into text and packed into a JSON content block next to your prompt. This lesson builds that payload. Every later lesson reuses it. What multimodal input means. A normal text

8 lessons, each with runnable code in the browser.

  1. Sending an Image to a Vision Model
  2. The Smallest Caption Call
  3. Caption vs Alt Text: Two Different Jobs
  4. Alt Text Has Rules
  5. Structuring the Output as JSON
  6. When the Upload or the Reply Breaks
  7. Vision Calls Cost More Than Text
  8. Ship the Captioner

Compilearn home