Caption & Alt-Text Generator
A vision model never opens a photo file. It reads bytes that have been base64-encoded into text and packed into a JSON content block next to your prompt. This lesson builds that payload. Every later lesson reuses it. What multimodal input means. A normal text
8 lessons, each with runnable code in the browser.