Words as Coordinates

Vectors

Part of: Embeddings & Semantic Search

Your phone's autocomplete knows "cat" and "kitten" are related but "cat" and "car" aren't, even though "cat" and "car" share more letters. How? It doesn't read letters. It reads numbers. The word "cat" becomes something like [4, 1] and the machine compares those numbers, not the spelling. What a vector is A vector is just an ordered list of numbers. Picture every word as a dot on a map. "cat" lives at coordinates [4, 1]. "kitten" lives nearby at [3, 1]. "car" is way off at [1, 5]. That list of coordinates is the vector. The map can have 2 dimensions like here, or 1,536 dimensions like a real embedding model. Same idea, more axes. The core idea of embeddings is one sentence: place similar things close together and different things far apart. Once meaning lives as coordinates, "find related words" turns into "find nearby dots," which is just arithmetic. How you compare two vectors You need two tools, and you build each one once: - Dot product , multiply matching components, then add them up. [4,1] · [3,1] = 4 3 + 1 1 = 13. A bigger result roughly means the two arrows point in a similar direction. - Magnitude , the length of the arrow from the origin to the dot. [4,1] = sqrt(4² + 1²)

Challenge: Nearest Neighbor in Embedding Space