Chapter 1 · 2013
How does a computer know similar words?
You type “kitten” into a search box and it also finds pages about cats. That feels obvious to you. To an old computer program, those were two different ID numbers. No family. No neighborhood. Just labels.
The old way treated words like locker numbers.
Locker 17 is not “near” locker 4,003 in any useful sense. So “cat” and “kitten” were as unrelated as “cat” and “microscope.”
same alphabet · zero neighborhood
What if a word’s meaning was just… who it hangs out with?
They taught the map with two little guessing games.
Not with a dictionary. With practice, over and over, on real sentences.
Cover one word. Show its neighbors. Guess the missing one.
neighbors in → missing word out
Show one word. Guess who usually stands next to it.
one word in → likely company out
Every wrong guess nudges the houses around. After enough sentences, words that keep the same company end up living next door. Words that never meet drift apart.
Cat and kitten share so much company that they become neighbors — without anyone writing a definition.
Researchers called these neighborhood maps word vectors. The two games got names too: continuous bag-of-words (hide the middle) and skip-gram (guess the neighbors).
The map could even take a walk.
Start at “king.” Step the way you’d go from “man” toward “woman.” You land near “queen.” Not because anyone typed a royal family tree — because those words keep analogously similar company in lots of text.
a walk, not a lookup table
It worked — and it was fast enough to train on a lot of text.
Similar words clustered. Analogies like the king–queen walk showed up. The point of the paper was not a prettier dictionary. It was a way to learn those neighborhoods efficiently.
That’s why this still matters.
Later machines still start by putting words on a map. ChatGPT does not look up locker numbers. It looks up neighborhoods.