A lesson on word embeddings, played as a game
Word Bocce is a free browser game played on a map of word meanings. Every word is a point; you add and subtract words to roll your ball toward a target word, the jack. It's a hands-on way into word embeddings, and into why it's hard to explain what a learned model is doing.
This page is a lesson plan for 30 to 45 minutes, written for intro NLP, machine learning, and data-science classes. It also works for curious high-school classes if you skip the formula in Activity 2.
Learning goals
By the end, students should be able to:
- Explain that a word embedding stores each word as a list of numbers, and that "close in meaning" is measured by cosine similarity.
- Do word arithmetic (a start word plus some words, minus others) and reason about where the result lands, including why rank and similarity can disagree.
- Describe how training data shapes the map, by comparing an embedding learned only from text with one that also uses a knowledge base of everyday facts.
- Tell the difference between an exact breakdown of a score and an explanation of why a model learned what it did.
- Name some limits: one point per word (so different senses get mixed), associations inherited from the training text, and how fixed word vectors differ from what's inside a modern language model.
Before class
- Open wordbocce.davidreinstein.org on the projector. A first visit starts a 30-second tutorial, which doubles as the warm-up.
- These links jump straight to a mode, handy on a slide or as a QR code:
/#tutorial,/#daily,/#puzzles,/#practice. - Each device downloads the word data once: about 6.5 MB for the default word set, about 4.4 MB for the other. On a busy school network, have students open the page as they come in.
- The game starts on the common-sense word set. To switch, use the Words button in the header; on a phone, tap Help, open "Why does adding words work?", and tap "common sense, or raw text". Each device remembers its choice.
- The Daily court is the same for everyone using the same word set on the same date.
- Privacy: there's no sign-up, and scores stay in the browser. One thing is optional and sent to the game's server: after a round, the "Try your own words" box asks whether other words would have made the court more fun. If a student answers, that suggestion (the court, their words, any note they type, and a random ID for the device) is sent so the word lists can be improved. Online rooms connect browsers directly, and nothing is stored.
Lesson outline
| Minutes | Activity | In the game |
|---|---|---|
| 0–5 | Warm-up: word arithmetic | Tutorial, on the projector |
| 5–15 | Activity 1: today's court | Daily |
| 15–27 | Activity 2: reading the "Why?" panel | Puzzles |
| 27–37 | Activity 3: two maps of meaning | Words switch, Puzzles |
| 37–45 | Discussion |
Warm-up: word arithmetic (5 minutes)
On the board, give four words two made-up coordinates:
| Word | Royal | Female |
|---|---|---|
| man | 0 | 0 |
| woman | 0 | 1 |
| king | 1 | 0 |
| queen | 1 | 1 |
Then king − man + woman = (1, 0) − (0, 0) + (0, 1) = (1, 1), which is queen.
Real word embeddings work the same way, with two differences. There are hundreds of coordinates (100 per word in the game's raw-text set, 300 in the common-sense set). And nobody chose what they mean: they're learned from data, and no single number means "royal".
Closeness is measured with cosine similarity: 1 means two vectors point the same way, 0 means unrelated. The game also shows rank, which is easier to read: #1 means the jack is the nearest word to where your ball stopped.
Now ask the class to predict hat + foot − head, and play the tutorial on the projector:
- At the start, shoe is the 249th nearest word to hat.
- After + foot, it's 5th. The ball stops near "toe".
- After + foot − head, it's 1st: a bacio, the bocce word for a ball kissing the jack.
Leave one question open for Activity 2: why did taking away "head" help?
Activity 1: today's court (10 minutes)
- Everyone opens Daily. They all get the same start word, jack, and nine tiles.
- In pairs, before throwing, sort the nine tiles into three piles: words that pull toward the jack, words worth subtracting, and traps. (The deal is built this way: two tiles pull toward the jack, two carry the start word's flavor and are worth subtracting, two are linked to both, and three look related to the jack but pull only weakly.)
- Each player gets four balls, up to three tiles per throw, and the best ball counts.
- At the end, the game shows par (the best throw the hand allowed), a few nearby words on the court, and the jack's nearest words. "Copy result" gives a short summary students can paste into the class chat.
Debrief: Who matched par? Which tile was the best trap? Did anyone's "obvious" throw go the wrong way, and can they say why?
Activity 2: reading the "Why?" panel (10–12 minutes)
After every throw there's a Why? link. It splits the ball's similarity to the jack into one share per word. The shares add up exactly, because the ball is just the sum of the word vectors (each of length 1), rescaled:
cos(ball, jack) = Σ sign × cos(word, jack) ÷ |Σ sign × word|
The start word counts as a + word. Here is the tutorial throw, with numbers from the game:
| Word | Similarity to shoe | Share |
|---|---|---|
| hat (start) | 0.23 | 0.16 |
| + foot | 0.50 | 0.34 |
| − head | 0.03 | −0.02 |
| Ball | 0.48 |
Now the open question from the warm-up. With + foot alone, the similarity to shoe was 0.51 and shoe was 5th. Subtracting head lowered the similarity to 0.48, yet shoe went to 1st. Head and shoe are barely linked, so subtracting it cost almost nothing; what changed was the competition. Before, "toe", "toes", and "beanie" sat nearer the ball than shoe. Afterward, none did. Rank counts how many words beat the jack, not just how close the jack is.
Why did those words fall back? The panel can't say. It can give the arithmetic exactly, but not why the model put "toe" where it did.
Student task (pairs, Puzzles, common-sense words)
Pick one: Weather Shift (hot → cold, medium), Day and Night (day → night, easy), or Frankenstein's Monster (dead → alive, hard). After each throw, open "Why?" and write down:
- Which word had the biggest share?
- Did a word you added barely point at the jack? Did subtracting a word cost you? (The panel flags this.)
- One thing the panel can't tell you.
All three puzzles are opposites, which is the twist. On this map, cold is the 16th-nearest word to hot, and alive is the 11th-nearest word to dead. In Weather Shift, adding warm moves the ball toward cold (warm and cold have a similarity of 0.62). Why might opposites sit close together? One answer: they turn up in the same kinds of sentences ("the water was ___").
Activity 3: two maps of meaning (10 minutes)
The game has two word sets. The idea is the same (each word is a list of numbers); the training data differs.
- Common sense (the default): ConceptNet Numberbatch 19.08, which blends text statistics with ConceptNet, a large collection of everyday facts people have written down ("a hen lays eggs", "bacon is a kind of meat"). 21,114 everyday words, 300 numbers each.
- Raw text: GloVe (2014), learned only from which words appear near each other in about six billion words of Wikipedia and news. About 40,000 words, 100 numbers each.
Here are the nearest words in each, taken from the game's own data:
| Word | Raw text | Common sense |
|---|---|---|
| orange | yellow, red, blue, green, pink | oranges, citrus, tangerine, yellow, purple |
| apple | microsoft, ibm, software, intel, dell | apples, orchards, ipad, iphone, pears |
| ham | sunderland, fulham, middlesbrough, wigan, burnley | bacon, pork, sausage, meats, sausages |
| cold | warm, dry, hot, cool, chilly | chilly, frigid, colds, frosty, colder |
| eggs | egg, fertilized, nests, nest, cooked | egg, hens, yolks, chickens, yolk |
The ham row is English football: news text mentions West Ham alongside Fulham, Sunderland, and other clubs. And in the raw-text set, "meat" is the 7th-nearest word to eggs, while "bacon" is 163rd. One guess is that news stories mention eggs alongside meat, milk, and fish (in prices, farming, food safety) more often than at breakfast. That's a hypothesis, though, and the numbers alone can't confirm it.
Student task
Play Fruit Transform (apple → orange) or Lock and Key (lock → key) once with each word set. Compare how far away the jack starts, what par is, and the list of words "closest in meaning" to the jack at the end. Then discuss:
- Which map is right? (Neither: they were trained on different data.)
- Where did the raw-text map learn that orange is mainly a color, and that key mostly means "important"?
- What would a map trained on recipes, sports pages, or your group chat look like?
Extension: a chatbot's own map
The game also has a third word set, AI tokens: GPT-2's token table, the list of numbers the model looks up for each token before anything else happens (49,745 tokens, reduced to 128 numbers each). Switch to it and look at the neighbours: a word's nearest "words" are often its own capitalised or spaced variants (" Shoe", "Shoe") and word pieces, because a language model reads tokens, not words. The Tokens tab shows GPT-2's tokenizer splitting any text you type, with a short "guess the split" quiz.
Discussion questions
What embeddings capture, and what they miss
- The raw-text map puts "cold" next to "warm" and "hot". Is that a bug? What does it say about learning meaning only from which words appear together?
- "Ham" gets one point on the map, whether it means meat or a football club. What goes wrong when a word has two meanings, and how might a model handle it?
- What can't be captured by one point per word? Try word order ("dog bites man"), negation, or sarcasm.
Bias
- Embeddings learned from text pick up the associations in that text, including stereotypes about gender, race, and jobs (Bolukbasi et al., 2016; Caliskan et al., 2017). Where could that show up in a game like this?
- The game keeps slurs and profanity out of its vocabulary with a blocklist. Does removing words remove the associations? (The remaining words were still placed using the same text.)
- Nissim, van Noord & van der Goot (2020) argue that some well-known analogy results, including biased-looking ones, were partly produced by how analogies are scored. How would you test a claim like "man is to doctor as woman is to nurse" fairly?
Attribution isn't explanation
- The "Why?" bars are exact: they add up to the similarity. So what don't they tell you?
- If someone showed you an exact breakdown of an AI model's output, what more would you want before saying you know why the model did it?
Fixed versus contextual
- In this game "bank" has one position. Inside a chatbot, the representation of "bank" changes with the sentence around it. What does that gain, and why does it make the model harder to inspect?
Honest caveats
These are worth saying out loud in class.
These are fixed ("static") embeddings, not a transformer's internals. Each word or token gets one list of numbers, from GloVe (2014), Numberbatch 19.08 (2019), or GPT-2's token table (2019). The GPT-2 map is the model's first step only: everything after it, across many layers, builds representations that change with the sentence. The game is a warm-up for those ideas, not a look inside a chatbot.
- The "Why?" panel is a reading of the numbers, not the model's reasons. The shares are exact arithmetic; the plain-language notes are guesses from those numbers, and the panel says so.
- Rank leaves out the words you threw. Otherwise the ball would often be nearest to its own start word or a tile: the ball from hat + foot − head is actually nearest to foot, then hat, then shoe. Classic analogy tests use the same rule (the ball from king − man + woman is nearest to king), and Nissim et al. show that it can flatter results.
- The common-sense courts are chosen to be explainable. A court is only dealt if one of its strongest throws can be explained word by word, and par is picked among such throws. So that set looks tidier than the raw embedding would on its own.
- Ranks aren't comparable across word sets: one has 21,114 words, the other about 40,000.
- Precision. To fit in a browser, the vectors are stored at 8-bit precision, so the numbers can differ slightly from the published files.
- Online rooms are new. Versus → "Play online with friends" takes up to six players per room, but it hasn't been tried with a whole class. Treat it as a bonus, with Daily as the fallback.
Going further
- The slides explain word vectors and vector arithmetic, and compare plain vector arithmetic with the "3CosAdd" analogy method. They were written for an earlier, multiplayer version of the game, so some details of play are out of date.
- In the game, "How to play" → "Why don't throws always make sense?" opens a short note on why the game can't fully explain itself, with further reading on interpretability.
- For a programming class: the source code is plain JavaScript. The engine (
web/engine.js) is about 300 lines, andtests/engine.test.jschecks it against the real vectors. One exercise: find where rank is computed, change it to count the thrown words, and see which puzzles survive. - Reading: Mikolov, Yih & Zweig (2013), the king − man + woman paper; GloVe (Pennington, Socher & Manning, 2014); ConceptNet 5.5 (Speer, Chin & Havasi, 2017); Mapping the mind of a large language model (Anthropic, 2024); Zoom In: an introduction to circuits (Olah et al., 2020).