Word Bocce

for teachers

A lesson on word embeddings, played as a game

Word Bocce is a free browser game played on a map of word meanings. Every word is a point; you add and subtract words to roll your ball toward a target word, the jack. It's a hands-on way into word embeddings, and into why it's hard to explain what a learned model is doing.

This page is a lesson plan for 30 to 45 minutes, written for intro NLP, machine learning, and data-science classes. It also works for curious high-school classes if you skip the formula in Activity 2.

Time30–45 minutes. For 30, drop Activity 3 or show it as a quick demo.
LevelIntro undergraduate, or interested high-school students.
You needA projector, and a phone or laptop per student or pair. No accounts or installs.
LinksThe game · slides · source

Learning goals

By the end, students should be able to:

  1. Explain that a word embedding stores each word as a list of numbers, and that "close in meaning" is measured by cosine similarity.
  2. Do word arithmetic (a start word plus some words, minus others) and reason about where the result lands, including why rank and similarity can disagree.
  3. Describe how training data shapes the map, by comparing an embedding learned only from text with one that also uses a knowledge base of everyday facts.
  4. Tell the difference between an exact breakdown of a score and an explanation of why a model learned what it did.
  5. Name some limits: one point per word (so different senses get mixed), associations inherited from the training text, and how fixed word vectors differ from what's inside a modern language model.

Before class

Lesson outline

MinutesActivityIn the game
0–5Warm-up: word arithmeticTutorial, on the projector
5–15Activity 1: today's courtDaily
15–27Activity 2: reading the "Why?" panelPuzzles
27–37Activity 3: two maps of meaningWords switch, Puzzles
37–45Discussion

Warm-up: word arithmetic (5 minutes)

On the board, give four words two made-up coordinates:

WordRoyalFemale
man00
woman01
king10
queen11

Then king − man + woman = (1, 0) − (0, 0) + (0, 1) = (1, 1), which is queen.

Real word embeddings work the same way, with two differences. There are hundreds of coordinates (100 per word in the game's raw-text set, 300 in the common-sense set). And nobody chose what they mean: they're learned from data, and no single number means "royal".

Closeness is measured with cosine similarity: 1 means two vectors point the same way, 0 means unrelated. The game also shows rank, which is easier to read: #1 means the jack is the nearest word to where your ball stopped.

Now ask the class to predict hat + foot − head, and play the tutorial on the projector:

Leave one question open for Activity 2: why did taking away "head" help?

Activity 1: today's court (10 minutes)

  1. Everyone opens Daily. They all get the same start word, jack, and nine tiles.
  2. In pairs, before throwing, sort the nine tiles into three piles: words that pull toward the jack, words worth subtracting, and traps. (The deal is built this way: two tiles pull toward the jack, two carry the start word's flavor and are worth subtracting, two are linked to both, and three look related to the jack but pull only weakly.)
  3. Each player gets four balls, up to three tiles per throw, and the best ball counts.
  4. At the end, the game shows par (the best throw the hand allowed), a few nearby words on the court, and the jack's nearest words. "Copy result" gives a short summary students can paste into the class chat.

Debrief: Who matched par? Which tile was the best trap? Did anyone's "obvious" throw go the wrong way, and can they say why?

Activity 2: reading the "Why?" panel (10–12 minutes)

After every throw there's a Why? link. It splits the ball's similarity to the jack into one share per word. The shares add up exactly, because the ball is just the sum of the word vectors (each of length 1), rescaled:

cos(ball, jack) = Σ sign × cos(word, jack) ÷ |Σ sign × word|

The start word counts as a + word. Here is the tutorial throw, with numbers from the game:

WordSimilarity to shoeShare
hat (start)0.230.16
+ foot0.500.34
− head0.03−0.02
Ball0.48

Now the open question from the warm-up. With + foot alone, the similarity to shoe was 0.51 and shoe was 5th. Subtracting head lowered the similarity to 0.48, yet shoe went to 1st. Head and shoe are barely linked, so subtracting it cost almost nothing; what changed was the competition. Before, "toe", "toes", and "beanie" sat nearer the ball than shoe. Afterward, none did. Rank counts how many words beat the jack, not just how close the jack is.

Why did those words fall back? The panel can't say. It can give the arithmetic exactly, but not why the model put "toe" where it did.

Student task (pairs, Puzzles, common-sense words)

Pick one: Weather Shift (hot → cold, medium), Day and Night (day → night, easy), or Frankenstein's Monster (dead → alive, hard). After each throw, open "Why?" and write down:

  1. Which word had the biggest share?
  2. Did a word you added barely point at the jack? Did subtracting a word cost you? (The panel flags this.)
  3. One thing the panel can't tell you.

All three puzzles are opposites, which is the twist. On this map, cold is the 16th-nearest word to hot, and alive is the 11th-nearest word to dead. In Weather Shift, adding warm moves the ball toward cold (warm and cold have a similarity of 0.62). Why might opposites sit close together? One answer: they turn up in the same kinds of sentences ("the water was ___").

Activity 3: two maps of meaning (10 minutes)

The game has two word sets. The idea is the same (each word is a list of numbers); the training data differs.

Here are the nearest words in each, taken from the game's own data:

WordRaw textCommon sense
orangeyellow, red, blue, green, pinkoranges, citrus, tangerine, yellow, purple
applemicrosoft, ibm, software, intel, dellapples, orchards, ipad, iphone, pears
hamsunderland, fulham, middlesbrough, wigan, burnleybacon, pork, sausage, meats, sausages
coldwarm, dry, hot, cool, chillychilly, frigid, colds, frosty, colder
eggsegg, fertilized, nests, nest, cookedegg, hens, yolks, chickens, yolk

The ham row is English football: news text mentions West Ham alongside Fulham, Sunderland, and other clubs. And in the raw-text set, "meat" is the 7th-nearest word to eggs, while "bacon" is 163rd. One guess is that news stories mention eggs alongside meat, milk, and fish (in prices, farming, food safety) more often than at breakfast. That's a hypothesis, though, and the numbers alone can't confirm it.

Student task

Play Fruit Transform (apple → orange) or Lock and Key (lock → key) once with each word set. Compare how far away the jack starts, what par is, and the list of words "closest in meaning" to the jack at the end. Then discuss:

Extension: a chatbot's own map

The game also has a third word set, AI tokens: GPT-2's token table, the list of numbers the model looks up for each token before anything else happens (49,745 tokens, reduced to 128 numbers each). Switch to it and look at the neighbours: a word's nearest "words" are often its own capitalised or spaced variants (" Shoe", "Shoe") and word pieces, because a language model reads tokens, not words. The Tokens tab shows GPT-2's tokenizer splitting any text you type, with a short "guess the split" quiz.

Discussion questions

What embeddings capture, and what they miss

Bias

Attribution isn't explanation

Fixed versus contextual

Honest caveats

These are worth saying out loud in class.

These are fixed ("static") embeddings, not a transformer's internals. Each word or token gets one list of numbers, from GloVe (2014), Numberbatch 19.08 (2019), or GPT-2's token table (2019). The GPT-2 map is the model's first step only: everything after it, across many layers, builds representations that change with the sentence. The game is a warm-up for those ideas, not a look inside a chatbot.

Going further