TLDRocket
Sign in

What exactly does word2vec learn?

BAIR

Researchers developed a closed-form mathematical theory explaining how word2vec learns word embeddings, proving that the algorithm reduces to PCA on a matrix defined by corpus co-occurrence statistics. Word2vec achieves 68% accuracy on analogy completion benchmarks while learning discrete, sequential linear concepts that correspond to interpretable topics. The theory enables predicting learned features beforehand from corpus statistics and provides insights into how neural language models develop meaningful representations.

Why it matters

What exactly does word2vec learn, and how? Answering this question amounts to understanding representation learning in a minimal yet interesting language modeling task. Despite the fact that word2vec is a well-known precursor to modern language models, for many years, researchers lacked a quantitative and predictive theory describing its learning process. In our new paper, we finally provide such a theory. We prove that there are realistic, practical regimes in which the learning problem reduces to unweighted least-squares matrix factorization. We solve the gradient flow dynamics in closed form; the final learned representations are simply given by PCA. Learning dynamics of word2vec. When trained from small initialization, word2vec learns in discrete, sequential steps. Left: rank-incrementing learning steps in the weight matrix, each decreasing the loss. Right: three time slices of the latent embedding space showing how embedding vectors expand into subspaces of increasing dimension at

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.