Contextual Embeddings Explained: How Transformers Represent Meaning in Modern NLP


Contextual embeddings

How models learn meaning and why Transformers changed the way we represent language

From static vectors to dynamic representations that understand context.

So far we’ve toured the whole machinery: the RNNs that tried to remember, the gates of LSTM/GRU, the attention mechanism, and finally the Transformers that process everything in parallel.

But there’s a foundational question we’ve been dodging for several chapters, and it’s the most basic one of all:

How does a machine turn language into math without throwing its meaning in the trash?

Before a Transformer can attend, compare, generate, or reason, it needs to translate each word into numbers. That’s where embeddings come in, and above all the leap that changed everything: contextual embeddings.

Diagram comparing static embeddings vs contextual embeddings, showing vectors shifting based on sentence context.


What is an embedding? The geometry of meaning

An embedding is a vector: a list of numbers (usually hundreds or thousands of positions) that represents the meaning of a word, a phrase, or an entire document.

But it’s not just any list, nor a simple code. It’s a geometric map of language.

The city-map analogy. Picture a giant map where each word is a city. Cities with a similar «climate» end up close to one another: dog and cat are almost neighbors; dog and microchip are on different continents. And here’s the fascinating part: the distances and directions carry meaning of their own. The arrow that goes from man to woman points in exactly the same direction as the one from king to queen.

It’s not magic or a trick: it’s well-calibrated semantic geometry. You subtract «masculinity,» add «femininity,» and you move to the correct point on the map.


Static embeddings: Word2Vec, GloVe, and FastText

Before Transformers, embeddings had an original sin: they were static.

One word = one single vector. Always. No matter the sentence.

Think of the word «bank»:

  • «I deposited my savings at the bank.»
  • «I sat on the bank by the plaza.» (a bench)
  • «A bank of fish crossed the reef.»

In a static model, bank has exactly the same coordinates in all three cases. Its vector is a forced average of all its uses: a fuzzy blend of finance, furniture, and marine biology. The model couldn’t disambiguate because it consulted a fixed dictionary, not the actual situation.

Static model How it learned Key limitation
Word2Vec Words appearing in similar contexts have similar meanings One vector per word for the whole corpus
GloVe Global co-occurrence statistics Fixed; blind to polysemy
FastText Word pieces (subwords) Handles typos and rare words, but still static

They were enormous advances, the foundation of everything that followed. But they crashed into the same wall: they couldn’t read the sentence.


The big leap: contextual embeddings

Transformers brought a radical idea that, deep down, is very human:

The meaning of a word doesn’t live in its dictionary entry; it emerges from the words that surround it.

Therefore, the embedding can’t be precooked: it has to be made live, every time.

Now «bank» in «I deposited money at the bank» and «bank» in «The fisherman sat by the river bank» have completely different vectors. The model no longer memorizes meanings: it builds them in real time.

 

The word stops being a motionless point and becomes a particle that shifts toward the finance zone or the geography zone depending on its company.


How is a contextual embedding built?

Unlike Word2Vec (where a lookup in a fixed table was enough), here the embedding is the result of a layer-by-layer refinement process:


Contextual embedding

  1. Initial embedding: the word enters with a generic starting vector, its «neutral self.»
  2. Positional Encoding: positional information is added (being at the start of the sentence isn’t the same as being at the end).
  3. Self-Attention: each token «converses» with all the others and decides which ones to look at.
  4. Multi-Head and Feed-Forward: in each block, the vector absorbs syntactic and semantic nuances.

The vector that comes out at the end is no longer the isolated word: it’s the word steeped in the whole sentence.


An intuitive example: the pronoun test

The clearest case shows up when resolving pronouns, exactly where static models crashed:

Sentence A: «The animal didn’t cross the street because it was tired.»

The word «it» directs its attention toward «animal,» ignores «street,» and adjusts its vector to reflect that it = animal (something alive, that gets tired).

Sentence B: «The street was closed because it was under construction.»

Now «it» points to «street,» and its embedding changes completely (something physical, under repair).

Same word, two radically different coordinates. The model doesn’t «understand» like a human, but its vectors adapt with surgical precision to the logic of the sentence.


Why were they a before and after?

Contextual embeddings didn’t improve one thing; they unlocked several at once.

  • Automatic disambiguation. A word no longer has a fixed meaning; the model builds it according to the sentence.
  • Deep representations. Each Transformer layer produces a richer, more nuanced version of the embedding, from syntax to pure semantics.
  • Massive transfer learning. Models like BERT showed that, once trained on giant corpora, those embeddings can be reused for classification, NER, question answering, sentiment analysis, or semantic search, without starting from scratch each time.
  • The foundation of modern LLMs. GPT, Claude, Gemini, Llama… all depend on contextual embeddings to generate coherent text.

From words to sentences, code, and images

By solving the dynamic-context problem, the technique stopped being limited to single words. Today we generate representations of much broader units:

Type of embedding What it represents What it’s for
Word A token in its context Fine-grained language understanding
Sentence / Document A whole paragraph condensed Semantic search, RAG, clustering
Code Functions and logic blocks (CodeBERT, StarCoder) Code search and autocompletion
Multimodal Text + image + audio together (CLIP, GPT-4o) Connecting a photo to its description

The multimodal example is the most striking: a photograph of a sunset and the phrase «reddish sunset» end up sharing very close vector coordinates, even though one is pixels and the other is letters.

 

They all share the same underlying idea: representing meaning as geometry in a latent space.


In summary

Contextual embeddings represent the maturity of computational linguistics:

  • They replaced static vectors with dynamic meaning that changes according to the sentence.
  • They are the direct product of attention: they wouldn’t exist without the self-attention that computes global relationships.
  • They enable deep understanding in Transformers, layer by layer.
  • They are the hidden engine of semantic search, RAG, and LLM reasoning.
  • They unified the modalities: text, code, audio, and image speaking the same geometric language.

It was the exact moment when language stopped being a static list of words… and became a living mathematical space, where meaning can be measured, compared, and manipulated in real time.