LLMs Explained: How Modern Language Models Work on the Inside


LLMs: how modern language models work on the inside

The real mechanism behind GPT, Claude, Gemini, and Llama

How modern models turn text into tokens, predict the future, and generate coherent language.

So far we’ve traced the whole evolutionary path:

Every piece brought us closer to an inevitable question, the one anyone using ChatGPT for the first time really wants answered:

What exactly happens inside the model when you type a sentence?

The answer is almost paradoxical: at its most intimate core, the idea is surprisingly simple; but the engineering needed to make it work is colossal.

An LLM does only one thing: predict the next token.

It doesn’t think like a human, nor does it query an internal database. It’s a massive statistical system that computes probabilities. For that prediction to look like genuine reasoning, it relies on a precise chain of gears: tokenization, embeddings, attention, thousands of layers, and billions of parameters. Let’s open them up one by one.

Diagram showing tokenization, embeddings, Transformer layers, and next‑token prediction inside an LLM.


1. What is an LLM? (And what a «token» really is)

An LLM (Large Language Model) is a giant neural network trained for a single task: given a preceding sequence of text, decide which is the most logical piece that should come next.

That «piece» isn’t always a whole word: it’s a token.

  • Common words are usually a single token:"cat","sun".
  • Long or technical words get split into chunks:"transformers"→["trans", "former", "s"].
  • In code or in languages with little online presence, a token can be just a couple of stray characters.

The LEGO analogy. The model doesn’t speak with words; it speaks with pieces of linguistic LEGO. It has a fixed drawer of pieces (its vocabulary, typically between 32,000 and 128,000) and builds any sentence by combining them. That’s why it can handle rare words, typos, and multiple languages: it always finds a combination of pieces that fits.

Those billions (or trillions) of parameters are the compressed memory where the language patterns, facts, and syntactic structures it learned are encoded.

As a rule of thumb: 100 tokens ≈ 75 words in English.


2. From text to math: tokenization and embeddings

Computers don’t understand letters; they only operate with numbers. The input goes through two chained stages:

  1. Tokenization (BPE or SentencePiece): breaks the text into unique numeric identifiers drawn from the fixed vocabulary.
  2. Embeddings: each ID is transformed into a vector (a list of numbers) in a space of thousands of dimensions.

As we saw in the previous chapter, these are not static Word2Vec-style vectors. They are contextual embeddings: the vector for «bank» is mathematically repositioned according to its neighbors (river vs. financial interest). Embeddings are, literally, the geometry of meaning.


3. The Transformer: the internal engine

Virtually every current LLM (GPT-4, Claude, Llama) is built on the Transformer architecture, in its Decoder-only variant: they process text causally, left to right, in order to continue the sentence.

Why did it win out over RNNs? Because the Transformer can:

  • process all tokens in parallel (instead of one by one),
  • connect any pair of words in a single step,
  • and capture long dependencies without losing the signal along the way.

Inside each layer, five key pieces are at work, and each layer delivers a richer version of the meaning than the previous one:

Component What it does Analogy
Self-Attention Decides which tokens look at which A reader underlining the connected words
Multi-Head Attention Repeats that gaze from dozens of angles Several experts reading the same sentence
Feed-Forward (FFN) Digests and stores facts and associations The model’s internal «library»
LayerNorm Keeps the numbers stable across layers The stabilizer that stops everything from exploding
Residual Connections Highways so the signal doesn’t degrade A shortcut that jumps from layer to layer

4. Self-Attention and Multi-Head: who looks at whom

The heart of the Transformer is attention: each token asks itself «which other words do I need to look at in order to understand myself?».

«The animal didn’t cross the street because it was tired.»

The word «it» directs its attention toward «animal» and ignores «street.» That’s how it resolves pronouns, ambiguities, agreement, and long dependencies without a single hand-coded rule.

Multi-Head Attention takes this further: instead of a single global gaze, it launches dozens of «heads» in parallel, each specialized:

  • one watches syntax (subject–verb),
  • another resolves pronouns,
  • another detects entities and proper nouns,
  • another follows verb tense or tone.

Then they’re all combined into a single, enriched representation. It’s like having a committee of expert readers analyzing the same sentence at once.


5. The fundamental principle: predicting the next token

Here’s the blunt truth, and it’s worth stating without hedging:

An LLM doesn’t «think,» doesn’t «understand,» and doesn’t «reason» like a human. It only predicts the most probable next token.

When the prompt crosses all the layers, the model doesn’t spit out a closed answer: it generates a probability table over its entire vocabulary.

Candidate token Estimated probability
"bright" 64.2%
"blue" 21.5%
"dark" 8.1%
… (rest of ~100k tokens) < 0.01% each

What’s fascinating is the consequence: to nail the next token across millions of different contexts (Python code, procedural law, medicine, poetry), the network is forced to internally compress deep logical, syntactic, and factual patterns. Massive prediction, at sufficient scale, gives rise to an apparent understanding.


6. How an LLM is built: Pretraining vs. Post-training

A model isn’t born knowing how to respond kindly. Its creation has two radically different phases:

  1. Pretraining. The network is fed trillions of tokens from the web, books, and code. The loop is simple: it receives a sequence, tries to predict the next token, compares with the real one, and adjusts its parameters through gradient descent. Repeated billions of times. The result is a base model: a pure completer. If you type «What is the capital of France?», it might not answer «Paris,» but instead continue with «And the one in Italy? And in Spain?», because it’s imitating a quiz it saw on the internet.
  2. Fine-tuning + RLHF. The base model is refined with model conversations (SFT) and then with reinforcement learning guided by humans (RLHF, Reinforcement Learning from Human Feedback). Here the statistical completer learns its role as an assistant: following instructions, being helpful, staying coherent, avoiding harmful responses, and admitting when it doesn’t know something.

7. Inference and sampling: why it doesn’t always say the same thing

When you use the model (inference), text is produced autoregressively: each chosen token is glued to the end and fed back in to compute the next one.

Important: the model generates token by token, not word by word. And it doesn’t always pick the number-one token. To regulate its behavior, sampling hyperparameters are used:

  • Temperature: near $0$ it becomes deterministic and conservative (ideal for math or data extraction); between $0.7$ and $1.2$ it flattens the differences and gains variety and creativity.
  • Top-k: restricts the choice to the $k$ most probable tokens (for example, the top 40).
  • Top-p (Nucleus Sampling): picks the cumulative group of tokens whose probability sums to a threshold (for example, the top 90%), adapting dynamically to the model’s certainty.

These three dials control the balance between coherence, diversity, creativity, and style.


8. Fundamental limits: what an LLM is NOT

Understanding the internal mechanism immediately explains its typical failures. They are not databases; they are statistical models, and that’s where their weaknesses come from:

Limitation Internal technical cause
Hallucinations It prioritizes the statistical plausibility of language over empirical verification of the facts.
Prompt sensitivity Small changes at the start alter the attention weights in early layers and shift the entire probabilistic trajectory.
No real memory It doesn’t «remember» continuously; it only processes the context window that fits in the current prompt.
Inherited biases If a bias is in the internet corpus, the model encodes it in its weight matrices.

9. The current ecosystem: beyond pure text

LLMs have stopped being simple predictive text boxes and have become software orchestrators and the universal interface between humans and machines:

  • RAG (Retrieval-Augmented Generation): connecting the model to vector search engines so it consults verified sources before answering.
  • Function calling and agents: the model predicts API calls or code snippets that an external environment executes to interact with databases and tools.
  • Native multimodality: processing text tokens, audio, and visual fragments (patches) within the same Transformer vector space.

On this foundation run today’s intelligent assistants, autonomous agents, document analysis, code generation, and semantic search.


In summary

Large Language Models operate under a systematic architecture:

  1. They fragment language into numeric units called tokens.
  2. They map those tokens into vectors with dynamic context (embeddings).
  3. They weigh relationships through multi-head attention across deep Transformer layers.
  4. They compute probabilities for the next token from massive patterns learned on the internet.
  5. They shape their behavior through supervised alignment (SFT) and human evaluation (RLHF).
  6. They generate text token by token, autoregressively.

They are not thinking minds or infallible oracles: they are statistical calculators on a colossal scale. Their historic achievement is that, at sufficient scale, predicting language with mathematical precision becomes indistinguishable from understanding it.