Review of RNN

Drawback of RNN
- RNN processes tokens one at a time — each step depends on the previous hidden state.
- This makes training slow and hard to parallelize on GPUs.
- RNN compresses all past information into a single hidden state at each step.
- ❌ May lose or distort important information over time.
Transformer
- Transformer allows full parallel computation across all tokens.
- ✅ Training is much faster and efficient on modern hardware.
- Transformers use attention over the full sequence, so each token can directly attend to any other token.
- ✅ Much better at modeling long contexts and dependencies.
- Context window.
Tutorial
Step by Step into Transformer
Code
Let's build GPT: from scratch, in code, spelled out.
Given a training set $(x_1, x_2, \ldots,x_T)$.
$\max_\theta\Pi_{t=1}^{T}p_\theta(x_t|x_1, \ldots,x_{t-1})$