Vanishing and Exploding Gradients + Beam Search: How Early Neural Networks Learned and Generated Text Before Transformers
A clear and intuitive explanation of vanishing and exploding gradients in deep networks, the techniques that stabilized training, and how Beam Search enabled coherent text generation before attention and Transformers.



