Tag: LLM training

Monitoring Loss and Perplexity: Reading Signals During LLM Training 26 September 2026

Monitoring Loss and Perplexity: Reading Signals During LLM Training

Learn how to interpret cross-entropy loss and perplexity during LLM training. Discover practical tips for reading training curves, avoiding common pitfalls, and diagnosing model health effectively.

Susannah Greenwood 5 Comments
Stochastic Depth and Regularization in Deep Transformer LLMs: A Practical Guide 21 August 2026

Stochastic Depth and Regularization in Deep Transformer LLMs: A Practical Guide

Learn how stochastic depth improves deep transformer LLMs by reducing overfitting and boosting efficiency. Discover practical tips for drop schedules and combining it with other regularization techniques.

Susannah Greenwood 7 Comments
Residual Connections and Layer Normalization in Large Language Models: Why They Keep Training Stable 2 January 2026

Residual Connections and Layer Normalization in Large Language Models: Why They Keep Training Stable

Residual connections and layer normalization are essential for training stable, deep large language models. Without them, transformers couldn't scale beyond a few layers. Here's how they work and why they're non-negotiable in modern AI.

Susannah Greenwood 7 Comments
Mixed-Precision Training for Large Language Models: FP16, BF16, and Beyond 16 December 2025

Mixed-Precision Training for Large Language Models: FP16, BF16, and Beyond

Mixed-precision training using FP16 and BF16 cuts LLM training time by up to 70% and reduces memory use by half. Learn how it works, why BF16 is now preferred over FP16, and how to implement it safely with PyTorch.

Susannah Greenwood 8 Comments