Tag: model compression

Stochastic Depth and Regularization in Deep Transformer LLMs: A Practical Guide 21 August 2026

Stochastic Depth and Regularization in Deep Transformer LLMs: A Practical Guide

Learn how stochastic depth improves deep transformer LLMs by reducing overfitting and boosting efficiency. Discover practical tips for drop schedules and combining it with other regularization techniques.

Susannah Greenwood 7 Comments
Production Guardrails for Compressed LLMs: Confidence and Abstention 21 June 2026

Production Guardrails for Compressed LLMs: Confidence and Abstention

Learn how production guardrails for compressed LLMs use confidence scores and abstention to balance safety and speed. Explore Defensive M2S, efficiency techniques, and implementation strategies.

Susannah Greenwood 0 Comments
How to Reduce Memory Footprint for Hosting Multiple Large Language Models 24 October 2025

How to Reduce Memory Footprint for Hosting Multiple Large Language Models

Learn how to reduce memory footprint for hosting multiple large language models using quantization, model parallelism, and hybrid techniques. Cut costs, run more models on less hardware, and avoid common pitfalls.

Susannah Greenwood 0 Comments