AI Briefing
KO

Integrated Expansion-Pruning Pipeline for Pre-training Generative Language Models: IDEA Prune

·2026.08.26 09:00

Key point

The IDEA Prune pipeline, which integrates pre-training of expanded models with pruning, improved LLM compression performance.

Details

Structured Pruning pipelines are gaining attention for the efficient deployment of Large Language Models (LLMs). This work explores whether integrating pre-training of expanded models, often overlooked in previous research, into the pruning process is valuable, and how the entire pipeline can be optimized.

The proposed IDEA Prune integrates the processes of expanded model training, pruning, and recovery under a single Cosine Annealing learning rate schedule. It also introduces a novel iterative structured pruning method for gradual parameter removal. This approach mitigates knowledge loss caused by learning rate increases in existing methods and effectively redistributes model capacity among surviving neurons, achieving smooth compression and performance improvements.

In experiments compressing a 2.8B model down to a 1.3B model, up to 2T tokens of pre-training data were utilized. The integrated approach provided insights into the token efficiency of expanded model pre-training and demonstrated the superiority of the pruned models.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.