N-gram Repetition Penalty
Confiance : high
repetition-penaltydoom-loopingsmall-modelsgeneration-controlliquid-aipost-training
A technique used during reinforcement learning training to prevent doom-looping in small language models. Applied alongside RL training to discourage repetitive n-gram patterns that lead to infinite loops or stuck generation states.
Implementation
The penalty is applied during Stage 2 of liquid-ai's training pipeline:
- Combined with reinforcement learning objectives
- Targets specific n-gram patterns that indicate repetitive behavior
- Helps maintain generation diversity while preserving task performance
Context
Part of the comprehensive solution to doom-looping-problem in small-language-models with reasoning traces. Used in conjunction with on-policy data generation and careful preference alignment to maintain model quality while eliminating pathological repetitive behaviors.
Effectiveness
Successfully reduced doom loop ratios in LFM2.5-1.2B-Thinking model, enabling reliable on-device reasoning without getting stuck in repetitive generation patterns.