~/wiki

N-gram Repetition Penalty

Confiance : high
repetition-penaltydoom-loopingsmall-modelsgeneration-controlliquid-aipost-training

A technique used during reinforcement learning training to prevent doom-looping in small language models. Applied alongside RL training to discourage repetitive n-gram patterns that lead to infinite loops or stuck generation states.

Implementation

The penalty is applied during Stage 2 of liquid-ai's training pipeline:

  • Combined with reinforcement learning objectives
  • Targets specific n-gram patterns that indicate repetitive behavior
  • Helps maintain generation diversity while preserving task performance

Context

Part of the comprehensive solution to doom-looping-problem in small-language-models with reasoning traces. Used in conjunction with on-policy data generation and careful preference alignment to maintain model quality while eliminating pathological repetitive behaviors.

Effectiveness

Successfully reduced doom loop ratios in LFM2.5-1.2B-Thinking model, enabling reliable on-device reasoning without getting stuck in repetitive generation patterns.

See also