~/wiki

Diffusion Text Models

Mis à jour le 2025-01-03Confiance : high
diffusion-text-modelsnon-autoregressivetext-generationiterative-refinementgoogle-diffusiongemmaparallel-generationvllm-supportinference-optimizationblock-denoisingapache-licensemoe-architectureresearch-directions1000-tokens-per-second4x-speedup256-token-blocksllama-cpp-supportlocal-execution

Alternative text generation paradigm that generates and refines blocks of text simultaneously rather than sequentially predicting next tokens. Represents departure from standard autoregressive language models with potential for significantly faster inference and novel capabilities.

Core Architecture

Block-Based Generation: Instead of autoregressive next-token prediction, diffusion text models generate and iteratively refine entire blocks of text simultaneously. google-diffusiongemma uses 256-token blocks with denoising processes to achieve up to 4x speedup over traditional approaches.

Iterative Refinement: Models start with noisy text blocks and progressively denoise them through multiple passes, similar to image diffusion models but adapted for discrete text tokens. This enables parallel processing across token positions within blocks.

Non-Sequential Decoding: Unlike autoregressive models that must generate tokens in strict left-to-right order, diffusion approaches can modify any position within a block during refinement, enabling more flexible editing and correction.

Performance Characteristics

Throughput Gains: google-diffusiongemma achieves 1000+ tokens per second on suitable hardware with reported 4x speedup compared to autoregressive baselines. Performance particularly strong on hardware optimized for parallel computation.

Serving Infrastructure: First diffusion LLM natively supported in vLLM achieving 1200+ output tokens/second at batch size 1 on H200 with FP8 quantization. Also supports local execution on consumer hardware (18GB+ VRAM).

Resource Requirements: google-diffusiongemma's 26B MoE architecture requires only 3.8B active parameters during inference, making it more accessible than full-scale models while maintaining competitive performance.

Research Directions

Constrained Generation: Diffusion approach enables novel capabilities including fill-in-the-middle generation, structured editing, and error correction that are difficult with autoregressive models.

Quality vs Speed Tradeoffs: Iterative refinement allows trading inference time for output quality by adjusting number of denoising steps, providing flexible deployment options.

Hybrid Architectures: Potential for combining autoregressive and diffusion approaches for different parts of generation pipeline, optimizing each component for its strengths.

Production Readiness

Open Source Availability: google-diffusiongemma released under Apache 2.0 license with open weights, enabling research and commercial deployment without licensing restrictions.

Inference Support: Native vLLM integration and llama.cpp compatibility provide production-ready serving infrastructure. GGUF quantization enables deployment on consumer hardware.

Experimental Status: Still early in development cycle compared to mature autoregressive models, but rapid progress in systems integration suggests near-term production viability.

See also