~/wiki

Three-Stage Pipeline

Confiance : high
three-stage-pipelineopenWakeWordaudio-processingmel-spectrogramembedding-modelsonnx-runtimewake-word-detectiontensor-processingreal-time-inferencerolling-buffersstreaming-audio

Audio processing architecture used in openwakeword systems that chains together three specialized ONNX models to transform raw audio into wake-word detection decisions. This design separates concerns and enables efficient real-time inference through progressive feature extraction.

Pipeline Architecture

Stage 1: Mel Spectrogram Conversion

  • Input: Raw PCM audio samples (1280 samples = 80ms at 16kHz)
  • Output: Mel-scale spectrogram with 32 frequency bins
  • Function: Converts time-domain audio to frequency-domain representation
  • Tensor Shape: [T, 32] where T varies based on audio chunk size

Stage 2: Feature Embedding

  • Input: Sliding window of 76 mel spectrogram frames
  • Output: 96-dimensional feature embedding vector
  • Function: Extracts semantic audio features from spectral data
  • Tensor Shape: [1, 76, 32, 1] → [96]

Stage 3: Wake-Word Classification

  • Input: Buffer of last 16 feature embeddings
  • Output: Single wake-word detection score
  • Function: Binary classification for specific wake-word
  • Tensor Shape: [1, 16, 96] → [1]

Streaming Implementation

The pipeline operates on continuous audio through rolling-buffers:

  1. Audio Accumulation: Maintain rolling buffer of raw PCM samples
  2. Mel Frame Buffer: Store last 76 mel spectrogram frames for embedding computation
  3. Embedding Buffer: Keep last 16 embeddings for classification
  4. Threshold Detection: Apply 0.5 threshold to classification output

Performance Characteristics

Latency Profile

  • Mel Computation: ~5ms per 80ms chunk
  • Embedding Inference: ~10ms for 76-frame window
  • Classification: ~1ms for 16-embedding input
  • Total Pipeline: ~16ms processing time per chunk

Memory Requirements

  • Raw Audio Buffer: ~3.5KB (1760 samples × 2 bytes)
  • Mel Frame Buffer: ~9.7KB (76 frames × 32 bins × 4 bytes)
  • Embedding Buffer: ~6.1KB (16 embeddings × 96 dims × 4 bytes)
  • Total Working Memory: ~20KB for buffers

Implementation Considerations

Startup Handling

The pipeline requires pre-filling buffers with zeros during initialization to prevent crashes when the rolling windows don't have enough historical data. This results in noisier detection during the first ~1.5 seconds of operation.

Tensor Transformations

Critical shape manipulations between stages:

  • Mel output squeeze: [T, 1, 1, 32] → [T, 32]
  • Preprocessing: Apply x/10 + 2 normalization
  • Embedding squeeze: [1, 96] → [96]

Concurrency Design

Each stage can be processed independently, enabling:

  • Pipelined execution across multiple chunks
  • Separate worker queues for each model
  • swift-concurrency integration with nonisolated workers

Advantages

  1. Modularity: Each stage serves a distinct purpose and can be optimized independently
  2. Efficiency: Progressive dimensionality reduction (audio → spectrogram → embedding → score)
  3. Flexibility: Wake-word classifier can be retrained without touching feature extraction stages
  4. Real-time Performance: Streaming design with predictable memory usage

See also