Three-Stage Pipeline
Confiance : high
three-stage-pipelineopenWakeWordaudio-processingmel-spectrogramembedding-modelsonnx-runtimewake-word-detectiontensor-processingreal-time-inferencerolling-buffersstreaming-audio
Audio processing architecture used in openwakeword systems that chains together three specialized ONNX models to transform raw audio into wake-word detection decisions. This design separates concerns and enables efficient real-time inference through progressive feature extraction.
Pipeline Architecture
Stage 1: Mel Spectrogram Conversion
- Input: Raw PCM audio samples (1280 samples = 80ms at 16kHz)
- Output: Mel-scale spectrogram with 32 frequency bins
- Function: Converts time-domain audio to frequency-domain representation
- Tensor Shape:
[T, 32]where T varies based on audio chunk size
Stage 2: Feature Embedding
- Input: Sliding window of 76 mel spectrogram frames
- Output: 96-dimensional feature embedding vector
- Function: Extracts semantic audio features from spectral data
- Tensor Shape:
[1, 76, 32, 1]→[96]
Stage 3: Wake-Word Classification
- Input: Buffer of last 16 feature embeddings
- Output: Single wake-word detection score
- Function: Binary classification for specific wake-word
- Tensor Shape:
[1, 16, 96]→[1]
Streaming Implementation
The pipeline operates on continuous audio through rolling-buffers:
- Audio Accumulation: Maintain rolling buffer of raw PCM samples
- Mel Frame Buffer: Store last 76 mel spectrogram frames for embedding computation
- Embedding Buffer: Keep last 16 embeddings for classification
- Threshold Detection: Apply 0.5 threshold to classification output
Performance Characteristics
Latency Profile
- Mel Computation: ~5ms per 80ms chunk
- Embedding Inference: ~10ms for 76-frame window
- Classification: ~1ms for 16-embedding input
- Total Pipeline: ~16ms processing time per chunk
Memory Requirements
- Raw Audio Buffer: ~3.5KB (1760 samples × 2 bytes)
- Mel Frame Buffer: ~9.7KB (76 frames × 32 bins × 4 bytes)
- Embedding Buffer: ~6.1KB (16 embeddings × 96 dims × 4 bytes)
- Total Working Memory: ~20KB for buffers
Implementation Considerations
Startup Handling
The pipeline requires pre-filling buffers with zeros during initialization to prevent crashes when the rolling windows don't have enough historical data. This results in noisier detection during the first ~1.5 seconds of operation.
Tensor Transformations
Critical shape manipulations between stages:
- Mel output squeeze:
[T, 1, 1, 32]→[T, 32] - Preprocessing: Apply
x/10 + 2normalization - Embedding squeeze:
[1, 96]→[96]
Concurrency Design
Each stage can be processed independently, enabling:
- Pipelined execution across multiple chunks
- Separate worker queues for each model
- swift-concurrency integration with
nonisolatedworkers
Advantages
- Modularity: Each stage serves a distinct purpose and can be optimized independently
- Efficiency: Progressive dimensionality reduction (audio → spectrogram → embedding → score)
- Flexibility: Wake-word classifier can be retrained without touching feature extraction stages
- Real-time Performance: Streaming design with predictable memory usage
See also
- openwakeword - Framework implementing this architecture
- rolling-buffers - Data structure for streaming audio
- mel-spectrogram-processing - First stage implementation
- tensor-processing - Shape manipulation techniques
- onnx-runtime - Inference engine for model execution