openWakeWord
Confiance : high
wake-word-detectionvoice-activationonnx-modelsaudio-processingreal-time-inferencemel-spectrogramembedding-modelsswift-integration16khz-audiorolling-bufferstensor-processingaudio-capturestreaming-inferencecustom-traininghackathon-integrationthree-stage-pipelinetensor-shapesworker-patternsserial-dispatch-queue
Open-source wake-word detection framework that enables custom voice activation triggers through lightweight ONNX model inference. Designed for real-time audio processing with minimal computational overhead.
Architecture
Three-Stage ONNX Pipeline
Audio (16kHz PCM) → Mel-Spectrogram → Speech Embedding → Custom Classifier
80ms chunks 32 mel bins 96-dim vectors confidence score
Stage 1: Mel-Spectrogram Model
- Input: Raw audio samples
[1, N](minimum 400 samples) - Output: Time-frequency representation
[mel_frames, 32] - Function: Converts PCM audio to mel-frequency domain
Stage 2: Embedding Model
- Input: Mel window
[batch, 76, 32, 1](76 frames × 32 bins) - Output: Speech embedding
[96]dimensions - Function: Extracts semantic features from mel spectrograms
Stage 3: Custom Classifier
- Input: Embedding sequence
[1, 16, 96](16 recent embeddings) - Output: Wake-word confidence score
[0.0, 1.0] - Function: Detects target wake-word from embeddings
Real-Time Processing
Streaming Implementation
- Chunk Size: 80ms audio windows (1280 samples at 16kHz)
- Rolling Buffers: Maintains context for continuous inference
- Detection Threshold: Typically 0.5 for balanced accuracy/false-positives
- Latency: Sub-100ms detection with proper buffering
Buffer Management
// Circular audio buffer for mel-spectrogram input
rawAudioBuffer.append(newChunk)
melInput = rawAudioBuffer.suffix(1760) // ~110ms context
// Rolling mel-frame buffer for embedding model
melFrameBuffer.append(newMelFrames)
embeddingInput = melFrameBuffer.suffix(76) // 76-frame window
// Embedding history for classifier
embeddingHistory.append(newEmbedding)
classifierInput = embeddingHistory.suffix(16) // Last 16 embeddings
Swift Integration
ONNX Runtime Setup
class OpenWakeWordPipeline {
private let ortEnvironment: ORTEnv
private let melSpectrogramSession: ORTSession
private let embeddingSession: ORTSession
private let classifierSession: ORTSession
private var rawAudioBuffer: [Float] = []
private var melFrameBuffer: Float = []
private var embeddingBuffer: Float = []
}
Worker Thread Architecture
class WakeWordInferenceWorker {
private let workerQueue = DispatchQueue(label: "wakeword.inference")
private let pipeline: OpenWakeWordPipeline
func enqueueAudioChunk(_ samples: [Float]) {
workerQueue.async { [weak self] in
let confidence = self?.pipeline.processAudioChunk(samples)
if confidence > threshold {
self?.fireWakeWordDetected()
}
}
}
}
Audio Capture Integration
// AVAudioEngine tap for real-time audio
audioEngine.inputNode.installTap(onBus: 0, bufferSize: 1024, format: audioFormat) {
[weak self] buffer, _ in
let samples = buffer.floatChannelData?[0]
self?.inferenceWorker.enqueueAudioChunk(Array(samples))
}
Tensor Shape Transformations
Mel-Spectrogram Processing
- Raw input:
[1280](80ms of 16kHz audio) - Model output:
[T, 1, 1, 32](time × channels × height × mel_bins) - Post-processing: Squeeze to
[T, 32], applyx/10 + 2normalization
Embedding Extraction
- Model input:
[1, 76, 32, 1](batched mel frames) - Model output:
[1, 1, 1, 96](batched embedding) - Post-processing: Squeeze to
[96]flat vector
Classification Input
- Rolling window: Last 16 embeddings
[16, 96] - Model input:
[1, 16, 96](batch dimension added) - Model output:
[1, 1](confidence score)
Custom Model Training
Training Pipeline
# Feature extraction using openWakeWord public models
mel_model = onnxruntime.InferenceSession("melspectrogram.onnx")
embedding_model = onnxruntime.InferenceSession("embedding_model.onnx")
# Custom classifier training
class WakeWordClassifier(nn.Module):
def __init__(self):
super().__init__()
self.classifier = nn.Sequential(
nn.Linear(96*16, 128),
nn.ReLU(),
nn.Dropout(0.2),
nn.Linear(128, 1),
nn.Sigmoid()
)
Training Configuration
- Threshold: 0.5 (balance precision/recall)
- Context Window: 16 embeddings (~1.28 seconds)
- Training Data: Positive/negative samples with augmentation
- Model Size: ~13KB for efficient deployment
Performance Considerations
Memory Efficiency
- Model Size: 13KB classifier + public feature models (~1MB total)
- Buffer Limits: Fixed-size rolling buffers prevent memory growth
- Batch Processing: Single-sample inference minimizes memory usage
Latency Optimization
- Pipeline Stages: Parallelizable with careful buffer management
- Detection Speed: Sub-100ms typical response time
- False Positive Rate: Tunable via threshold adjustment
Resource Usage
- CPU: Lightweight inference suitable for real-time processing
- Memory: <10MB total footprint including buffers
- Power: Efficient for always-on scenarios
Integration Patterns
Hackathon Development
// Branch isolation for risk management
git checkout feat/clicky-fork // Stable demo path
git checkout -b feat/wakeword-realtime // Innovation branch
Production Considerations
- Model Versioning: Bundle models in app resources
- Graceful Degradation: Fallback to push-to-talk if wake-word fails
- Privacy: Local processing avoids cloud audio transmission
- Customization: Per-user threshold tuning for optimal experience
See also
- onnx-runtime - Cross-platform ML inference
- tensor-processing - Multi-dimensional array operations
- rolling-buffers - Streaming data management
- swift-concurrency - Thread-safe audio processing
- AVAudioEngine - Real-time audio capture
- mel-spectrogram-processing - Audio feature extraction