~/wiki

openWakeWord

Confiance : high
wake-word-detectionvoice-activationonnx-modelsaudio-processingreal-time-inferencemel-spectrogramembedding-modelsswift-integration16khz-audiorolling-bufferstensor-processingaudio-capturestreaming-inferencecustom-traininghackathon-integrationthree-stage-pipelinetensor-shapesworker-patternsserial-dispatch-queue

Open-source wake-word detection framework that enables custom voice activation triggers through lightweight ONNX model inference. Designed for real-time audio processing with minimal computational overhead.

Architecture

Three-Stage ONNX Pipeline

Audio (16kHz PCM) → Mel-Spectrogram → Speech Embedding → Custom Classifier
    80ms chunks      32 mel bins      96-dim vectors    confidence score

Stage 1: Mel-Spectrogram Model

  • Input: Raw audio samples [1, N] (minimum 400 samples)
  • Output: Time-frequency representation [mel_frames, 32]
  • Function: Converts PCM audio to mel-frequency domain

Stage 2: Embedding Model

  • Input: Mel window [batch, 76, 32, 1] (76 frames × 32 bins)
  • Output: Speech embedding [96] dimensions
  • Function: Extracts semantic features from mel spectrograms

Stage 3: Custom Classifier

  • Input: Embedding sequence [1, 16, 96] (16 recent embeddings)
  • Output: Wake-word confidence score [0.0, 1.0]
  • Function: Detects target wake-word from embeddings

Real-Time Processing

Streaming Implementation

  • Chunk Size: 80ms audio windows (1280 samples at 16kHz)
  • Rolling Buffers: Maintains context for continuous inference
  • Detection Threshold: Typically 0.5 for balanced accuracy/false-positives
  • Latency: Sub-100ms detection with proper buffering

Buffer Management

// Circular audio buffer for mel-spectrogram input
rawAudioBuffer.append(newChunk)
melInput = rawAudioBuffer.suffix(1760) // ~110ms context

// Rolling mel-frame buffer for embedding model  
melFrameBuffer.append(newMelFrames)
embeddingInput = melFrameBuffer.suffix(76) // 76-frame window

// Embedding history for classifier
embeddingHistory.append(newEmbedding)
classifierInput = embeddingHistory.suffix(16) // Last 16 embeddings

Swift Integration

ONNX Runtime Setup

class OpenWakeWordPipeline {
    private let ortEnvironment: ORTEnv
    private let melSpectrogramSession: ORTSession
    private let embeddingSession: ORTSession  
    private let classifierSession: ORTSession
    
    private var rawAudioBuffer: [Float] = []
    private var melFrameBuffer: Float = []
    private var embeddingBuffer: Float = []
}

Worker Thread Architecture

class WakeWordInferenceWorker {
    private let workerQueue = DispatchQueue(label: "wakeword.inference")
    private let pipeline: OpenWakeWordPipeline
    
    func enqueueAudioChunk(_ samples: [Float]) {
        workerQueue.async { [weak self] in
            let confidence = self?.pipeline.processAudioChunk(samples)
            if confidence > threshold {
                self?.fireWakeWordDetected()
            }
        }
    }
}

Audio Capture Integration

// AVAudioEngine tap for real-time audio
audioEngine.inputNode.installTap(onBus: 0, bufferSize: 1024, format: audioFormat) { 
    [weak self] buffer, _ in
    let samples = buffer.floatChannelData?[0]
    self?.inferenceWorker.enqueueAudioChunk(Array(samples))
}

Tensor Shape Transformations

Mel-Spectrogram Processing

  • Raw input: [1280] (80ms of 16kHz audio)
  • Model output: [T, 1, 1, 32] (time × channels × height × mel_bins)
  • Post-processing: Squeeze to [T, 32], apply x/10 + 2 normalization

Embedding Extraction

  • Model input: [1, 76, 32, 1] (batched mel frames)
  • Model output: [1, 1, 1, 96] (batched embedding)
  • Post-processing: Squeeze to [96] flat vector

Classification Input

  • Rolling window: Last 16 embeddings [16, 96]
  • Model input: [1, 16, 96] (batch dimension added)
  • Model output: [1, 1] (confidence score)

Custom Model Training

Training Pipeline

# Feature extraction using openWakeWord public models
mel_model = onnxruntime.InferenceSession("melspectrogram.onnx")
embedding_model = onnxruntime.InferenceSession("embedding_model.onnx")

# Custom classifier training
class WakeWordClassifier(nn.Module):
    def __init__(self):
        super().__init__()
        self.classifier = nn.Sequential(
            nn.Linear(96*16, 128),
            nn.ReLU(),
            nn.Dropout(0.2),
            nn.Linear(128, 1),
            nn.Sigmoid()
        )

Training Configuration

  • Threshold: 0.5 (balance precision/recall)
  • Context Window: 16 embeddings (~1.28 seconds)
  • Training Data: Positive/negative samples with augmentation
  • Model Size: ~13KB for efficient deployment

Performance Considerations

Memory Efficiency

  • Model Size: 13KB classifier + public feature models (~1MB total)
  • Buffer Limits: Fixed-size rolling buffers prevent memory growth
  • Batch Processing: Single-sample inference minimizes memory usage

Latency Optimization

  • Pipeline Stages: Parallelizable with careful buffer management
  • Detection Speed: Sub-100ms typical response time
  • False Positive Rate: Tunable via threshold adjustment

Resource Usage

  • CPU: Lightweight inference suitable for real-time processing
  • Memory: <10MB total footprint including buffers
  • Power: Efficient for always-on scenarios

Integration Patterns

Hackathon Development

// Branch isolation for risk management
git checkout feat/clicky-fork          // Stable demo path
git checkout -b feat/wakeword-realtime  // Innovation branch

Production Considerations

  • Model Versioning: Bundle models in app resources
  • Graceful Degradation: Fallback to push-to-talk if wake-word fails
  • Privacy: Local processing avoids cloud audio transmission
  • Customization: Per-user threshold tuning for optimal experience

See also