Wake-Word Integration
System design pattern for replacing manual activation methods (like push-to-talk) with voice-triggered conversation initiation in AI assistant applications. Requires careful orchestration of audio processing, permission management, and session lifecycle.
Integration Architecture
Session Lifecycle Management
Wake-word systems must handle distinct conversation states:
- Listening State: Continuous background audio monitoring
- Activation Trigger: Wake-word detection fires conversation start
- Recording State: Active speech capture for transcription
- Processing State: LLM generation and response playback
- Return to Listening: Either explicit close-word or timeout
Permission Orchestration
Always-on listening requires robust permission handling:
// Check microphone access before wake-word activation
func refreshAllPermissions() {
// Existing permission logic...
if microphonePermissionGranted {
startWakeWordDetection()
}
}
Conversation Flow Mapping
The integration maps wake-word events to existing conversation infrastructure:
- Wake-word Detection → Same flow as keyboard shortcut activation
- Close-word Recognition → Same flow as push-to-talk release
- Natural Endpointing → AssemblyAI's end-of-turn detection
Implementation Patterns
Event-Driven Architecture
Wake-word detectors emit discrete events that trigger state transitions:
protocol WakeWordDetectorDelegate: AnyObject {
func wakeWordDetected()
func closeWordDetected()
}
Background Processing
Continuous audio analysis requires careful thread management:
- Audio capture on dedicated serial queue
- ONNX inference on separate worker queue
- UI updates dispatched to main actor
- Memory management for rolling buffers
Keyterm Enhancement
Speech recognition systems benefit from biasing toward expected phrases:
// AssemblyAI keyterm configuration
let keyterms = [
"Xiexie", // Wake-word for better transcription
"thank you", // Close-word detection
"merci" // Multi-language support
]
Privacy and Performance Considerations
Local Processing
Wake-word detection typically runs entirely on-device:
- No audio data transmitted to servers
- Lightweight models (sub-20KB) for real-time inference
- Battery optimization through efficient algorithms
False Positive Management
Production systems need strategies for handling incorrect activations:
- Confidence thresholds (typically 0.5-0.7)
- Contextual filtering (speaker verification, environmental noise)
- User feedback mechanisms for model improvement
Audio Buffer Management
Continuous listening requires bounded memory usage:
- Fixed-size rolling buffers for audio streams
- Efficient tensor operations to minimize CPU usage
- Startup handling for incomplete audio contexts
User Experience Design
Feedback Mechanisms
Clear indication of system state:
- Visual indicators during listening/processing states
- Audio confirmation of wake-word detection
- Error handling for failed activations
Accessibility Considerations
Alternative activation methods for users who cannot use voice:
- Maintain keyboard shortcuts as fallback option
- Visual wake-word detection for sign language
- Customizable sensitivity settings
Multi-Modal Integration
Wake-words often combine with other interaction patterns:
- Gesture-based activation for hands-free scenarios
- Screen presence detection for privacy
- Integration with smart home ecosystems
Technical Challenges
Latency Requirements
Real-time wake-word detection demands:
- Sub-100ms processing latency per audio chunk
- Efficient model architectures (three-stage-pipeline)
- Optimized tensor operations and memory access patterns
Concurrency Management
Swift concurrency considerations:
- MainActor isolation for UI state
nonisolatedworkers for audio processing- Proper resource cleanup on detection lifecycle
Model Distribution
Packaging and deployment of ONNX models:
- Bundle inclusion via pbxproj-editing
- Resource management with file system synchronized groups
- Version control for model updates
See also
- openwakeword - Framework for custom wake-word training
- three-stage-pipeline - Audio processing architecture
- swift-concurrency - Threading model for real-time audio
- AssemblyAI Keyterms - Speech recognition enhancement
- rolling-buffers - Streaming audio data structures