~/wiki

Wake-Word Integration

Confiance : high
wake-word-integrationvoice-interfacealways-on-listeningpush-to-talk-replacementconversation-flowsession-managementaudio-permissionsprivacy-considerationsreal-time-audio

System design pattern for replacing manual activation methods (like push-to-talk) with voice-triggered conversation initiation in AI assistant applications. Requires careful orchestration of audio processing, permission management, and session lifecycle.

Integration Architecture

Session Lifecycle Management

Wake-word systems must handle distinct conversation states:

  1. Listening State: Continuous background audio monitoring
  2. Activation Trigger: Wake-word detection fires conversation start
  3. Recording State: Active speech capture for transcription
  4. Processing State: LLM generation and response playback
  5. Return to Listening: Either explicit close-word or timeout

Permission Orchestration

Always-on listening requires robust permission handling:

// Check microphone access before wake-word activation
func refreshAllPermissions() {
    // Existing permission logic...
    if microphonePermissionGranted {
        startWakeWordDetection()
    }
}

Conversation Flow Mapping

The integration maps wake-word events to existing conversation infrastructure:

  • Wake-word Detection → Same flow as keyboard shortcut activation
  • Close-word Recognition → Same flow as push-to-talk release
  • Natural Endpointing → AssemblyAI's end-of-turn detection

Implementation Patterns

Event-Driven Architecture

Wake-word detectors emit discrete events that trigger state transitions:

protocol WakeWordDetectorDelegate: AnyObject {
    func wakeWordDetected()
    func closeWordDetected()
}

Background Processing

Continuous audio analysis requires careful thread management:

  • Audio capture on dedicated serial queue
  • ONNX inference on separate worker queue
  • UI updates dispatched to main actor
  • Memory management for rolling buffers

Keyterm Enhancement

Speech recognition systems benefit from biasing toward expected phrases:

// AssemblyAI keyterm configuration
let keyterms = [
    "Xiexie",      // Wake-word for better transcription
    "thank you",   // Close-word detection
    "merci"        // Multi-language support
]

Privacy and Performance Considerations

Local Processing

Wake-word detection typically runs entirely on-device:

  • No audio data transmitted to servers
  • Lightweight models (sub-20KB) for real-time inference
  • Battery optimization through efficient algorithms

False Positive Management

Production systems need strategies for handling incorrect activations:

  • Confidence thresholds (typically 0.5-0.7)
  • Contextual filtering (speaker verification, environmental noise)
  • User feedback mechanisms for model improvement

Audio Buffer Management

Continuous listening requires bounded memory usage:

  • Fixed-size rolling buffers for audio streams
  • Efficient tensor operations to minimize CPU usage
  • Startup handling for incomplete audio contexts

User Experience Design

Feedback Mechanisms

Clear indication of system state:

  • Visual indicators during listening/processing states
  • Audio confirmation of wake-word detection
  • Error handling for failed activations

Accessibility Considerations

Alternative activation methods for users who cannot use voice:

  • Maintain keyboard shortcuts as fallback option
  • Visual wake-word detection for sign language
  • Customizable sensitivity settings

Multi-Modal Integration

Wake-words often combine with other interaction patterns:

  • Gesture-based activation for hands-free scenarios
  • Screen presence detection for privacy
  • Integration with smart home ecosystems

Technical Challenges

Latency Requirements

Real-time wake-word detection demands:

  • Sub-100ms processing latency per audio chunk
  • Efficient model architectures (three-stage-pipeline)
  • Optimized tensor operations and memory access patterns

Concurrency Management

Swift concurrency considerations:

  • MainActor isolation for UI state
  • nonisolated workers for audio processing
  • Proper resource cleanup on detection lifecycle

Model Distribution

Packaging and deployment of ONNX models:

  • Bundle inclusion via pbxproj-editing
  • Resource management with file system synchronized groups
  • Version control for model updates

See also