~/wiki

Concepts — vue longue

retour à la liste

Toutes les pages concaténées sur un seul document, pour un Ctrl-F direct.

Audio Buffer Management

page dédiée →

Systematic approach to handling streaming audio data in real-time applications, balancing memory efficiency with processing requirements. Critical for voice AI applications that require continuous audio monitoring without unbounded memory growth.

Core Principles

Memory Constraints

Fixed-Size Buffers:

  • Prevent memory leaks in long-running audio applications
  • Maintain predictable memory footprint
  • Enable deterministic processing latency

Circular Buffer Pattern:

private var audioBuffer: [Float] = []
private let maxBufferSize = 16000 // 1 second at 16kHz

func appendSamples(_ newSamples: [Float]) {
    audioBuffer.append(contentsOf: newSamples)
    if audioBuffer.count > maxBufferSize {
        audioBuffer.removeFirst(audioBuffer.count - maxBufferSize)
    }
}

Real-Time Requirements

Low-Latency Processing:

  • Audio callbacks execute on dedicated audio thread
  • Minimize processing in callback to prevent dropouts
  • Forward samples to worker queue for heavy computation

Thread Safety:

  • Audio engine callbacks not on main thread
  • Synchronize buffer access across threads
  • Use swift-concurrency patterns for coordination

Implementation Patterns

AVAudioEngine Integration

Audio Tap Setup:

let audioEngine = AVAudioEngine()
let inputNode = audioEngine.inputNode
let recordingFormat = inputNode.outputFormat(forBus: 0)

inputNode.installTap(onBus: 0, 
                    bufferSize: 1024,
                    format: recordingFormat) { buffer, time in
    // Process audio buffer
    self.forwardToProcessor(buffer)
}

Buffer Conversion:

  • AVAudioPCMBuffer → Float array conversion
  • Handle different audio formats (16-bit, 24-bit, float)
  • Maintain consistent sample rate across pipeline

Streaming Context Windows

Rolling Window Strategy:

  • Maintain sufficient context for model inference
  • Example: openwakeword needs 1760 samples for mel-spectrogram
  • Efficient memory usage through sliding window

Chunk-Based Processing:

private let chunkSize = 1280 // 80ms at 16kHz
private let contextSize = 1760 // Required for mel computation

func processChunk(_ chunk: [Float]) {
    buffer.append(contentsOf: chunk)
    
    if buffer.count >= contextSize {
        let processingWindow = Array(buffer.suffix(contextSize))
        performInference(processingWindow)
        
        // Keep only what we need for next iteration
        if buffer.count > contextSize {
            buffer.removeFirst(buffer.count - contextSize)
        }
    }
}

Performance Optimization

Memory Allocation

Pre-Allocation Strategy:

  • Reserve buffer capacity to avoid repeated allocation
  • Use Array.reserveCapacity() for known maximum sizes
  • Minimize allocations in audio callback path

Copy Minimization:

  • Use withUnsafeBytes for zero-copy data access
  • Direct memory mapping where possible
  • Avoid unnecessary format conversions

Concurrency Patterns

Producer-Consumer Model:

  • Audio thread produces samples (lightweight)
  • Worker thread consumes for processing (heavy)
  • Lock-free queue for inter-thread communication

Backpressure Handling:

  • Drop samples if processing can't keep up
  • Maintain real-time responsiveness over accuracy
  • Monitor queue depth to detect performance issues

Error Handling

Buffer Overflow Protection

Graceful Degradation:

func safeBu appendSamples(_ samples: [Float]) {
    guard samples.count <= maxBufferSize else {
        // Drop samples that would overflow
        let keepCount = min(samples.count, maxBufferSize)
        buffer = Array(samples.suffix(keepCount))
        return
    }
    
    buffer.append(contentsOf: samples)
    if buffer.count > maxBufferSize {
        let excessCount = buffer.count - maxBufferSize
        buffer.removeFirst(excessCount)
    }
}

Recovery Strategies:

  • Reset buffer state on processing errors
  • Maintain minimum viable buffer for continued operation
  • Log buffer statistics for performance monitoring

Audio Interruption Handling

System Events:

  • Handle phone calls, notifications, other app audio
  • Gracefully pause/resume audio processing
  • Rebuild buffer state after interruption

Microphone Permissions:

  • Handle denied/revoked microphone access
  • Provide user feedback for permission issues
  • Graceful fallback when audio unavailable

Use Cases

Wake-Word Detection

Continuous Monitoring:

  • 24/7 audio capture with minimal memory footprint
  • Buffer management for openwakeword inference pipeline
  • Balance between detection accuracy and resource usage

Voice Activity Detection

Dynamic Buffering:

  • Expand buffer during speech segments
  • Compress during silence periods
  • Adaptive algorithms based on audio characteristics

Real-Time Transcription

Streaming ASR:

  • Buffer audio chunks for API submission
  • Manage overlapping windows for continuous transcription
  • Handle network latency without audio loss

See also

Mel-Spectrogram Processing

page dédiée →

Frequency-domain representation of audio signals that converts time-domain waveforms into mel-scale spectrograms, essential for speech recognition and audio analysis applications. Forms the preprocessing foundation for most modern voice AI systems.

Core Concepts

Mel Scale Transformation

Mel Scale Properties:

  • Perceptually linear frequency scale
  • Better aligns with human auditory perception
  • Compresses higher frequencies more than lower ones
  • Formula: mel = 2595 * log10(1 + hz/700)

Frequency Binning:

  • Standard configurations: 32, 64, 80, or 128 mel bins
  • Each bin represents a frequency range
  • Lower frequencies get more bins (higher resolution)
  • Higher frequencies get fewer bins (compression)

Spectrogram Computation

Time-Frequency Analysis:

  • Short-Time Fourier Transform (STFT) applied to audio windows
  • Hop length determines temporal resolution
  • Window size affects frequency resolution
  • Overlapping windows for smooth transitions

Processing Pipeline:

  1. Windowing: Apply Hann/Hamming window to audio segments
  2. FFT: Compute frequency spectrum for each window
  3. Mel Filtering: Apply mel-scale filter bank
  4. Log Transform: Convert to logarithmic scale (dB)

Implementation Patterns

ONNX Model Integration

Model Input/Output:

Input:  [batch_size, audio_samples] (PCM16 audio)
Output: [time_frames, mel_bins, 1, 1] (mel-spectrogram)

Typical Configurations:

  • Audio: 16 kHz mono, 1280 samples per chunk (80ms)
  • Output: 32 mel bins, variable time frames
  • Context: Minimum 400 samples for stable computation

Real-Time Processing

Streaming Considerations:

  • Overlapping audio windows for continuity
  • Buffer management for context windows
  • Memory-efficient frame extraction
  • Padding strategies for startup/shutdown

Buffer Management:

# Conceptual streaming pattern
class MelSpectrogramStreamer:
    def __init__(self):
        self.audio_buffer = []
        self.context_size = 1760  # samples for full context
        
    def process_chunk(self, new_samples):
        self.audio_buffer.extend(new_samples)
        
        # Extract with context
        if len(self.audio_buffer) >= self.context_size:
            window = self.audio_buffer[-self.context_size:]
        else:
            # Pad with zeros on startup
            window = [0.0] * (self.context_size - len(self.audio_buffer)) + self.audio_buffer
        
        return self.compute_mel_spectrogram(window)

Audio Preprocessing Requirements

Sample Rate and Format

Standard Configurations:

  • 16 kHz mono: Most common for speech recognition
  • 8 kHz: Telephony applications
  • 44.1/48 kHz: High-fidelity audio applications

Format Conversion:

  • PCM16 → Float32 normalization
  • Multi-channel → mono downmixing
  • Resampling for rate conversion
  • DC offset removal

Frame Processing

Temporal Parameters:

  • Frame size: Usually 25ms (400 samples at 16kHz)
  • Hop size: Usually 10ms (160 samples at 16kHz)
  • Context window: May require 1-2 seconds for stable features
  • Overlap: Typically 50-75% between consecutive frames

Feature Engineering

Post-Processing Transformations

Normalization Strategies:

# Common transformation patterns
def normalize_mel_features(mel_frames):
    # Method 1: Per-utterance normalization
    mean = np.mean(mel_frames, axis=0)
    std = np.std(mel_frames, axis=0) + 1e-8
    normalized = (mel_frames - mean) / std
    
    # Method 2: Global statistics (from training)
    # normalized = (mel_frames - global_mean) / global_std
    
    # Method 3: Min-max scaling
    # normalized = (mel_frames - min_val) / (max_val - min_val)
    
    return normalized

Domain-Specific Transformations:

  • Delta features (velocity): First-order time derivatives
  • Delta-delta features (acceleration): Second-order derivatives
  • Mean subtraction for robustness
  • Variance normalization across time/frequency

Feature Augmentation

Training-Time Enhancements:

  • SpecAugment: Frequency/time masking
  • Noise addition for robustness
  • Speed perturbation (time stretching)
  • Volume augmentation

openWakeWord Integration

Pipeline Architecture

Three-Stage Processing:

  1. Mel-Spectrogram: Raw audio → frequency features
  2. Embedding: Mel frames

Multi-Source Ingestion

page dédiée →

Architecture pattern for knowledge management systems that automatically captures and processes content from diverse sources and formats. Essential for comprehensive knowledge accumulation without manual overhead, enabling systematic learning from all information channels.

Core Architecture

Source Categories

1. Development Projects

  • Automated sync from ~/code/*/ using rsync
  • Intelligent filtering (exclude node_modules, .git, build artifacts)
  • Git diff tracking to process only changed files
  • Captures READMEs, documentation, and learning artifacts

2. Visual Content

  • iOS Shortcuts integration for frictionless screenshot capture
  • iCloud Drive sync for automatic mobile-to-desktop transfer
  • LLM vision for OCR and semantic classification
  • Context-aware processing of social media content

3. Audio Content

  • Conference talks and meeting recordings
  • Voice memos and interview transcripts
  • Whisper-based transcription with MLX optimization
  • Automatic speaker identification and topic segmentation

4. Web Content

  • RSS feed aggregation from key sources
  • Manual article addition with URL parsing
  • Research paper ingestion from arXiv and academic sources
  • Blog posts and technical documentation

5. Conversational Content

  • AI agent conversations from Claude Code, Cursor IDE
  • Chat transcripts with technical discussions
  • Code review conversations and architectural decisions

Processing Pipeline

Raw Sources → Triage → Classification → Extraction → Integration
     ↓           ↓          ↓           ↓          ↓
  Diverse     Quality   Content-Type   Key Info   Wiki Pages
  Formats     Filter    Detection      Capture    & References

Triage System

Automated Quality Assessment:

  • Content length and depth analysis
  • Source credibility scoring
  • Relevance to existing knowledge base
  • Novelty detection against existing pages

Classification Outcomes:

  • High: Immediate processing and integration
  • Medium: Queue for manual review in raw/pending/
  • Low: Archive to raw/discarded/ with reasoning

Content Type Detection

Intelligent Format Handling:

  • Markdown parsing and structure analysis
  • Image OCR with context understanding
  • Audio transcription with speaker diarization
  • Code extraction and documentation linking

Implementation Patterns

Git-Based Change Detection

Revolutionary approach using git-diff-tracking to eliminate manifest complexity:

# Detect changes since last ingest
git diff --name-status HEAD~1 raw/projects/
# Process only modified files
find raw/ -newer .last_ingest -type f

iOS Shortcuts Integration

Seamless mobile capture through ios-shortcuts-integration:

iPhone Screenshot → Share → "Brain Wiki" Shortcut → iCloud Drive → Processing Queue

Batch Processing

Efficient handling of large content volumes:

  • Process screenshots in batches for OCR efficiency
  • Parallel transcription of multiple audio files
  • Async classification of web articles
  • Incremental project synchronization

Cross-Modal Synthesis

Automated Cross-Referencing

  • Screenshot insights link to related projects
  • Conference talks reference technical documentation
  • Project learnings connect to research papers
  • Agent conversations enhance concept pages

Compound Learning Effects

Multi-source integration creates knowledge compounds:

  • Visual content + code examples + academic papers = comprehensive understanding
  • Social media insights + project experience + formal documentation = practical wisdom

Real-World Implementation

Epic Brain Wiki Results

From the complete brain-wiki implementation:

Sources Processed:

  • 31 AI/ML projects (4,060 markdown files)
  • Screenshot pipeline with iOS Shortcuts integration
  • Audio transcription using MLX Whisper
  • Web content through RSS and manual addition
  • Agent conversations from Claude Code sessions

Automation Achieved:

  • Zero manual overhead for routine ingestion
  • Daily orchestration processing all sources
  • Intelligent filtering eliminating 90% of noise
  • Cross-modal integration creating compound insights

Measured Benefits

  • 10x faster knowledge capture compared to manual curation
  • 5x more cross-references discovered through automated analysis
  • 90% reduction in manual content processing time
  • Continuous operation without human intervention

Technical Implementation

Directory Structure

multi-source-system/
├── raw/                    # Immutable source storage
│   ├── projects/          # Development project files
│   ├── screenshots/       # Visual content capture
│   ├── talks/            # Audio transcriptions
│   ├── articles/         # Web content archive
│   ├── conversations/    # Agent chat logs
│   ├── pending/          # Medium-quality sources
│   └── discarded/        # Low-quality archive
├── scripts/
│   ├── collect-projects.sh    # Project synchronization
│   ├── process-screenshots.sh # Image handling
│   ├── process-recordings.sh  # Audio transcription
│   └── daily-ingest.md       # Orchestration guide
└── processed/             # Integrated wiki content

Automation Scripts

Project Collection:

# Sync all code projects
rsync -av --exclude="node_modules" --exclude=".git" \
      ~/code/ raw/projects/

Screenshot Processing:

# Move from iCloud to processing queue
mv ~/Library/Mobile\ Documents/com~apple~CloudDocs/brain-wiki-inbox/* \
   raw/screenshots/inbox/

Audio Transcription:

# Whisper transcription
mlx_whisper audio_file.mp3 --output-format txt

Quality Control

Automated Validation

  • Minimum content length thresholds
  • Duplicate detection across sources
  • Link validation and reference checking
  • Format consistency verification

Human Oversight

  • Weekly review of medium-quality sources
  • Manual curation of cross-references
  • System performance monitoring
  • Pipeline optimization decisions

Scaling Considerations

Performance Optimization

  • Incremental processing to handle growth
  • Batch operations for efficiency
  • Resource pooling for concurrent tasks
  • Cache layers for repeated operations

Storage Management

  • Compression for archived content
  • Automated cleanup of outdated sources
  • Backup strategies for critical content
  • Version control for all processed data

Advanced Features

Content Enrichment

  • Automatic tagging based on content analysis
  • Entity extraction and relationship mapping
  • Topic modeling for content clustering
  • Sentiment analysis for conversational content

Adaptive Processing

  • Learning from user feedback on content quality
  • Dynamic threshold adjustment for triage
  • Personalized relevance scoring
  • Context-aware classification improvement

Implementation Challenges

Content Quality Variability

  • Handling low-signal sources (screenshots of memes)
  • Dealing with incomplete or corrupted files
  • Managing different content formats and structures
  • Balancing automation with quality control

Technical Complexity

  • Coordinating multiple processing pipelines
  • Error handling across diverse input types
  • Resource management for intensive operations
  • Maintaining system reliability and uptime

Privacy and Security

  • Sensitive content identification and handling
  • Access control for different source types
  • Secure storage of personal information
  • Compliance with data protection requirements

See Also

Multimodal AI

page dédiée →

AI systems capable of understanding, processing, and generating content across multiple modalities including text, images, audio, and video. Represents a significant advancement over single-modality AI systems.

Core Concepts

Modality Integration

Multimodal AI systems can:

  • Process inputs from multiple modalities simultaneously
  • Generate outputs in different modalities than the input
  • Understand relationships and correspondences between modalities
  • Maintain coherent understanding across modality boundaries

Common Modalities

  • Text: Natural language processing and generation
  • Vision: Image understanding, object detection, scene analysis
  • Audio: Speech recognition, audio understanding, music analysis
  • Video: Temporal visual understanding, motion analysis

Industry Examples

Google's Approach

google has advanced multimodal capabilities through:

  • gemini-omni: Latest comprehensive multimodal model
  • Gemini 3.5: Enhanced multimodal capabilities
  • Integration across Google's product ecosystem

Competitive Landscape

Major players developing multimodal AI include:

  • google with Gemini family
  • openai with GPT-4V and beyond
  • anthropic with Claude's vision capabilities

Applications

  • Cross-Modal Search: Finding content across different media types
  • Content Creation: Generating images from text, videos from descriptions
  • Accessibility: Converting between modalities for different user needs
  • Interactive Agents: voice-agents with visual understanding

Technical Challenges

  • Alignment: Ensuring consistent understanding across modalities
  • Efficiency: Processing multiple input types without excessive compute
  • Training: Developing datasets and methods for multimodal learning
  • Evaluation: Benchmarking performance across diverse tasks

See also

ONNX Runtime

page dédiée →

Cross-platform, high-performance machine learning inference engine that executes ONNX (Open Neural Network Exchange) models. Provides native bindings for multiple programming languages including Swift, enabling efficient ML model deployment in production applications.

Core Features

Cross-Platform Deployment

  • Universal Format: ONNX models run consistently across platforms
  • Hardware Optimization: Automatic acceleration using available hardware (CPU, GPU, specialized chips)
  • Language Bindings: Native APIs for C++, Python, C#, Java, Swift, and others
  • Mobile Optimization: Lightweight inference for iOS/Android applications

Performance Characteristics

  • Optimized Inference: Graph optimization and kernel fusion
  • Memory Efficiency: Minimal memory footprint for edge deployment
  • Batching Support: Process multiple inputs simultaneously
  • Precision Options: FP32, FP16, INT8 quantization support

Swift Integration

Objective-C Bridge Architecture

ONNX Runtime Swift support comes through Objective-C bindings that bridge to the native C++ runtime:

import onnxruntime_objc

class ONNXInferenceEngine {
    private let ortEnvironment: ORTEnv
    private let session: ORTSession
    
    init(modelPath: String) throws {
        ortEnvironment = try ORTEnv(loggingLevel: .warning)
        session = try ORTSession(env: ortEnvironment, modelPath: modelPath)
    }
}

Session Management

Model Loading:

// Load model from bundle
guard let modelPath = Bundle.main.path(forResource: "model", ofType: "onnx") else {
    throw ModelError.fileNotFound
}
let session = try ORTSession(env: environment, modelPath: modelPath)

Session Configuration:

  • Set execution providers (CPU, CoreML, etc.)
  • Configure memory patterns and optimization level
  • Set thread count for CPU inference

Tensor Operations

Input Preparation:

func prepareInput(_ audioSamples: [Float]) throws -> ORTValue {
    let inputTensor = try ORTValue(
        tensorData: NSMutableData(bytes: audioSamples, length: audioSamples.count * 4),
        elementType: .float,
        shape: [1, NSNumber(value: audioSamples.count)]
    )
    return inputTensor
}

**Running

openWakeWord

page dédiée →

Open-source wake-word detection framework that enables custom voice activation triggers through lightweight ONNX model inference. Designed for real-time audio processing with minimal computational overhead.

Architecture

Three-Stage ONNX Pipeline

Audio (16kHz PCM) → Mel-Spectrogram → Speech Embedding → Custom Classifier
    80ms chunks      32 mel bins      96-dim vectors    confidence score

Stage 1: Mel-Spectrogram Model

  • Input: Raw audio samples [1, N] (minimum 400 samples)
  • Output: Time-frequency representation [mel_frames, 32]
  • Function: Converts PCM audio to mel-frequency domain

Stage 2: Embedding Model

  • Input: Mel window [batch, 76, 32, 1] (76 frames × 32 bins)
  • Output: Speech embedding [96] dimensions
  • Function: Extracts semantic features from mel spectrograms

Stage 3: Custom Classifier

  • Input: Embedding sequence [1, 16, 96] (16 recent embeddings)
  • Output: Wake-word confidence score [0.0, 1.0]
  • Function: Detects target wake-word from embeddings

Real-Time Processing

Streaming Implementation

  • Chunk Size: 80ms audio windows (1280 samples at 16kHz)
  • Rolling Buffers: Maintains context for continuous inference
  • Detection Threshold: Typically 0.5 for balanced accuracy/false-positives
  • Latency: Sub-100ms detection with proper buffering

Buffer Management

// Circular audio buffer for mel-spectrogram input
rawAudioBuffer.append(newChunk)
melInput = rawAudioBuffer.suffix(1760) // ~110ms context

// Rolling mel-frame buffer for embedding model  
melFrameBuffer.append(newMelFrames)
embeddingInput = melFrameBuffer.suffix(76) // 76-frame window

// Embedding history for classifier
embeddingHistory.append(newEmbedding)
classifierInput = embeddingHistory.suffix(16) // Last 16 embeddings

Swift Integration

ONNX Runtime Setup

class OpenWakeWordPipeline {
    private let ortEnvironment: ORTEnv
    private let melSpectrogramSession: ORTSession
    private let embeddingSession: ORTSession  
    private let classifierSession: ORTSession
    
    private var rawAudioBuffer: [Float] = []
    private var melFrameBuffer: Float = []
    private var embeddingBuffer: Float = []
}

Worker Thread Architecture

class WakeWordInferenceWorker {
    private let workerQueue = DispatchQueue(label: "wakeword.inference")
    private let pipeline: OpenWakeWordPipeline
    
    func enqueueAudioChunk(_ samples: [Float]) {
        workerQueue.async { [weak self] in
            let confidence = self?.pipeline.processAudioChunk(samples)
            if confidence > threshold {
                self?.fireWakeWordDetected()
            }
        }
    }
}

Audio Capture Integration

// AVAudioEngine tap for real-time audio
audioEngine.inputNode.installTap(onBus: 0, bufferSize: 1024, format: audioFormat) { 
    [weak self] buffer, _ in
    let samples = buffer.floatChannelData?[0]
    self?.inferenceWorker.enqueueAudioChunk(Array(samples))
}

Tensor Shape Transformations

Mel-Spectrogram Processing

  • Raw input: [1280] (80ms of 16kHz audio)
  • Model output: [T, 1, 1, 32] (time × channels × height × mel_bins)
  • Post-processing: Squeeze to [T, 32], apply x/10 + 2 normalization

Embedding Extraction

  • Model input: [1, 76, 32, 1] (batched mel frames)
  • Model output: [1, 1, 1, 96] (batched embedding)
  • Post-processing: Squeeze to [96] flat vector

Classification Input

  • Rolling window: Last 16 embeddings [16, 96]
  • Model input: [1, 16, 96] (batch dimension added)
  • Model output: [1, 1] (confidence score)

Custom Model Training

Training Pipeline

# Feature extraction using openWakeWord public models
mel_model = onnxruntime.InferenceSession("melspectrogram.onnx")
embedding_model = onnxruntime.InferenceSession("embedding_model.onnx")

# Custom classifier training
class WakeWordClassifier(nn.Module):
    def __init__(self):
        super().__init__()
        self.classifier = nn.Sequential(
            nn.Linear(96*16, 128),
            nn.ReLU(),
            nn.Dropout(0.2),
            nn.Linear(128, 1),
            nn.Sigmoid()
        )

Training Configuration

  • Threshold: 0.5 (balance precision/recall)
  • Context Window: 16 embeddings (~1.28 seconds)
  • Training Data: Positive/negative samples with augmentation
  • Model Size: ~13KB for efficient deployment

Performance Considerations

Memory Efficiency

  • Model Size: 13KB classifier + public feature models (~1MB total)
  • Buffer Limits: Fixed-size rolling buffers prevent memory growth
  • Batch Processing: Single-sample inference minimizes memory usage

Latency Optimization

  • Pipeline Stages: Parallelizable with careful buffer management
  • Detection Speed: Sub-100ms typical response time
  • False Positive Rate: Tunable via threshold adjustment

Resource Usage

  • CPU: Lightweight inference suitable for real-time processing
  • Memory: <10MB total footprint including buffers
  • Power: Efficient for always-on scenarios

Integration Patterns

Hackathon Development

// Branch isolation for risk management
git checkout feat/clicky-fork          // Stable demo path
git checkout -b feat/wakeword-realtime  // Innovation branch

Production Considerations

  • Model Versioning: Bundle models in app resources
  • Graceful Degradation: Fallback to push-to-talk if wake-word fails
  • Privacy: Local processing avoids cloud audio transmission
  • Customization: Per-user threshold tuning for optimal experience

See also

Rolling Buffers

page dédiée →

Data structures that maintain a fixed-size sliding window of elements, automatically discarding old data as new data arrives. Essential for streaming applications where you need to maintain context over time while controlling memory usage.

Core Concept

Sliding Window: Fixed-size buffer that "rolls" forward as new data arrives, keeping only the most recent N elements.

class RollingBuffer<T> {
    private var buffer: [T] = []
    private let maxSize: Int
    
    init(maxSize: Int) {
        self.maxSize = maxSize
    }
    
    func append(_ element: T) {
        buffer.append(element)
        if buffer.count > maxSize {
            buffer.removeFirst()
        }
    }
    
    func suffix(_ count: Int) -> ArraySlice<T> {
        return buffer.suffix(count)
    }
}

Audio Processing Applications

Wake-Word Detection Pipeline

Rolling buffers enable continuous audio processing by maintaining multiple overlapping context windows:

// Raw audio buffer for mel-spectrogram computation
rawAudioBuffer.append(newChunk)           // Add 80ms chunk
melInput = rawAudioBuffer.suffix(1760)    // Use last ~110ms for context

// Mel-frame buffer for embedding model
melFrameBuffer.append(newMelFrames)       // Add new frames  
embeddingInput = melFrameBuffer.suffix(76) // 76-frame window

// Embedding buffer for classification
embeddingBuffer.append(newEmbedding)      // Add 96-dim vector
classifierInput = embeddingBuffer.suffix(16) // Last 16 embeddings

Cascading Buffer Chain

Each stage maintains its own rolling buffer with appropriate window sizes:

  1. Audio Samples: 1760 samples (~110ms) for mel-spectrogram context
  2. Mel Frames: 76 frames for embedding model input window
  3. Embeddings: 16 vectors (~1.28s) for wake-word classification

Memory Management Benefits

Bounded Memory Usage

// Without rolling buffers - memory grows unbounded
var allAudioSamples: [Float] = []  // Grows forever
allAudioSamples.append(contentsOf: newChunk)

// With rolling buffers - fixed memory footprint
let audioBuffer = RollingBuffer<Float>(maxSize: 1760)
audioBuffer.append(contentsOf: newChunk)  // Auto-discards old data

Predictable Resource Usage

  • Audio Buffer: 1760 × 4 bytes = ~7KB
  • Mel Buffer: 76 × 32 × 4 bytes = ~10KB
  • Embedding Buffer: 16 × 96 × 4 bytes = ~6KB
  • Total: <25KB for complete pipeline context

Implementation Patterns

Startup Padding

Handle insufficient data during initialization:

func processAudioChunk(_ chunk: [Float]) -> Float? {
    rawAudioBuffer.append(contentsOf: chunk)
    
    // Need minimum samples for mel-spectrogram
    guard rawAudioBuffer.count >= minSamplesRequired else {
        return nil  // Skip inference until buffer fills
    }
    
    let melInput = rawAudioBuffer.suffix(1760)
    // Proceed with inference...
}

Buffer Synchronization

Coordinate multiple rolling buffers for pipeline consistency:

class StreamingPipeline {
    private let audioBuffer = RollingBuffer<Float>(maxSize: 1760)
    private let melBuffer = RollingBuffer<[Float]>(maxSize: 76)  
    private let embeddingBuffer = RollingBuffer<[Float]>(maxSize: 16)
    
    func process(_ chunk: [Float]) {
        // Stage 1: Audio → Mel
        audioBuffer.append(contentsOf: chunk)
        let newMelFrames = computeMelSpectrogram(audioBuffer.suffix(1760))
        
        // Stage 2: Mel → Embedding  
        melBuffer.append(contentsOf: newMelFrames)
        let newEmbedding = computeEmbedding(melBuffer.suffix(76))
        
        // Stage 3: Embedding → Classification
        embeddingBuffer.append(newEmbedding)
        if embeddingBuffer.count >= 16 {
            let confidence = classify(embeddingBuffer.suffix(16))
            return confidence
        }
    }
}

Performance Characteristics

Time Complexity

  • Append: O(1) amortized (Array.append + conditional removeFirst)
  • Suffix Access: O(k) where k is suffix length
  • Space: O(maxSize) fixed memory footprint

Optimization Strategies

// Circular buffer for O(1) operations
class CircularRollingBuffer<T> {
    private var buffer: [T?]
    private var head: Int = 0
    private var count: Int = 0
    private let capacity: Int
    
    func append(_ element: T) {
        buffer[head] = element
        head = (head + 1) % capacity
        count = min(count + 1, capacity)
    }
    
    func recentElements(_ k: Int) -> [T] {
        // Extract last k elements in order
        let start = (head - min(k, count) + capacity) % capacity
        // Implementation details...
    }
}

Real-Time Streaming Patterns

Continuous Processing Loop

func startStreaming() {
    audioEngine.inputNode.installTap(onBus: 0, bufferSize: 1024, format: format) { 
        [weak self] buffer, _ in
        
        let samples = Array(buffer.floatChannelData![0][0..<Int(buffer.frameLength)])
        
        self?.workerQueue.async {
            if let confidence = self?.pipeline.processChunk(samples) {
                if confidence > threshold {
                    DispatchQueue.main.async {
                        self?.onWakeWordDetected()
                    }
                }
            }
        }
    }
}

Buffer Warm-Up Strategy

// Pre-fill buffers with zeros to avoid startup delays
func initializeBuffers() {
    // Fill audio buffer with silence
    let silence = [Float](repeating: 0.0, count: 1760)
    audioBuffer.append(contentsOf: silence)
    
    // Pre-compute initial mel frames
    let initialMel = computeMelSpectrogram(silence)
    melBuffer.append(contentsOf: initialMel)
    
    // Note: First ~1.5s of detection may be noisier
}

Error Handling and Edge Cases

Buffer Underflow

func safeSuffix<T>(_ buffer: [T], _ count: Int) -> ArraySlice<T> {
    let availableCount = min(count, buffer.count)
    return buffer.suffix(availableCount)
}

Thread Safety

class ThreadSafeRollingBuffer<T> {
    private let queue = DispatchQueue(label: "rolling-buffer")
    private var buffer: [T] = []
    
    func append(_ element: T) {
        queue.sync {
            buffer.append(element)
            if buffer.count > maxSize {
                buffer.removeFirst()
            }
        }
    }
    
    func recentElements(_ count: Int) -> [T] {
        return queue.sync {
            return Array(buffer.suffix(count))
        }
    }
}

Use Cases Beyond Audio

Time Series Analytics

  • Metric Monitoring: Rolling window for moving averages
  • Anomaly Detection: Recent data points for trend analysis
  • Real-Time Dashboards: Latest N data points for visualization

Network Streaming

  • Video Buffering: Fixed-size frame buffer for smooth playback
  • Chat Systems: Recent message history with memory bounds
  • Game State: Rolling history for replay and prediction

Machine Learning

  • Online Learning: Recent samples for model updates
  • Feature Engineering: Temporal windows for sequence models
  • Real-Time Inference: Context maintenance for streaming predictions

See also

Swift Concurrency

page dédiée →

Modern concurrency system in Swift providing actor-based isolation, structured concurrency, and compile-time thread safety guarantees. Essential for building responsive iOS/macOS applications that handle multiple concurrent operations safely.

Core Concepts

Actor Isolation

  • Actors: Reference types that protect their mutable state by serializing access
  • MainActor: Global actor representing the main thread, required for UI updates
  • Isolated Methods: Can only be called from within the same actor context
  • Nonisolated Methods: Can be called from any context without actor hopping

Async/Await

  • Async Functions: Methods that can suspend execution and resume later
  • Await: Keyword for calling async functions and potentially yielding control
  • Structured Concurrency: Task hierarchies with automatic cancellation propagation

Real-World Challenges

MainActor Default Isolation

In projects with SWIFT_DEFAULT_ACTOR_ISOLATION = MainActor, all classes become main-actor-isolated by default unless explicitly marked otherwise. This creates challenges when:

// This class is implicitly @MainActor
class AudioProcessor {
    func processAudio() { /* Must run on main thread */ }
}

// Worker queue access requires careful handling
let processor = AudioProcessor()
DispatchQueue.global().async {
    // ERROR: Main actor-isolated instance cannot be accessed
    processor.processAudio()
}

Audio Processing Patterns

Real-time audio processing requires background execution but must coordinate with main-actor UI updates:

@MainActor
class WakeWordDetector {
    private let inferenceWorker = WakeWordInferenceWorker()
    
    // Audio tap runs on audio thread, needs careful bridging
    private func setupAudioTap() {
        audioEngine.inputNode.installTap(onBus: 0) { [weak self] buffer, _ in
            // This callback runs on audio thread
            self?.inferenceWorker.processAudio(buffer)
        }
    }
}

// Separate worker for background processing
class WakeWordInferenceWorker {
    private let processingQueue = DispatchQueue(label: "wake-word")
    
    func processAudio(_ buffer: AVAudioPCMBuffer) {
        processingQueue.async { [weak self] in
            self?.runInference(buffer)
        }
    }
}

Nonisolated Patterns

nonisolated-methods allow safe access across actor boundaries:

@MainActor
class CompanionManager {
    private nonisolated(unsafe) var pipeline: OpenWakeWordPipeline?
    
    nonisolated func handleWakeWordDetected() {
        // Can be called from any thread/queue
        DispatchQueue.main.async { [weak self] in
            self?.startConversation()
        }
    }
}

Worker Patterns in Practice

Serial Queue Architecture

worker-patterns for real-time processing while maintaining thread safety:

class WakeWordInferenceWorker {
    private let processingQueue = DispatchQueue(
        label: "wake-word-inference",
        qos: .userInitiated
    )
    private var pipeline: OpenWakeWordPipeline?
    
    func processAudioSamples(_ samples: [Float]) {
        processingQueue.async { [weak self] in
            guard let pipeline = self?.pipeline else { return }
            
            if pipeline.detectWakeWord(samples) {
                // Fire callback on main thread
                DispatchQueue.main.async {
                    self?.delegate?.wakeWordDetected()
                }
            }
        }
    }
}

Memory Management

Careful weak references prevent retain cycles between actors:

// Audio callback uses weak self to prevent cycles
audioEngine.inputNode.installTap { [weak self] buffer, _ in
    self?.handleAudioBuffer(buffer)
}

// Worker callbacks also use weak references
processingQueue.async { [weak self] in
    guard let self = self else { return }
    // Safe to use self here
}

Integration with ONNX Runtime

Thread Safety Considerations

onnx-runtime models are not thread-safe and require careful coordination:

class OpenWakeWordPipeline {
    private var melModel: ORTSession?
    private var embeddingModel: ORTSession?
    private var classifierModel: ORTSession?
    
    // All inference must happen on same serial queue
    func detectWakeWord(_ samples: [Float]) -> Bool {
        // This method assumes it's called from a serial queue
        precondition(DispatchQueue.getSpecific(key: processingKey) != nil)
        
        // Safe to use models sequentially
        let melOutput = try melModel?.run(...)
        let embeddingOutput = try embeddingModel?.run(...)
        let score = try classifierModel?.run(...)
        
        return score > threshold
    }
}

Best Practices

Isolation Strategy

  1. MainActor: UI components, user interaction handlers
  2. Background Queues: Heavy computation, I/O operations
  3. Serial Queues: Stateful processing like audio pipelines
  4. Nonisolated: Thread-safe data access and callbacks

Error Handling

// Async methods should handle isolation errors
@MainActor
func startRecording() async {
    do {
        try await audioEngine.start()
        setupWakeWordDetection()
    } catch {
        // Handle on main actor for UI updates
        showError(error)
    }
}

Performance Optimization

  • Use nonisolated(unsafe) sparingly and only for truly thread-safe access
  • Minimize actor hopping with strategic await placement
  • Batch operations to reduce context switching overhead

See also

Tensor Processing

page dédiée →

Mathematical operations on multi-dimensional arrays (tensors) that form the foundation of machine learning inference pipelines. Critical for real-time audio processing where tensor shape transformations and buffer management directly impact performance and accuracy.

Core Concepts

Tensor Dimensions and Shapes

// Common tensor shapes in audio ML pipelines:
let audioSamples: [Float] = [...]        // Shape: [N] - 1D array
let melSpectrogram: Float = [...]    // Shape: [T, F] - Time × Frequency
let batchedInput: [Float] = [...]    // Shape: [B, T, F] - Batch × Time × Frequency

Shape Transformations

Critical operations for preparing data between model stages:

// Squeeze: Remove dimensions of size 1
// [T, 1, 1, 32] → [T, 32]
func squeeze4Dto2D(_ tensor: [[Float]]) -> Float {
    return tensor.map { timeFrame in
        return timeFrame[0][0]  // Extract [32] from [1, 1, 32]
    }
}

// Expand: Add batch dimension
// [16, 96] → [1, 16, 96]  
func addBatchDimension(_ tensor: Float) -> [Float] {
    return [tensor]
}

ONNX Runtime Integration

Input/Output Tensor Handling

// Create input tensor from Swift array
let inputData = Data(bytes: floatArray, count: floatArray.count * 4)
let inputTensor = try ORTValue(
    tensorData: NSMutableData(data: inputData),
    elementType: .float,
    shape: [1, timeFrames, melBins]
)

// Extract output tensor data
let outputTensor = results[outputName]
let outputData = try outputTensor.tensorData() as Data
let outputFloats = outputData.withUnsafeBytes {
    Array($0.bindMemory(to: Float.self))
}

Dynamic Shape Handling

class OpenWakeWordPipeline {
    func processAudioChunk(_ samples: [Float]) -> Float? {
        // Stage 1: Audio → Mel-Spectrogram [N] → [T, 32]
        let melFrames = try melSpectrogramModel.run(samples)
        let squeezedMel = squeeze4Dto2D(melFrames)  // [T, 1, 1, 32] → [T, 32]
        let normalizedMel = squeezedMel.map { $0.map { $0 / 10 + 2 } }
        
        // Stage 2: Mel → Embedding [76, 32] → [96]
        melFrameBuffer.append(contentsOf: normalizedMel)
        guard melFrameBuffer.count >= 76 else { return nil }
        
        let melWindow = Array(melFrameBuffer.suffix(76))
        let batchedMel = [melWindow]  // Add batch dimension: [1, 76, 32]
        let embedding = try embeddingModel.run(batchedMel)
        let flatEmbedding = squeeze3Dto1D(embedding)  // [1, 1, 96] → [96]
        
        // Stage 3: Embeddings → Classification [16, 96] → [1]
        embeddingBuffer.append(flatEmbedding)
        guard embeddingBuffer.count >= 16 else { return nil }
        
        let embeddingSequence = Array(embeddingBuffer.suffix(16))
        let batchedSequence = [embeddingSequence]  // [1, 16, 96]
        let confidence = try classifierModel.run(batchedSequence)
        
        return confidence[0]  // Extract scalar confidence score
    }
}

Real-Time Buffer Management

Rolling Tensor Windows

// Maintain rolling windows at different tensor dimensions
class MultiDimensionalRollingBuffer {
    private var audioBuffer: [Float] = []              // 1D: [N]
    private var melFrameBuffer: Float = []         // 2D: [T, F]  
    private var embeddingBuffer: Float = []        // 2D: [T, E]
    
    func appendAudio(_ samples: [Float]) {
        audioBuffer.append(contentsOf: samples)
        if audioBuffer.count > maxAudioSamples {
            audioBuffer.removeFirst(audioBuffer.count - maxAudioSamples)
        }
    }
    
    func appendMelFrames(_ frames: Float) {
        melFrameBuffer.append(contentsOf: frames)
        if melFrameBuffer.count > maxMelFrames {
            melFrameBuffer.removeFirst(melFrameBuffer.count - maxMelFrames)
        }
    }
    
    func getAudioWindow(_ samples: Int) -> [Float] {
        return Array(audioBuffer.suffix(samples))
    }
    
    func getMelWindow(_ frames: Int) -> Float {
        return Array(melFrameBuffer.suffix(frames))
    }
}

Memory-Efficient Tensor Operations

// Avoid unnecessary copying for large tensors
extension Array where Element == Float {
    func withContiguousStorage<R>(_ body: (UnsafeBufferPointer<Float>) -> R) -> R {
        return self.withUnsafeBufferPointer(body)
    }
}

// Process tensors in-place when possible
func normalizeInPlace(_ tensor: inout Float) {
    for i in 0..<tensor.count {
        for j in 0..<tensor[i].count {
            tensor[i][j] = tensor[i][j] / 10.0 +

Three-Stage Pipeline

page dédiée →

Audio processing architecture used in openwakeword systems that chains together three specialized ONNX models to transform raw audio into wake-word detection decisions. This design separates concerns and enables efficient real-time inference through progressive feature extraction.

Pipeline Architecture

Stage 1: Mel Spectrogram Conversion

  • Input: Raw PCM audio samples (1280 samples = 80ms at 16kHz)
  • Output: Mel-scale spectrogram with 32 frequency bins
  • Function: Converts time-domain audio to frequency-domain representation
  • Tensor Shape: [T, 32] where T varies based on audio chunk size

Stage 2: Feature Embedding

  • Input: Sliding window of 76 mel spectrogram frames
  • Output: 96-dimensional feature embedding vector
  • Function: Extracts semantic audio features from spectral data
  • Tensor Shape: [1, 76, 32, 1] → [96]

Stage 3: Wake-Word Classification

  • Input: Buffer of last 16 feature embeddings
  • Output: Single wake-word detection score
  • Function: Binary classification for specific wake-word
  • Tensor Shape: [1, 16, 96] → [1]

Streaming Implementation

The pipeline operates on continuous audio through rolling-buffers:

  1. Audio Accumulation: Maintain rolling buffer of raw PCM samples
  2. Mel Frame Buffer: Store last 76 mel spectrogram frames for embedding computation
  3. Embedding Buffer: Keep last 16 embeddings for classification
  4. Threshold Detection: Apply 0.5 threshold to classification output

Performance Characteristics

Latency Profile

  • Mel Computation: ~5ms per 80ms chunk
  • Embedding Inference: ~10ms for 76-frame window
  • Classification: ~1ms for 16-embedding input
  • Total Pipeline: ~16ms processing time per chunk

Memory Requirements

  • Raw Audio Buffer: ~3.5KB (1760 samples × 2 bytes)
  • Mel Frame Buffer: ~9.7KB (76 frames × 32 bins × 4 bytes)
  • Embedding Buffer: ~6.1KB (16 embeddings × 96 dims × 4 bytes)
  • Total Working Memory: ~20KB for buffers

Implementation Considerations

Startup Handling

The pipeline requires pre-filling buffers with zeros during initialization to prevent crashes when the rolling windows don't have enough historical data. This results in noisier detection during the first ~1.5 seconds of operation.

Tensor Transformations

Critical shape manipulations between stages:

  • Mel output squeeze: [T, 1, 1, 32] → [T, 32]
  • Preprocessing: Apply x/10 + 2 normalization
  • Embedding squeeze: [1, 96] → [96]

Concurrency Design

Each stage can be processed independently, enabling:

  • Pipelined execution across multiple chunks
  • Separate worker queues for each model
  • swift-concurrency integration with nonisolated workers

Advantages

  1. Modularity: Each stage serves a distinct purpose and can be optimized independently
  2. Efficiency: Progressive dimensionality reduction (audio → spectrogram → embedding → score)
  3. Flexibility: Wake-word classifier can be retrained without touching feature extraction stages
  4. Real-time Performance: Streaming design with predictable memory usage

See also

Voice AI Pipelines

page dédiée →

Comprehensive architecture patterns for building voice-enabled AI applications, covering speech recognition, text-to-speech, and real-time audio processing. Demonstrated through the lesphinx voice game implementation with production-ready error handling and cross-language normalization.

Core Architecture

Voice AI pipelines typically consist of three main components working in concert:

  1. Speech-to-Text (STT): Converting audio input to text
  2. Natural Language Processing: Understanding and generating responses
  3. Text-to-Speech (TTS): Converting text responses back to audio

The key challenge is maintaining state consistency and handling errors gracefully across all three stages while providing real-time user feedback.

Production Implementation Patterns

Web Speech API Integration

Modern voice applications can leverage browser-native capabilities for both speech recognition and synthesis:

// Speech recognition setup
const recognition = new webkitSpeechRecognition();
recognition.continuous = false;
recognition.interimResults = false;
recognition.lang = getCurrentLanguage(); // 'fr-FR' or 'en-US'

recognition.onresult = (event) => {
    const transcript = event.results[0][0].transcript;
    processVoiceInput(transcript);
};

// Text-to-speech synthesis
const synth = window.speechSynthesis;
const utterance = new SpeechSynthesisUtterance(text);
utterance.lang = getCurrentLanguage();
synth.speak(utterance);

Cross-Language Normalization

Voice input requires robust text normalization to handle variations in pronunciation and recognition accuracy:

def normalize_answer(text: str, target_language: str = "fr") -> str:
    """Normalize voice input with longest-match precedence."""
    text = text.lower().strip()
    
    # Language-specific normalization patterns
    fr_patterns = {
        "absolument pas": "non",
        "pas du tout": "non", 
        "bien sûr": "oui",
        "exactement": "oui"
    }
    
    en_patterns = {
        "absolutely not": "no",
        "not at all": "no",
        "of course": "yes", 
        "exactly": "yes"
    }
    
    # Apply longest-match first to avoid partial matches
    patterns = fr_patterns if target_language == "fr" else en_patterns
    for pattern in sorted(patterns.keys(), key=len, reverse=True):
        if pattern in text:
            return patterns[pattern]
    
    return text

Error Handling and Fallback Systems

Production voice pipelines require comprehensive error handling:

class VoiceProcessor:
    def __init__(self):
        self.fallback_questions = [
            "Tell me more about this character.",
            "What else can you share?",
            "Any other details?"
        ]
    
    async def process_voice_input(self, audio_data):
        try:
            # Primary STT processing
            transcript = await self.stt_service.transcribe(audio_data)
            normalized = self.normalize_text(transcript)
            return await self.llm_service.process(normalized)
        
        except STTException as e:
            logger.warning(f"STT failed: {e}")
            # Fallback to text input prompt
            return self.prompt_text_input()
            
        except LLMException as e:
            logger.error(f"LLM processing failed: {e}")
            # Use fallback question
            return self.get_fallback_response()

State Management in Voice Systems

Voice applications require careful state management to handle the asynchronous nature of speech processing:

Game State Synchronization

class VoiceGameEngine:
    def __init__(self):
        self.state = GameState.WAITING
        self.pending_audio = False
        self.last_utterance = None
    
    async def process_voice_turn(self, session_id: str, audio_input: str):
        session = self.get_session(session_id)
        
        # Check if this is a guess confirmation
        if session.last_turn_type == "guess":
            return await self.handle_guess_confirmation(session, audio_input)
        
        # Normal game flow
        normalized_input = normalize_answer(audio_input, session.language)
        return await self.process_answer(session, normalized_input)

Real-Time Feedback

Voice interfaces require immediate feedback to maintain user engagement:

class VoiceController {
    constructor() {
        this.isListening = false;
        this.isProcessing = false;
        this.feedbackTimer = null;
    }
    
    startListening() {
        this.isListening = true;
        this.updateUI('listening');
        this.recognition.start();
        
        // Provide feedback for long processing
        this.feedbackTimer = setTimeout(() => {
            if (this.isProcessing) {
                this.showMessage("Thinking...", "info");
            }
        }, 2000);
    }
    
    async processResult(transcript) {
        this.isProcessing = true;
        this.updateUI('processing');
        
        try {
            const response = await this.sendVoiceAnswer(transcript);
            this.handleGameResponse(response);
        } catch (error) {
            this.showError("Processing failed. Please try again.");
        } finally {
            this.isProcessing = false;
            this.updateUI('ready');
        }
    }
}

Performance Optimization

Streaming Audio Processing

For real-time applications, implement streaming audio processing:

import asyncio
from asyncio import Queue

class StreamingVoiceProcessor:
    def __init__(self):
        self.audio_queue = Queue()
        self.transcript_queue = Queue()
    
    async def stream_audio(self, websocket):
        """Process audio chunks in real-time."""
        while True:
            chunk = await websocket.receive_bytes()
            await self.audio_queue.put(chunk)
    
    async def transcribe_stream(self):
        """Convert audio stream to text continuously."""
        buffer = b""
        while True:
            chunk = await self.audio_queue.get()
            buffer += chunk
            
            if len(buffer) >= self.min_chunk_size:
                transcript = await self.stt_service.transcribe_chunk(buffer)
                if transcript:
                    await self.transcript_queue.put(transcript)
                buffer = b""

Memory Management

Voice applications can accumulate significant memory usage:

class VoiceSessionManager:
    def __init__(self, max_sessions: int = 1000):
        self.sessions = {}
        self.max_sessions = max_sessions
        self.last_activity = {}
    
    def cleanup_stale_sessions(self):
        """Remove inactive sessions to prevent memory bloat."""
        cutoff = time.time() - 3600  # 1 hour timeout
        stale_sessions = [
            sid for sid, last_seen in self.last_activity.items()
            if last_seen < cutoff
        ]
        
        for sid in stale_sessions:
            self.sessions.pop(sid, None)
            self.last_activity.pop(sid, None)

Advanced Patterns

Multi-Modal Integration

Combine voice with other input modalities:

class MultiModalProcessor:
    def __init__(self):
        self.voice_processor = VoiceProcessor()
        self.text_processor = TextProcessor()
        self.gesture_processor = GestureProcessor()
    
    async def process_input(self, input_data):
        # Determine input type and route accordingly
        if input_data.get('audio'):
            return await self.voice_processor.process(input_data['audio'])
        elif input_data.get('text'):
            return await self.text_processor.process(input_data['text'])
        elif input_data.get('gesture'):
            return await self.gesture_processor.process(input_data['gesture'])

Voice Analytics and Monitoring

Track voice system performance:

class VoiceAnalytics:
    def __init__(self):
        self.metrics = {
            'recognition_accuracy': [],
            'processing_latency': [],
            'user_satisfaction': []
        }
    
    def track_interaction(self, start_time, transcript, confidence, user_feedback):
        latency = time.time() - start_time
        self.metrics['processing_latency'].append(latency)
        self.metrics['recognition_accuracy'].append(confidence)
        
        if user_feedback:
            self.metrics['user_satisfaction'].append(user_feedback)

Best Practices

  1. Always Provide Visual Feedback: Users need to know when the system is listening, processing, or has encountered an error
  2. Implement Graceful Fallbacks: Voice recognition will fail; have text input alternatives ready
  3. Normalize Extensively: Handle various ways users might express the same intent
  4. Manage Session State: Keep track of conversation context and user preferences
  5. Monitor Performance: Track recognition accuracy, processing latency, and user satisfaction
  6. Handle Noise Gracefully: Implement confidence thresholds and noise filtering
  7. Support Multiple Languages: Plan for internationalization from the start

Common Pitfalls

  • Over-reliance on Voice: Always provide alternative input methods
  • Poor Error Messages: Users need clear feedback when voice processing fails
  • Memory Leaks: Voice sessions can accumulate significant memory usage
  • Language Confusion: Users may switch languages mid-conversation
  • Processing Delays: Long processing times without feedback frustrate users
  • State Inconsistency: Audio processing is asynchronous and can create race conditions

See also

Voice Effects Processing

page dédiée →

Digital signal processing techniques for transforming synthesized or recorded voice into character voices, particularly for robotic applications and AI character development. Essential for creating distinctive synthetic personalities without relying on voice cloning of existing persons.

Core Audio Effects

Pitch Correction/Autotune

Forces voice onto specific musical scales, creating the characteristic "robotic" singing effect. Can be tuned to pentatonic or minor scales for different emotional tones.

Vocoder Processing

Transforms voice using synthesizer carrier waves, producing the classic robot voice effect. The carrier wave determines the harmonic content while the voice provides the modulation.

Formant Shifting

Modifies vocal tract resonance characteristics to create smaller, metallic, or android-like voices. Essential for non-human character voices.

Spatial Effects

  • Chorus: Creates width and ensemble effect
  • Flanger/Phaser: Adds movement and musical quality
  • Reverb/Delay: Provides spatial character

Dynamic Processing

  • Compression: Evens out volume levels for consistent output
  • EQ: Optimizes frequency response for small speakers
  • Limiting: Prevents clipping and distortion

Digital Degradation

  • Bitcrushing: Reduces bit depth for retro digital artifacts
  • Saturation: Adds harmonic distortion for warmth or aggression

Implementation Strategies

Real-time Processing

Direct microphone input through effects chain with live output. Requires careful latency management and feedback prevention.

TTS + Effects Pipeline

  1. Text generation from LLM
  2. Voice synthesis (Gradium/ElevenLabs)
  3. Effects processing
  4. Playback through robot speaker

Hybrid Approach

Pre-process common phrases with effects, use real-time for dynamic content.

Platform Integration

ReachyMini Implementation

Uses mini.media.push_audio_sample() API for custom audio playback. Recommends external processing due to Raspberry Pi computational limitations.

Development Considerations

  • Latency requirements: <300ms for natural conversation
  • Processing power: Complex effects chains require dedicated hardware
  • Feedback prevention: Careful microphone/speaker isolation
  • Effect presets: Create character-specific processing chains

Character Voice Archetypes

Daft Punk Style

Combination of autotune, vocoder, chorus, and compression for electronic music robot aesthetic.

Protocol Droid

Formal speech patterns with slight metallic coloration and measured delivery.

Beep/Chirp Synthesis

Pure synthetic tones and frequency sweeps for non-verbal robot communication.

Mechanical Voice

Bitcrushing, formant shifting, and servo-like artifacts for industrial robot character.

See also

Wake Word Detection

page dédiée →

Audio processing technique that enables devices to activate or respond to specific spoken trigger phrases while maintaining low-power, always-on listening capabilities. Essential for voice-activated AI systems and smart devices that need to differentiate activation commands from ambient conversation.

Technical Architecture

Core Components

  • Audio Buffer Management: Continuous audio stream processing with rolling buffers
  • Feature Extraction: Mel-spectrogram analysis for audio pattern recognition
  • Model Inference: Lightweight neural networks (typically ONNX) for real-time detection
  • Threshold Management: Configurable confidence levels to balance accuracy and false positives

Implementation Patterns

  • Edge Processing: Local model inference to avoid cloud dependency and privacy concerns
  • Low Latency: Sub-second detection times for natural user interaction
  • Power Efficiency: Optimized for continuous operation without significant battery drain

openWakeWord

Open-source framework providing:

  • Custom wake word training capabilities
  • ONNX model deployment for cross-platform compatibility
  • 16kHz audio processing with minimal computational overhead
  • Integration with various audio frameworks and platforms

Platform-Specific Solutions

  • Reachy Mini: "Hey Reachy" detection integrated into the robotics platform
  • Smart Speakers: "Alexa", "Hey Google", "Hey Siri" implementations
  • Mobile Devices: Always-on voice activation for assistants and apps

Application Domains

Robotics Platforms

  • reachy-mini uses wake word detection for voice activation
  • Enables hands-free robot interaction and command initiation
  • Combines with directional microphone arrays for spatial awareness

Smart Home Integration

  • Device activation without physical interaction
  • Multi-device coordination with unique wake phrases
  • Privacy-preserving local processing

Mobile Applications

  • Voice memo activation
  • Navigation and accessibility features
  • Background app triggering

Technical Challenges

Accuracy vs. Efficiency Trade-offs

  • False Positives: Unwanted activations from similar-sounding phrases
  • False Negatives: Missed detections due to accent, noise, or pronunciation variations
  • Resource Usage: Balancing detection accuracy with computational requirements

Environmental Considerations

  • Noise Robustness: Performance in challenging acoustic environments
  • Multi-Speaker Scenarios: Distinguishing target speakers from background voices
  • Acoustic Variability: Handling different room acoustics and distances

Privacy and Security

Local Processing Benefits

  • Audio processing without cloud transmission
  • Reduced privacy concerns for always-listening devices
  • Lower latency and offline capability

Security Considerations

  • Protection against adversarial audio attacks
  • Secure model updates and validation
  • User control over wake word sensitivity and activation

Integration with Voice AI Pipelines

Wake word detection typically serves as the entry point for more complex voice-ai systems:

  1. Detection: Wake word triggers system activation
  2. Recording: Full audio capture begins after detection
  3. Processing: Speech-to-text conversion of command or query
  4. Response: AI processing and text-to-speech output
  5. Return: System returns to wake word listening state

See also

Whisper Transcription Workflows

page dédiée →

Automated audio-to-text processing system using OpenAI's Whisper model for capturing knowledge from conferences, meetings, and voice recordings. Essential component of comprehensive knowledge ingestion pipelines.

Core Architecture

Local Processing: Use MLX Whisper for on-device transcription to maintain privacy and avoid cloud service costs for large audio volumes.

Batch Processing: Design workflows to handle multiple audio files efficiently with appropriate resource management for GPU/CPU intensive transcription tasks.

Integration Pipeline: Connect transcription output to knowledge management systems for further processing and integration.

Implementation Pattern

Brain Wiki Example:

# process-recordings.sh workflow
1. Monitor iCloud Drive/brain-wiki-inbox/ for audio files
2. Move audio to raw/talks/pending/
3. Run MLX Whisper transcription: audio.mp3 → audio_transcript.txt
4. Generate metadata: date, duration, source, confidence scores
5. Create structured markdown with transcript and metadata
6. Queue for LLM processing and wiki integration

Technical Stack:

  • MLX Whisper: Apple Silicon optimized for fast local processing
  • Voice Memos: Native iOS app for recording with iCloud sync
  • iCloud Drive: Seamless transfer from mobile recording to desktop processing
  • Batch Processing: Handle multiple recordings in scheduled runs

Content Types

Conference Talks: Capture keynotes, technical presentations, Q&A sessions with speaker identification and slide correlation.

Voice Memos: Personal reflections, ideas, walking thoughts converted to searchable text for later integration.

Meeting Recordings: Team discussions, client calls, interview transcripts with participant identification.

Podcast Processing: Extract insights from technical podcasts for knowledge base integration.

Quality Considerations

Audio Preprocessing: Noise reduction, volume normalization, and format standardization for optimal transcription accuracy.

Confidence Scoring: Track Whisper confidence levels to identify segments requiring human review or re-processing.

Speaker Identification: Use additional models or manual annotation for multi-speaker recordings.

Punctuation and Formatting: Post-process raw transcription for readability and semantic structure.

Mobile Integration

Capture Workflow:

  1. Use Voice Memos app during conferences or meetings
  2. Share recorded audio to iOS Shortcut "Brain Wiki Audio"
  3. Shortcut saves to iCloud Drive with structured naming
  4. Desktop processing automatically handles transcription and integration

Naming Convention: YYYY-MM-DD-event-speaker-topic.m4a for organized batch processing

Processing Optimization

Resource Management: MLX Whisper utilizes Apple Silicon efficiently but requires memory consideration for long recordings.

Batch Scheduling: Process recordings during low-activity periods to avoid interfering with interactive work.

Error Recovery: Handle corrupted audio files, network interruptions, and processing failures gracefully.

Quality Thresholds: Set minimum confidence scores for automatic processing vs. human review queues.

Integration Patterns

Structured Output: Generate markdown with frontmatter containing metadata, confidence scores, and processing notes.

Semantic Segmentation: Break long transcripts into logical sections for better knowledge base integration.

Cross-Reference Generation: Automatically identify concepts, people, and topics mentioned for wiki linking.

Summary Generation: Use LLM processing to create abstracts and key insights from raw transcripts.

Success Metrics

Transcription Accuracy: Word error rate and semantic understanding quality Processing Speed: Time from audio capture to searchable text availability Integration Rate: Percentage of transcripts successfully incorporated into knowledge base User Adoption: Frequency of audio capture vs. alternative note-taking methods

Whisper transcription workflows enable comprehensive knowledge capture from audio sources, particularly valuable for conference attendance, meeting documentation, and voice-based reflection practices.

See also