~/wiki

Voice AI Pipelines

Confiance : high
voice-aispeech-recognitiontext-to-speechweb-speech-apiaudio-processingreal-time-voicepipeline-architecturecross-language-normalizationvoice-uiclick-to-speakerror-handlingfallback-systemsbrowser-compatibilityvoice-controlsproduction-voice-systemsvoice-state-managementvoice-feedback-loops

Comprehensive architecture patterns for building voice-enabled AI applications, covering speech recognition, text-to-speech, and real-time audio processing. Demonstrated through the lesphinx voice game implementation with production-ready error handling and cross-language normalization.

Core Architecture

Voice AI pipelines typically consist of three main components working in concert:

  1. Speech-to-Text (STT): Converting audio input to text
  2. Natural Language Processing: Understanding and generating responses
  3. Text-to-Speech (TTS): Converting text responses back to audio

The key challenge is maintaining state consistency and handling errors gracefully across all three stages while providing real-time user feedback.

Production Implementation Patterns

Web Speech API Integration

Modern voice applications can leverage browser-native capabilities for both speech recognition and synthesis:

// Speech recognition setup
const recognition = new webkitSpeechRecognition();
recognition.continuous = false;
recognition.interimResults = false;
recognition.lang = getCurrentLanguage(); // 'fr-FR' or 'en-US'

recognition.onresult = (event) => {
    const transcript = event.results[0][0].transcript;
    processVoiceInput(transcript);
};

// Text-to-speech synthesis
const synth = window.speechSynthesis;
const utterance = new SpeechSynthesisUtterance(text);
utterance.lang = getCurrentLanguage();
synth.speak(utterance);

Cross-Language Normalization

Voice input requires robust text normalization to handle variations in pronunciation and recognition accuracy:

def normalize_answer(text: str, target_language: str = "fr") -> str:
    """Normalize voice input with longest-match precedence."""
    text = text.lower().strip()
    
    # Language-specific normalization patterns
    fr_patterns = {
        "absolument pas": "non",
        "pas du tout": "non", 
        "bien sûr": "oui",
        "exactement": "oui"
    }
    
    en_patterns = {
        "absolutely not": "no",
        "not at all": "no",
        "of course": "yes", 
        "exactly": "yes"
    }
    
    # Apply longest-match first to avoid partial matches
    patterns = fr_patterns if target_language == "fr" else en_patterns
    for pattern in sorted(patterns.keys(), key=len, reverse=True):
        if pattern in text:
            return patterns[pattern]
    
    return text

Error Handling and Fallback Systems

Production voice pipelines require comprehensive error handling:

class VoiceProcessor:
    def __init__(self):
        self.fallback_questions = [
            "Tell me more about this character.",
            "What else can you share?",
            "Any other details?"
        ]
    
    async def process_voice_input(self, audio_data):
        try:
            # Primary STT processing
            transcript = await self.stt_service.transcribe(audio_data)
            normalized = self.normalize_text(transcript)
            return await self.llm_service.process(normalized)
        
        except STTException as e:
            logger.warning(f"STT failed: {e}")
            # Fallback to text input prompt
            return self.prompt_text_input()
            
        except LLMException as e:
            logger.error(f"LLM processing failed: {e}")
            # Use fallback question
            return self.get_fallback_response()

State Management in Voice Systems

Voice applications require careful state management to handle the asynchronous nature of speech processing:

Game State Synchronization

class VoiceGameEngine:
    def __init__(self):
        self.state = GameState.WAITING
        self.pending_audio = False
        self.last_utterance = None
    
    async def process_voice_turn(self, session_id: str, audio_input: str):
        session = self.get_session(session_id)
        
        # Check if this is a guess confirmation
        if session.last_turn_type == "guess":
            return await self.handle_guess_confirmation(session, audio_input)
        
        # Normal game flow
        normalized_input = normalize_answer(audio_input, session.language)
        return await self.process_answer(session, normalized_input)

Real-Time Feedback

Voice interfaces require immediate feedback to maintain user engagement:

class VoiceController {
    constructor() {
        this.isListening = false;
        this.isProcessing = false;
        this.feedbackTimer = null;
    }
    
    startListening() {
        this.isListening = true;
        this.updateUI('listening');
        this.recognition.start();
        
        // Provide feedback for long processing
        this.feedbackTimer = setTimeout(() => {
            if (this.isProcessing) {
                this.showMessage("Thinking...", "info");
            }
        }, 2000);
    }
    
    async processResult(transcript) {
        this.isProcessing = true;
        this.updateUI('processing');
        
        try {
            const response = await this.sendVoiceAnswer(transcript);
            this.handleGameResponse(response);
        } catch (error) {
            this.showError("Processing failed. Please try again.");
        } finally {
            this.isProcessing = false;
            this.updateUI('ready');
        }
    }
}

Performance Optimization

Streaming Audio Processing

For real-time applications, implement streaming audio processing:

import asyncio
from asyncio import Queue

class StreamingVoiceProcessor:
    def __init__(self):
        self.audio_queue = Queue()
        self.transcript_queue = Queue()
    
    async def stream_audio(self, websocket):
        """Process audio chunks in real-time."""
        while True:
            chunk = await websocket.receive_bytes()
            await self.audio_queue.put(chunk)
    
    async def transcribe_stream(self):
        """Convert audio stream to text continuously."""
        buffer = b""
        while True:
            chunk = await self.audio_queue.get()
            buffer += chunk
            
            if len(buffer) >= self.min_chunk_size:
                transcript = await self.stt_service.transcribe_chunk(buffer)
                if transcript:
                    await self.transcript_queue.put(transcript)
                buffer = b""

Memory Management

Voice applications can accumulate significant memory usage:

class VoiceSessionManager:
    def __init__(self, max_sessions: int = 1000):
        self.sessions = {}
        self.max_sessions = max_sessions
        self.last_activity = {}
    
    def cleanup_stale_sessions(self):
        """Remove inactive sessions to prevent memory bloat."""
        cutoff = time.time() - 3600  # 1 hour timeout
        stale_sessions = [
            sid for sid, last_seen in self.last_activity.items()
            if last_seen < cutoff
        ]
        
        for sid in stale_sessions:
            self.sessions.pop(sid, None)
            self.last_activity.pop(sid, None)

Advanced Patterns

Multi-Modal Integration

Combine voice with other input modalities:

class MultiModalProcessor:
    def __init__(self):
        self.voice_processor = VoiceProcessor()
        self.text_processor = TextProcessor()
        self.gesture_processor = GestureProcessor()
    
    async def process_input(self, input_data):
        # Determine input type and route accordingly
        if input_data.get('audio'):
            return await self.voice_processor.process(input_data['audio'])
        elif input_data.get('text'):
            return await self.text_processor.process(input_data['text'])
        elif input_data.get('gesture'):
            return await self.gesture_processor.process(input_data['gesture'])

Voice Analytics and Monitoring

Track voice system performance:

class VoiceAnalytics:
    def __init__(self):
        self.metrics = {
            'recognition_accuracy': [],
            'processing_latency': [],
            'user_satisfaction': []
        }
    
    def track_interaction(self, start_time, transcript, confidence, user_feedback):
        latency = time.time() - start_time
        self.metrics['processing_latency'].append(latency)
        self.metrics['recognition_accuracy'].append(confidence)
        
        if user_feedback:
            self.metrics['user_satisfaction'].append(user_feedback)

Best Practices

  1. Always Provide Visual Feedback: Users need to know when the system is listening, processing, or has encountered an error
  2. Implement Graceful Fallbacks: Voice recognition will fail; have text input alternatives ready
  3. Normalize Extensively: Handle various ways users might express the same intent
  4. Manage Session State: Keep track of conversation context and user preferences
  5. Monitor Performance: Track recognition accuracy, processing latency, and user satisfaction
  6. Handle Noise Gracefully: Implement confidence thresholds and noise filtering
  7. Support Multiple Languages: Plan for internationalization from the start

Common Pitfalls

  • Over-reliance on Voice: Always provide alternative input methods
  • Poor Error Messages: Users need clear feedback when voice processing fails
  • Memory Leaks: Voice sessions can accumulate significant memory usage
  • Language Confusion: Users may switch languages mid-conversation
  • Processing Delays: Long processing times without feedback frustrate users
  • State Inconsistency: Audio processing is asynchronous and can create race conditions

See also