Voice AI Pipelines
Comprehensive architecture patterns for building voice-enabled AI applications, covering speech recognition, text-to-speech, and real-time audio processing. Demonstrated through the lesphinx voice game implementation with production-ready error handling and cross-language normalization.
Core Architecture
Voice AI pipelines typically consist of three main components working in concert:
- Speech-to-Text (STT): Converting audio input to text
- Natural Language Processing: Understanding and generating responses
- Text-to-Speech (TTS): Converting text responses back to audio
The key challenge is maintaining state consistency and handling errors gracefully across all three stages while providing real-time user feedback.
Production Implementation Patterns
Web Speech API Integration
Modern voice applications can leverage browser-native capabilities for both speech recognition and synthesis:
// Speech recognition setup
const recognition = new webkitSpeechRecognition();
recognition.continuous = false;
recognition.interimResults = false;
recognition.lang = getCurrentLanguage(); // 'fr-FR' or 'en-US'
recognition.onresult = (event) => {
const transcript = event.results[0][0].transcript;
processVoiceInput(transcript);
};
// Text-to-speech synthesis
const synth = window.speechSynthesis;
const utterance = new SpeechSynthesisUtterance(text);
utterance.lang = getCurrentLanguage();
synth.speak(utterance);
Cross-Language Normalization
Voice input requires robust text normalization to handle variations in pronunciation and recognition accuracy:
def normalize_answer(text: str, target_language: str = "fr") -> str:
"""Normalize voice input with longest-match precedence."""
text = text.lower().strip()
# Language-specific normalization patterns
fr_patterns = {
"absolument pas": "non",
"pas du tout": "non",
"bien sûr": "oui",
"exactement": "oui"
}
en_patterns = {
"absolutely not": "no",
"not at all": "no",
"of course": "yes",
"exactly": "yes"
}
# Apply longest-match first to avoid partial matches
patterns = fr_patterns if target_language == "fr" else en_patterns
for pattern in sorted(patterns.keys(), key=len, reverse=True):
if pattern in text:
return patterns[pattern]
return text
Error Handling and Fallback Systems
Production voice pipelines require comprehensive error handling:
class VoiceProcessor:
def __init__(self):
self.fallback_questions = [
"Tell me more about this character.",
"What else can you share?",
"Any other details?"
]
async def process_voice_input(self, audio_data):
try:
# Primary STT processing
transcript = await self.stt_service.transcribe(audio_data)
normalized = self.normalize_text(transcript)
return await self.llm_service.process(normalized)
except STTException as e:
logger.warning(f"STT failed: {e}")
# Fallback to text input prompt
return self.prompt_text_input()
except LLMException as e:
logger.error(f"LLM processing failed: {e}")
# Use fallback question
return self.get_fallback_response()
State Management in Voice Systems
Voice applications require careful state management to handle the asynchronous nature of speech processing:
Game State Synchronization
class VoiceGameEngine:
def __init__(self):
self.state = GameState.WAITING
self.pending_audio = False
self.last_utterance = None
async def process_voice_turn(self, session_id: str, audio_input: str):
session = self.get_session(session_id)
# Check if this is a guess confirmation
if session.last_turn_type == "guess":
return await self.handle_guess_confirmation(session, audio_input)
# Normal game flow
normalized_input = normalize_answer(audio_input, session.language)
return await self.process_answer(session, normalized_input)
Real-Time Feedback
Voice interfaces require immediate feedback to maintain user engagement:
class VoiceController {
constructor() {
this.isListening = false;
this.isProcessing = false;
this.feedbackTimer = null;
}
startListening() {
this.isListening = true;
this.updateUI('listening');
this.recognition.start();
// Provide feedback for long processing
this.feedbackTimer = setTimeout(() => {
if (this.isProcessing) {
this.showMessage("Thinking...", "info");
}
}, 2000);
}
async processResult(transcript) {
this.isProcessing = true;
this.updateUI('processing');
try {
const response = await this.sendVoiceAnswer(transcript);
this.handleGameResponse(response);
} catch (error) {
this.showError("Processing failed. Please try again.");
} finally {
this.isProcessing = false;
this.updateUI('ready');
}
}
}
Performance Optimization
Streaming Audio Processing
For real-time applications, implement streaming audio processing:
import asyncio
from asyncio import Queue
class StreamingVoiceProcessor:
def __init__(self):
self.audio_queue = Queue()
self.transcript_queue = Queue()
async def stream_audio(self, websocket):
"""Process audio chunks in real-time."""
while True:
chunk = await websocket.receive_bytes()
await self.audio_queue.put(chunk)
async def transcribe_stream(self):
"""Convert audio stream to text continuously."""
buffer = b""
while True:
chunk = await self.audio_queue.get()
buffer += chunk
if len(buffer) >= self.min_chunk_size:
transcript = await self.stt_service.transcribe_chunk(buffer)
if transcript:
await self.transcript_queue.put(transcript)
buffer = b""
Memory Management
Voice applications can accumulate significant memory usage:
class VoiceSessionManager:
def __init__(self, max_sessions: int = 1000):
self.sessions = {}
self.max_sessions = max_sessions
self.last_activity = {}
def cleanup_stale_sessions(self):
"""Remove inactive sessions to prevent memory bloat."""
cutoff = time.time() - 3600 # 1 hour timeout
stale_sessions = [
sid for sid, last_seen in self.last_activity.items()
if last_seen < cutoff
]
for sid in stale_sessions:
self.sessions.pop(sid, None)
self.last_activity.pop(sid, None)
Advanced Patterns
Multi-Modal Integration
Combine voice with other input modalities:
class MultiModalProcessor:
def __init__(self):
self.voice_processor = VoiceProcessor()
self.text_processor = TextProcessor()
self.gesture_processor = GestureProcessor()
async def process_input(self, input_data):
# Determine input type and route accordingly
if input_data.get('audio'):
return await self.voice_processor.process(input_data['audio'])
elif input_data.get('text'):
return await self.text_processor.process(input_data['text'])
elif input_data.get('gesture'):
return await self.gesture_processor.process(input_data['gesture'])
Voice Analytics and Monitoring
Track voice system performance:
class VoiceAnalytics:
def __init__(self):
self.metrics = {
'recognition_accuracy': [],
'processing_latency': [],
'user_satisfaction': []
}
def track_interaction(self, start_time, transcript, confidence, user_feedback):
latency = time.time() - start_time
self.metrics['processing_latency'].append(latency)
self.metrics['recognition_accuracy'].append(confidence)
if user_feedback:
self.metrics['user_satisfaction'].append(user_feedback)
Best Practices
- Always Provide Visual Feedback: Users need to know when the system is listening, processing, or has encountered an error
- Implement Graceful Fallbacks: Voice recognition will fail; have text input alternatives ready
- Normalize Extensively: Handle various ways users might express the same intent
- Manage Session State: Keep track of conversation context and user preferences
- Monitor Performance: Track recognition accuracy, processing latency, and user satisfaction
- Handle Noise Gracefully: Implement confidence thresholds and noise filtering
- Support Multiple Languages: Plan for internationalization from the start
Common Pitfalls
- Over-reliance on Voice: Always provide alternative input methods
- Poor Error Messages: Users need clear feedback when voice processing fails
- Memory Leaks: Voice sessions can accumulate significant memory usage
- Language Confusion: Users may switch languages mid-conversation
- Processing Delays: Long processing times without feedback frustrate users
- State Inconsistency: Audio processing is asynchronous and can create race conditions
See also
- llm-integration-patterns
- real-time-audio-processing
- web-speech-api
- cross-language-normalization
- lesphinx