~/wiki

Cached Retrieval

Confiance : high
cached-retrievalperformance-optimizationsession-statememory-managementretrieval-architecturestreamlit-cachingcorpus-loadingtf-idf-caching

Performance optimization strategy that caches retrieval components and search results to minimize computational overhead and improve response times. Particularly important for applications with repeated queries against the same document corpus.

Core Principles

Multi-Level Caching

Effective caching strategies operate at multiple levels:

  • Retriever instances: Cache expensive model loading and database connections
  • Corpus data: Cache loaded document collections in memory
  • Search indices: Pre-compute and cache TF-IDF matrices, embeddings, or search structures
  • Query results: Cache frequent query results for instant response

Cache Granularity

Different caching levels serve different purposes:

  • Session-level: Data persistent across user interactions within a session
  • Application-level: Shared data across multiple users and sessions
  • Query-level: Individual search results with parameter-based keys
  • Component-level: Individual retrieval components cached independently

Implementation Patterns

Streamlit Integration

The assistant-rh project demonstrates effective caching with streamlit-ui:

Retriever Instance Caching

@st.cache_resource
def load_csv_retriever(csv_path):
    """Cache retriever instance for session reuse"""
    return CSVRetriever(csv_path)

Corpus Statistics Caching

@st.cache_data
def get_corpus_stats(csv_path):
    """Cache lightweight corpus overview"""
    return {"total_chunks": count, "sources": unique_sources}

CSV-Based Caching

For csv-based-retrieval systems:

  • Single CSV read: File loaded once per session, cached in memory
  • Tokenization caching: Pre-computed TF-IDF weights stored for reuse
  • Index persistence: Search structures maintained across queries
  • Metadata caching: Source information and statistics cached separately

Performance Benefits

Response Time Optimization

  • Eliminates redundant I/O: Avoid repeated file reads or database queries
  • Reduces computation: Skip expensive tokenization and similarity calculations
  • Improves user experience: Near-instant response for cached queries
  • Scales with usage: Frequently accessed data becomes faster over time

Resource Efficiency

  • Memory optimization: Share expensive data structures across queries
  • CPU savings: Avoid repeated computation of search indices
  • Network efficiency: Reduce database and API calls
  • Startup acceleration: Warm caches improve application responsiveness

Cache Management

Invalidation Strategies

  • Time-based expiration: Cache entries expire after fixed duration
  • Content-based invalidation: Cache cleared when underlying data changes
  • Size-based eviction: Least recently used (LRU) eviction for memory management
  • Manual refresh: User-triggered cache clearing for data updates

Memory Considerations

  • Cache size limits: Prevent unbounded memory growth
  • Selective caching: Cache only frequently accessed or expensive data
  • Compression: Reduce memory footprint of cached data
  • Monitoring: Track cache hit rates and memory usage

Use Cases

Interactive Applications

Cached retrieval is essential for:

  • Chat interfaces: Fast response to user queries
  • Search applications: Instant results for common searches
  • Dashboard applications: Quick data visualization updates
  • Development tools: Rapid iteration during testing and debugging

Large Corpus Applications

  • Document search: Cache expensive corpus loading and indexing
  • Knowledge bases: Persistent search indices for large datasets
  • Multi-user systems: Shared caches across concurrent users
  • Real-time applications: Pre-computed results for immediate response

Architecture Patterns

Layered Caching

User Query → Query Cache → Search Index Cache → Corpus Cache → Storage

Each layer provides different optimization benefits:

  • Query cache: Exact query matches for instant response
  • Search index cache: Pre-computed similarity structures
  • Corpus cache: In-memory document storage
  • Storage: Persistent data source (files, databases)

Hybrid Strategies

  • Hot/Cold data: Frequently accessed data in fast cache, archive in slow storage
  • Prefetching: Anticipate likely queries and pre-compute results
  • Lazy loading: Cache data on first access, maintain for future use
  • Background refresh: Update caches asynchronously to maintain freshness

Implementation Considerations

Cache Key Design

  • Parameter sensitivity: Include all relevant query parameters in cache keys
  • Versioning: Handle data version changes in cache keys
  • Normalization: Consistent key format for equivalent queries
  • Collision avoidance: Unique keys for different query contexts

Consistency Management

  • Cache coherence: Ensure cached data remains consistent with source
  • Update propagation: Coordinate cache updates across system components
  • Conflict resolution: Handle simultaneous updates to cached data
  • Rollback handling: Recover from failed cache updates

Monitoring and Debugging

Performance Metrics

  • Hit rate: Percentage of queries served from cache
  • Response time: Latency comparison between cached and uncached queries
  • Memory usage: Cache size and growth patterns
  • Eviction frequency: How often cache entries are removed

Debugging Tools

  • Cache inspection: Ability to examine cached data and statistics
  • Cache warming: Tools to pre-populate caches for testing
  • Invalidation testing: Verify cache clearing works correctly
  • Performance profiling: Identify cache bottlenecks and optimization opportunities

See also