~/wiki

CSV-Based Retrieval

Confiance : high
csv-retrievaltf-idftext-searchlightweight-ragcorpus-processingdocument-chunksstreamlit-integrationcached-retrievalperformance-optimizationfallback-architecturetokenizationsimilarity-search

Lightweight retrieval approach using CSV files to store pre-processed document chunks with TF-IDF based similarity search. Particularly effective for small to medium-sized document corpora where vector database overhead is unnecessary.

Core Architecture

Data Structure

  • CSV format: Document chunks stored as rows with metadata columns
  • Standardized schema: Consistent chunk representation across different corpora
  • Source tracking: Metadata linking chunks back to original documents
  • Pre-processing: Chunks already cleaned, segmented, and formatted for search

Search Implementation

  • TF-IDF tokenization: Term frequency analysis for relevance scoring
  • Similarity calculation: Fast text-based matching without vector embeddings
  • Ranked results: Sorted by relevance score with configurable result limits
  • Chunk objects: Standardized return format compatible with vector retrievers

Performance Characteristics

Caching Strategy

The assistant-rh implementation demonstrates effective caching patterns:

  • Single CSV read: File loaded once per session and cached in memory
  • Tokenization caching: Pre-computed TF-IDF weights stored for reuse
  • Corpus statistics: Overview metrics cached separately from search index
  • Session persistence: Retriever instance reused across multiple queries

Scalability Considerations

  • Memory footprint: Entire corpus held in memory for fast access
  • Search latency: Sub-second response times for typical document sizes
  • Corpus size limits: Practical upper bound around 1000-10000 chunks
  • Startup time: Initial CSV parsing adds session initialization delay

Integration Patterns

Streamlit UI Integration

Effective patterns for web application integration:

  • Sidebar configuration: Corpus selection and statistics display
  • Fallback handling: Graceful degradation when other retrievers unavailable
  • Result formatting: Consistent chunk display across retrieval methods
  • Error recovery: Robust handling of CSV parsing and search failures

Interface Consistency

Key architectural principle ensuring interchangeable retrievers:

  • Standardized return types: All retrievers return Chunk objects
  • Common parameters: Consistent search API across backends
  • Metadata preservation: Source information maintained through search pipeline
  • Error handling: Uniform exception patterns for UI error display

Use Cases

Lightweight RAG Applications

Ideal scenarios for CSV-based retrieval:

  • Small document corpora: 100-1000 pre-processed chunks
  • Rapid prototyping: Quick setup without vector database infrastructure
  • Demo applications: Simplified deployment with file-based storage
  • Development testing: Local development without external dependencies

Hybrid Architectures

CSV retrieval as part of larger systems:

  • Fallback mechanism: Backup when vector databases unavailable
  • Development phase: Initial implementation before vector migration
  • A/B testing: Comparison baseline for vector search performance
  • Edge deployment: Resource-constrained environments

Implementation Examples

The assistant-rh project used CSV retrieval for:

  • 130 decree chunks: French legal documents from corpus_130.csv
  • TF-IDF search: Term-based relevance without embeddings
  • Streamlit chat: Interactive Q&A with retrieved context
  • Performance optimization: Cached loading with session persistence

Advantages and Limitations

Advantages

  • Simple deployment: No database setup or vector infrastructure required
  • Fast development: Rapid iteration with file-based storage
  • Transparent debugging: CSV contents easily inspected and modified
  • Low resource usage: Minimal memory and compute requirements
  • Portable: Easy backup, version control, and distribution

Limitations

  • Scalability ceiling: Performance degrades with large corpora
  • Memory constraints: Entire dataset must fit in application memory
  • Limited search sophistication: No semantic similarity or neural ranking
  • Update complexity: Modifying corpus requires file regeneration
  • Concurrent access: File-based storage not suitable for multi-user systems

See also