TF-IDF (Term Frequency-Inverse Document Frequency)
Classic information retrieval technique that weights terms based on their frequency within individual documents and their rarity across the entire corpus. Fundamental to traditional search systems and still relevant for lightweight retrieval implementations.
Mathematical Foundation
Term Frequency (TF): Measures how frequently a term appears in a document, often normalized by document length.
Inverse Document Frequency (IDF): Measures how rare a term is across the entire corpus, giving higher weight to discriminative terms.
Combined Score: TF-IDF = TF(term, document) × IDF(term, corpus)
Implementation Patterns
Basic TF-IDF Calculation
import math
from collections import Counter
def compute_tf(term_counts, total_terms):
return {term: count/total_terms for term, count in term_counts.items()}
def compute_idf(term, all_documents):
containing_docs = sum(1 for doc in all_documents if term in doc)
return math.log(len(all_documents) / containing_docs)
Document Similarity Scoring
- Query Processing: Convert user query into TF-IDF weighted term vector
- Document Matching: Compute cosine similarity between query and document vectors
- Result Ranking: Sort documents by similarity score in descending order
Use Cases in RAG Systems
CSV-Based Retrieval
Lightweight alternative to vector embeddings for small document corpora, as demonstrated in assistant-rh:
- Preprocessing: Build TF-IDF indices during CSV loading
- Query Processing: Score document chunks against user queries
- Fallback Architecture: Secondary retrieval when vector databases unavailable
Hybrid Search Systems
- Lexical Component: TF-IDF provides keyword-based matching
- Semantic Component: Combined with dense embeddings for comprehensive retrieval
- Score Fusion: Weighted combination of lexical and semantic similarity scores
Advantages
Interpretability: Clear mathematical foundation and explainable scoring mechanism.
Computational Efficiency: Fast processing for medium-sized corpora without GPU requirements.
No Training Required: Works out-of-the-box without pre-trained models or embedding generation.
Language Agnostic: Effective across different languages with appropriate tokenization.
Limitations
Semantic Gap: Cannot capture semantic similarity between synonymous terms.
Context Insensitivity: Single-term focus misses multi-word concepts and phrasal meaning.
Vocabulary Mismatch: Query terms must match document terms exactly for retrieval.
Sparse Representations: High-dimensional but sparse vectors can be inefficient for large corpora.
Modern Applications
Document Preprocessing
- Feature Extraction: Generate term importance weights for downstream ML models
- Keyword Extraction: Identify most discriminative terms in document collections
- Content Analysis: Discover topics and themes through high TF-IDF terms
Search Engine Components
- Query Expansion: Identify related terms based on corpus-wide term statistics
- Result Explanation: Highlight matching terms with their importance scores
- Quality Assessment: Evaluate document relevance based on term overlap
Integration with Neural Methods
- Initial Retrieval: Fast first-stage retrieval before neural reranking
- Feature Engineering: TF-IDF scores as input features for learning-to-rank models
- Baseline Comparison: Performance benchmark for evaluating semantic search systems
Implementation Considerations
Tokenization Strategy: Choice of tokenization (word-level, subword, n-grams) significantly impacts performance.
Stopword Filtering: Remove common words that provide little discriminative value.
Normalization Schemes: Different TF and IDF calculation variants (log normalization, sublinear scaling).
Memory Management: Efficient sparse matrix representations for large vocabulary sizes.
See also
- csv-based-retrieval
- information-retrieval
- hybrid-retrieval-systems
- lexical-search
- assistant-rh