~/wiki

TF-IDF (Term Frequency-Inverse Document Frequency)

Confiance : high
tf-idfinformation-retrievaltext-searchdocument-rankinglexical-similaritytraditional-irterm-weightingcorpus-statistics

Classic information retrieval technique that weights terms based on their frequency within individual documents and their rarity across the entire corpus. Fundamental to traditional search systems and still relevant for lightweight retrieval implementations.

Mathematical Foundation

Term Frequency (TF): Measures how frequently a term appears in a document, often normalized by document length.

Inverse Document Frequency (IDF): Measures how rare a term is across the entire corpus, giving higher weight to discriminative terms.

Combined Score: TF-IDF = TF(term, document) × IDF(term, corpus)

Implementation Patterns

Basic TF-IDF Calculation

import math
from collections import Counter

def compute_tf(term_counts, total_terms):
    return {term: count/total_terms for term, count in term_counts.items()}

def compute_idf(term, all_documents):
    containing_docs = sum(1 for doc in all_documents if term in doc)
    return math.log(len(all_documents) / containing_docs)

Document Similarity Scoring

  • Query Processing: Convert user query into TF-IDF weighted term vector
  • Document Matching: Compute cosine similarity between query and document vectors
  • Result Ranking: Sort documents by similarity score in descending order

Use Cases in RAG Systems

CSV-Based Retrieval

Lightweight alternative to vector embeddings for small document corpora, as demonstrated in assistant-rh:

  • Preprocessing: Build TF-IDF indices during CSV loading
  • Query Processing: Score document chunks against user queries
  • Fallback Architecture: Secondary retrieval when vector databases unavailable

Hybrid Search Systems

  • Lexical Component: TF-IDF provides keyword-based matching
  • Semantic Component: Combined with dense embeddings for comprehensive retrieval
  • Score Fusion: Weighted combination of lexical and semantic similarity scores

Advantages

Interpretability: Clear mathematical foundation and explainable scoring mechanism.

Computational Efficiency: Fast processing for medium-sized corpora without GPU requirements.

No Training Required: Works out-of-the-box without pre-trained models or embedding generation.

Language Agnostic: Effective across different languages with appropriate tokenization.

Limitations

Semantic Gap: Cannot capture semantic similarity between synonymous terms.

Context Insensitivity: Single-term focus misses multi-word concepts and phrasal meaning.

Vocabulary Mismatch: Query terms must match document terms exactly for retrieval.

Sparse Representations: High-dimensional but sparse vectors can be inefficient for large corpora.

Modern Applications

Document Preprocessing

  • Feature Extraction: Generate term importance weights for downstream ML models
  • Keyword Extraction: Identify most discriminative terms in document collections
  • Content Analysis: Discover topics and themes through high TF-IDF terms

Search Engine Components

  • Query Expansion: Identify related terms based on corpus-wide term statistics
  • Result Explanation: Highlight matching terms with their importance scores
  • Quality Assessment: Evaluate document relevance based on term overlap

Integration with Neural Methods

  • Initial Retrieval: Fast first-stage retrieval before neural reranking
  • Feature Engineering: TF-IDF scores as input features for learning-to-rank models
  • Baseline Comparison: Performance benchmark for evaluating semantic search systems

Implementation Considerations

Tokenization Strategy: Choice of tokenization (word-level, subword, n-grams) significantly impacts performance.

Stopword Filtering: Remove common words that provide little discriminative value.

Normalization Schemes: Different TF and IDF calculation variants (log normalization, sublinear scaling).

Memory Management: Efficient sparse matrix representations for large vocabulary sizes.

See also