~/wiki

TF-IDF Retrieval

Mis à jour le 2026-06-11Confiance : medium
tf-idfinformation-retrievaltext-searchvectorizationdocument-rankinglexical-search

Term Frequency-Inverse Document Frequency (TF-IDF) is a classical information retrieval technique that scores documents based on term importance within individual documents relative to their rarity across the entire corpus.

Core Mechanism

TF-IDF Calculation:

  • Term Frequency (TF): How often a term appears in a document, usually normalized by document length
  • Inverse Document Frequency (IDF): Logarithmic measure of how rare a term is across the corpus
  • Combined Score: TF × IDF gives higher weights to terms that are frequent in specific documents but rare overall

Advantages:

  • Fast computation and retrieval
  • No training required - works with raw text
  • Interpretable scoring based on term statistics
  • Effective baseline for lexical matching
  • Lightweight memory footprint

Implementation Patterns

CSV-Based Implementation:

  • Tokenize document chunks during initialization
  • Build term-document matrices with TF-IDF weights
  • Cache computed vectors for fast query-time retrieval
  • Return structured results with source metadata

Use Cases:

  • Lightweight retrieval systems without neural infrastructure
  • Baseline comparison for semantic search systems
  • Domain-specific search where exact term matching is crucial
  • Resource-constrained environments requiring fast responses

Limitations

Semantic Blindness:

  • Cannot understand synonyms, context, or meaning
  • Relies purely on exact term matching
  • Struggles with paraphrased queries or concept-level search
  • No understanding of semantic relationships between terms

Modern Alternatives:

  • Dense vector embeddings for semantic understanding
  • Hybrid systems combining TF-IDF with neural search
  • BM25 as an improved probabilistic ranking function

See also