~/wiki

Retrieval Augmented Generation (RAG)

Confiance : high
ragretrievalgenerationknowledge-retrievalvector-searchembedding-modelscontext-augmentationllm-applicationsdocument-qaknowledge-management

A technique that enhances language model responses by retrieving relevant information from external knowledge bases at query time. The retrieved context is then used to ground the model's generation, reducing hallucinations and enabling access to information beyond the training data.

Core Process

  1. Query Processing: User question is converted to embeddings
  2. Retrieval: Similar document chunks are found using vector search
  3. Context Formation: Retrieved chunks are assembled as context
  4. Generation: LLM generates response using retrieved context
  5. Response: Final answer combines retrieved facts with model knowledge

Architecture Components

Vector Store

  • Document embeddings stored in vector database (chroma, pinecone, weaviate)
  • Enables semantic similarity search
  • Supports hybrid search combining dense and sparse retrieval

Embedding Models

  • Convert text to high-dimensional vectors
  • Popular options: OpenAI Ada-002, sentence-transformers, Cohere
  • Critical for retrieval quality

Chunking Strategy

  • Documents split into manageable segments
  • Balance between context preservation and retrieval precision
  • Common approaches: fixed-size, semantic, hierarchical

Advanced Patterns

Multi-Query RAG

Generate multiple query variants to improve retrieval coverage:

def multi_query_rag(question):
    variants = llm.generate_query_variants(question)
    all_chunks = []
    for variant in variants:
        chunks = vector_store.search(variant)
        all_chunks.extend(chunks)
    return llm.generate(question, dedupe(all_chunks))

rag-fusion

Combines multiple retrieval strategies and reranks results for improved relevance.

conversational-rag

Maintains conversation history to enable multi-turn interactions with context awareness.

agentic-rag

Uses AI agents to orchestrate complex retrieval workflows, including tool use and multi-step reasoning.

Limitations

Knowledge Rediscovery Problem

RAG systems rediscover knowledge from scratch on every query. As andrej-karpathy notes in the llm-wiki-pattern, this prevents knowledge accumulation - subtle questions requiring synthesis across multiple documents must repeatedly piece together fragments without building persistent understanding.

Context Window Constraints

  • Limited by model's maximum context length
  • Must balance breadth vs. depth of retrieved information
  • May miss relevant information if not in top-k results

Retrieval Quality Issues

  • Embedding similarity doesn't always match semantic relevance
  • Struggles with concepts spanning multiple documents
  • May retrieve contradictory information without resolution

Alternative Approaches

llm-wiki-pattern

Instead of retrieving from raw documents, LLMs maintain persistent wikis that incrementally integrate knowledge over time. This creates compounding-artifacts where cross-references exist permanently and contradictions are pre-resolved.

knowledge-graphs

Structured representation of entities and relationships can provide more precise retrieval paths than vector similarity.

Use Cases

  • Document Q&A: Customer support, internal documentation
  • Research Assistance: Scientific literature review, fact-checking
  • Content Generation: Blog writing with factual grounding
  • Educational Tools: Personalized tutoring with curriculum materials

Implementation Considerations

Evaluation Metrics

  • Retrieval Accuracy: Precision/recall of relevant documents
  • Answer Quality: Factual correctness, completeness, relevance
  • Latency: End-to-end response time
  • Cost: Embedding generation and storage costs

Production Challenges

  • Keeping knowledge base current with new information
  • Handling contradictory or outdated information
  • Scaling retrieval performance with growing document corpus
  • Managing embedding model updates and reindexing

See also