~/wiki

LLM Wiki Pattern

Mis à jour le 2025-01-04Confiance : high
llm-wiki-patternpersistent-learningknowledge-managementcompounding-artifactsrag-alternativewiki-maintenanceandrej-karpathymemex-visionincremental-knowledgeimplementation-architectureobsidian-integrationthree-layer-architectureqmd-searchmarp-integrationdataview-pluginoriginal-specification

A paradigm for building personal knowledge bases where LLMs incrementally build and maintain persistent wikis rather than retrieving from raw documents at query time. Developed by andrej-karpathy as an alternative to traditional RAG systems that rediscover knowledge from scratch on every interaction.

Core Innovation

Persistent vs Ephemeral Knowledge: Most people's experience with LLMs and documents follows the RAG pattern - upload files, retrieve relevant chunks at query time, generate answers. This works but requires rediscovering knowledge from scratch on every question. The LLM Wiki Pattern instead creates persistent, compounding artifacts where knowledge is compiled once and kept current.

Key Difference: The wiki sits between you and raw sources. When adding new sources, the LLM doesn't just index for later retrieval - it reads, extracts key information, and integrates into existing wiki structure. Updates entity pages, revises summaries, flags contradictions, strengthens synthesis. Cross-references already exist. Contradictions already flagged. Synthesis already reflects everything read.

Three-Layer Architecture

  1. Raw Sources: Immutable curated collection (articles, papers, images, data files). LLM reads but never modifies. Source of truth.

  2. The Wiki: Directory of LLM-generated markdown files. Summaries, entity pages, concept pages, comparisons, synthesis. LLM owns this layer entirely - creates, updates, maintains cross-references, ensures consistency.

  3. The Schema: Configuration document (CLAUDE.md, AGENTS.md) defining wiki structure, conventions, workflows. Co-evolved between human and LLM over time as patterns emerge.

Core Operations

Ingest Workflow

Drop new source into raw collection. LLM reads source, discusses takeaways, writes summary page, updates index, updates relevant entity/concept pages across wiki, appends log entry. Single source might touch 10-15 wiki pages. Can be done one-at-a-time with supervision or batch-processed.

Query Workflow

Ask questions against wiki. LLM searches relevant pages, reads them, synthesizes answer with citations. Answers can take multiple forms - markdown pages, comparison tables, slide decks (marp-integration), charts, canvas. Critical insight: Good answers get filed back as new wiki pages. Explorations compound in knowledge base like ingested sources.

Lint Workflow

Periodic health-checking. Look for contradictions between pages, stale claims superseded by newer sources, orphan pages with no inbound links, important concepts mentioned but lacking pages, missing cross-references, data gaps. LLM suggests new questions to investigate and sources to find.

Index and Logging

  • index.md: Content-oriented catalog of all wiki pages with links, summaries, metadata. Organized by category. Updated on every ingest. LLM reads index first to find relevant pages for queries. Works well at moderate scale (~100 sources, hundreds of pages).

  • log.md: Chronological append-only record of ingests, queries, lint passes. Parseable with consistent prefixes (## [2026-04-02] ingest | Article Title). Provides timeline of wiki evolution.

Optional Tooling

qmd-search engine recommended for scaling beyond index-based navigation. Local search for markdown with hybrid BM25/vector search and LLM re-ranking. Available as CLI tool and MCP server.

Implementation Recommendations

Obsidian Integration

Recommended interface: obsidian-integration where human browses in Obsidian while LLM makes live edits to markdown files. "Obsidian is the IDE; the LLM is the programmer; the wiki is the codebase."

Useful Obsidian Features:

  • Web Clipper: Browser extension converting articles to markdown
  • Image handling: Download attachments locally, bind to hotkey (Ctrl+Shift+D)
  • Graph view: Visualize wiki connections, identify hubs and orphans
  • marp-integration: Generate slide decks from wiki content
  • dataview-plugin: Query page frontmatter for dynamic tables/lists
  • Git integration: Version history, branching, collaboration

Use Cases

Personal: Goals, health, psychology tracking. File journal entries, articles, podcast notes into structured self-picture over time.

Research: Deep topic exploration over weeks/months. Build comprehensive wiki with evolving thesis.

Reading Companion: File chapters as you go. Build pages for characters, themes, plot threads. Like fan wikis (Tolkien Gateway) but personal with LLM maintenance.

Business/Team: Internal wiki fed by Slack threads, meeting transcripts, project docs, customer calls. Humans review updates.

Other: Competitive analysis, due diligence, trip planning, course notes, hobby deep-dives.

Historical Foundation

Explicitly references vannevar-bush's memex-vision (1945) as spiritual predecessor. Bush envisioned personal, curated knowledge store with associative-trails between documents. His vision was private, actively curated, with connections between documents as valuable as documents themselves.

Key Insight: Bush couldn't solve the maintenance problem. LLMs handle that through bookkeeping-automation. Humans abandon wikis because maintenance burden grows faster than value. LLMs don't get bored, don't forget cross-references, can touch 15 files in one pass. Maintenance cost approaches zero.

Division of Labor

Human Role: Curate sources, direct analysis, ask good questions, think about meaning.

LLM Role: Summarizing, cross-referencing, filing, bookkeeping that makes knowledge base useful over time.

Design Philosophy

Intentionally abstract specification describing the pattern, not specific implementation. Directory structure, schema conventions, page formats, tooling all depend on domain, preferences, LLM choice. Everything modular and optional - pick what's useful. Designed to be shared with LLM agents for collaborative instantiation.

See also