~/wiki

Concepts — vue longue

retour à la liste

Toutes les pages concaténées sur un seul document, pour un Ctrl-F direct.

Bookkeeping Automation

page dédiée →

The automation of tedious knowledge management tasks that typically cause humans to abandon personal wikis. Core insight behind andrej-karpathy's llm-wiki-pattern - LLMs excel at the systematic maintenance work that humans find burdensome but that's essential for knowledge base health.

The Abandonment Problem

Human Pattern: Start enthusiastic personal wiki → maintenance burden grows → abandon due to tedium Root Cause: "The maintenance burden grows faster than the value"

Tedious Tasks That Cause Abandonment

  • Updating cross-references when adding new information
  • Keeping summaries current as understanding evolves
  • Noting when new data contradicts old claims
  • Maintaining consistency across dozens of pages
  • Filing and organizing information systematically
  • Creating and maintaining category structures

LLM Advantages for Bookkeeping

No Boredom: LLMs don't get tired of repetitive maintenance tasks Perfect Memory: Never forget to update a cross-reference or flag a contradiction
Batch Processing: Can touch 15 files in one pass during integration Systematic Approach: Consistently apply rules and conventions across entire knowledge base Pattern Recognition: Identify maintenance needs humans might miss

Automated Bookkeeping Tasks

During ingest-workflow

  • Update entity pages when new sources mention existing entities
  • Create new entity pages for newly mentioned people/organizations/tools
  • Revise concept summaries with new information
  • Add cross-references between related topics
  • Flag contradictions with existing content
  • Update index and categorization

During query-workflow

  • File valuable answers as new wiki pages
  • Update existing pages with insights from queries
  • Create new cross-references discovered through questioning

During lint-workflow

  • Find and fix broken internal links
  • Identify orphan pages needing better integration
  • Suggest missing cross-references
  • Flag outdated claims for human review

Cost-Benefit Transformation

Traditional Wiki:

  • High human maintenance cost
  • Maintenance burden compounds with scale
  • Value growth plateaus while cost accelerates

LLM Wiki Pattern:

  • Near-zero maintenance cost (automated)
  • Value compounds while cost stays flat
  • Sustainable knowledge accumulation

Human-LLM Role Division

Human Responsibilities:

  • Curate sources (what to include)
  • Direct analysis (what to emphasize)
  • Ask good questions (what to explore)
  • Think about meaning (what it all means)

LLM Responsibilities:

  • All bookkeeping and maintenance
  • Cross-referencing and consistency
  • Systematic organization and filing
  • Routine updates and corrections

Implementation Strategy

The key is documenting bookkeeping requirements in the schema layer so LLMs can execute them consistently. This includes:

  • Cross-reference conventions
  • Page update triggers
  • Consistency checking rules
  • Integration workflows
  • Quality standards

See also

Compounding Artifacts

page dédiée →

Knowledge artifacts that grow in value and sophistication over time through incremental addition and integration of new information, rather than through simple accumulation. The key principle underlying andrej-karpathy's llm-wiki-pattern and modern approaches to persistent-learning.

Core Concept

Value Multiplication: Each new source doesn't just add isolated information - it enhances the entire knowledge base by creating cross-references, resolving contradictions, strengthening synthesis, and revealing new connections between existing concepts.

Persistent Structure: Unlike query-time retrieval systems, compounding artifacts maintain their organized structure between interactions. The cross-references are already there, contradictions already flagged, synthesis already complete and continuously updated.

Semantic Integration: Information from multiple sources synthesizes into coherent understanding that becomes greater than the sum of individual sources through systematic integration rather than simple accumulation.

Compounding Mechanisms

Cross-Reference Network

Each new source creates links to existing entities and concepts, building a knowledge graph where value emerges from connections as much as content.

Contradiction Resolution

New information challenges existing claims, forcing refinement of understanding and highlighting areas requiring deeper investigation.

Synthesis Evolution

Overarching themes and analyses evolve continuously as new perspectives integrate with existing frameworks.

Context Enrichment

Entity profiles and concept definitions deepen with each mention across sources, building comprehensive understanding over time.

Implementation in LLM Wikis

Incremental Integration: Single sources typically update 10-15 existing wiki pages while creating new concept/entity pages, maximizing integration density.

Automated Maintenance: LLMs handle the tedious cross-referencing and consistency work that causes humans to abandon knowledge bases, enabling continuous compounding.

Query Amplification: Valuable query answers get filed back as new wiki pages, turning exploration into knowledge base expansion.

Contrast with Document Storage

Traditional Approach: Each document exists in isolation, requiring reassembly of knowledge for each new question or analysis.

Compounding Approach: Knowledge integrates continuously, with each addition making the entire base more valuable through enhanced connections and refined understanding.

Examples

Personal Research Wiki: Papers and articles integrate into evolving concept pages with cross-referenced relationships and emerging thesis development.

Book Companion Wiki: Characters, themes, and plot threads build incrementally while reading, creating rich companion resource comparable to community-built fan wikis.

Business Intelligence Wiki: Customer calls, meeting transcripts, and market reports synthesize into continuously updated competitive analysis and strategic insight.

Historical Foundation

Builds on vannevar-bush's memex-vision where the connections between documents become as valuable as the documents themselves. Modern LLM automation solves Bush's unsolved maintenance problem.

See also

Human-LLM Division of Labor

page dédiée →

The strategic allocation of responsibilities between humans and LLMs in llm-wiki-pattern systems, optimizing each party's strengths while solving the traditional wiki abandonment problem. Core insight: humans excel at curation and synthesis; LLMs excel at maintenance and bookkeeping.

Task Allocation

Human Responsibilities

  • Source curation: Selecting valuable documents to ingest
  • Direction setting: Guiding analysis emphasis and priorities
  • Question asking: Driving exploration through strategic queries
  • Meaning synthesis: Understanding implications and significance
  • Quality oversight: Reviewing summaries and checking updates
  • Schema evolution: Adapting system configuration based on needs

LLM Responsibilities

  • Content maintenance: Updating cross-references, keeping summaries current
  • Consistency management: Noting contradictions, maintaining coherence
  • Bookkeeping automation: Filing, indexing, logging operations
  • Cross-referencing: Building and maintaining wikilinks networks
  • Structure creation: Generating new pages, organizing content
  • Workflow execution: Following schema-defined operational procedures

Solving the Abandonment Problem

Traditional Wiki Failure Pattern

  • Initial enthusiasm: Humans start with high motivation
  • Growing burden: Maintenance tasks accumulate faster than value
  • Cognitive overhead: Cross-referencing becomes mentally taxing
  • Inevitable abandonment: Effort required exceeds perceived benefit

LLM Solution

  • Zero maintenance fatigue: LLMs don't experience tedium or boredom
  • Consistent execution: Never forget to update cross-references
  • Parallel processing: Can touch 15 files in one pass without cognitive load
  • Near-zero cost: Maintenance burden becomes negligible

Cognitive Complementarity

Human Cognitive Strengths

  • Contextual judgment: Understanding significance and relevance
  • Creative synthesis: Making novel connections and insights
  • Domain expertise: Applying specialized knowledge and intuition
  • Strategic thinking: Long-term planning and goal-oriented exploration

LLM Cognitive Strengths

  • Systematic processing: Consistent application of rules and procedures
  • Pattern recognition: Identifying structural relationships across content
  • Parallel attention: Managing multiple interconnected updates simultaneously
  • Infinite patience: Performing repetitive tasks without degradation

Operational Boundaries

Human Decision Points

  • Source selection: What documents deserve ingestion?
  • Emphasis guidance: What aspects need highlighting?
  • Quality gates: Are summaries accurate and useful?
  • Exploration direction: What questions should drive further investigation?

LLM Execution Points

  • Content integration: How to incorporate new information?
  • Link maintenance: Which pages need cross-reference updates?
  • Consistency checks: Where do contradictions need flagging?
  • Structure organization: How to categorize and file content?

Interface Design

Human-LLM Interaction Model

  • LLM agent open on one side of screen
  • Obsidian (or wiki browser) open on other side
  • Real-time collaboration: Human monitors LLM edits live
  • Immediate feedback: Human can guide and correct during operation

Communication Patterns

  • Explicit instructions: Human provides clear direction for emphasis
  • Progress reporting: LLM describes what updates are being made
  • Quality confirmation: Human reviews and approves significant changes
  • Schema discussion: Collaborative evolution of system configuration

Benefits of Clear Division

Efficiency Optimization

  • Leverage strengths: Each party focuses on optimal tasks
  • Minimize waste: Avoid humans doing tedious work, LLMs making judgment calls
  • Sustainable workflow: Maintenance burden doesn't grow with scale

Quality Assurance

  • Human oversight: Strategic decisions remain under human control
  • LLM consistency: Mechanical tasks executed reliably
  • Complementary validation: Different cognitive approaches catch different errors

Implementation Considerations

Trust Building

  • Gradual automation: Start with supervised workflows, increase autonomy
  • Transparency: LLM reports all changes and reasoning
  • Reversibility: Git versioning enables rollback of problematic updates

Workflow Evolution

  • Usage-driven refinement: Division of labor adapts based on experience
  • Domain customization: Different fields may require different task allocations
  • Tool integration: Technical capabilities influence responsibility boundaries

See also

Incremental Knowledge Building

page dédiée →

The process of systematically adding new information to existing knowledge bases through integration rather than simple accumulation. Core methodology underlying compounding-artifacts and the llm-wiki-pattern. Contrasts sharply with traditional document storage approaches where each source exists in isolation.

Core Process

When new information arrives, incremental knowledge building involves:

  1. Reading and extraction: Parse source for key information, insights, entities
  2. Integration analysis: Identify how new information relates to existing knowledge
  3. Synthesis updating: Revise summaries and overviews to reflect new understanding
  4. Cross-referencing: Create links between new information and relevant existing pages
  5. Contradiction detection: Flag where new data challenges existing claims
  6. Entity maintenance: Update person, organization, concept pages with new details

Implementation in LLM Wikis

In andrej-karpathy's framework, incremental knowledge building happens during the ingest operation:

  • LLM reads new source document
  • Extracts key information and discusses takeaways
  • Writes summary page for the source
  • Updates index with new entry
  • Updates relevant entity and concept pages across wiki
  • Creates new pages for previously uncovered topics
  • Maintains cross-references and flags contradictions

Single source might touch 10-15 wiki pages during integration.

Contrast with Document Accumulation

Traditional approach (Document Libraries):

  • Add new document to collection
  • Document exists in isolation
  • No integration with existing knowledge
  • Connections must be discovered anew each time
  • Knowledge scattered across disconnected files

Incremental Knowledge Building:

  • New information integrates into existing structure
  • Cross-references automatically maintained
  • Contradictions flagged for resolution
  • Synthesis evolves to reflect growing understanding
  • Knowledge compounds rather than accumulates

Maintenance Challenge

The key insight: humans abandon knowledge bases because maintenance burden grows faster than value. Updating cross-references, keeping summaries current, noting when new data contradicts old claims - this bookkeeping is tedious but essential.

LLMs excel at incremental knowledge building because they:

  • Don't get bored with repetitive maintenance tasks
  • Can update multiple files in single operation
  • Maintain consistency across large collections
  • Don't forget to update cross-references

Quality Control

Effective incremental knowledge building requires:

  • Schema adherence: Consistent formatting and organization
  • Citation tracking: Clear source attribution for all claims
  • Contradiction flagging: Explicit noting of conflicting information
  • Quality assessment: Confidence ratings for different claims
  • Regular linting: Periodic health checks for consistency

Applications

Any domain where knowledge accumulates over time:

  • Research synthesis across multiple papers
  • Personal learning from diverse sources
  • Business intelligence aggregation
  • Competitive analysis updates
  • Course note integration

See also

Karpathy-Inspired Claude Code Guidelines

page dédiée →

A practical implementation of andrej-karpathy's observations about LLM coding pitfalls, packaged as guidelines for improving Claude Code behavior. Created by Forrest Chang as a single CLAUDE.md file that addresses common issues in LLM-generated code.

Origin and Problem Statement

Based on Karpathy's insights about LLM coding problems:

  • Models make wrong assumptions and run with them without checking
  • They don't manage confusion, seek clarifications, or surface inconsistencies
  • They overcomplicate code with bloated abstractions
  • They change/remove code orthogonal to the task without understanding

The Four Core Principles

1. Think Before Coding

Addresses: Wrong assumptions, hidden confusion, missing tradeoffs

Implementation:

  • State assumptions explicitly - if uncertain, ask rather than guess
  • Present multiple interpretations - don't pick silently when ambiguity exists
  • Push back when warranted - if simpler approach exists, say so
  • Stop when confused - name what's unclear and ask for clarification

2. Simplicity First

Addresses: Overcomplication, bloated abstractions

Implementation:

  • No features beyond what was asked
  • No abstractions for single-use code
  • No "flexibility" or "configurability" that wasn't requested
  • No error handling for impossible scenarios
  • Test: Would a senior engineer call this overcomplicated?

3. Surgical Changes

Addresses: Orthogonal edits, touching code you shouldn't

When editing existing code:

  • Don't "improve" adjacent code, comments, or formatting
  • Don't refactor things that aren't broken
  • Match existing style, even if you'd do it differently
  • Test: Every changed line should trace directly to user's request

4. Goal-Driven Execution

Addresses: Lack of verifiable success criteria

Transform imperatives into verifiable goals:

  • "Add validation" → "Write tests for invalid inputs, then make them pass"
  • "Fix the bug" → "Write a test that reproduces it, then make it pass"
  • "Refactor X" → "Ensure tests pass before and after"

Installation Methods

/plugin marketplace add forrestchang/andrej-karpathy-skills
/plugin install andrej-karpathy-skills@karpathy-skills

Per-Project CLAUDE.md

# New project
curl -o CLAUDE.md https://raw.githubusercontent.com/forrestchang/andrej-karpathy-skills/main/CLAUDE.md

# Existing project (append)
curl https://raw.githubusercontent.com/forrestchang/andrej-karpathy-skills/main/CLAUDE.md >> CLAUDE.md

Key Architectural Insight

Leverages Karpathy's observation: "LLMs are exceptionally good at looping until they meet specific goals... Don't tell it what to do, give it success criteria and watch it go."

This shifts from imperative instructions to declarative goals with verification loops - a fundamental change in how to interact with coding LLMs.

Success Indicators

Guidelines are working when you observe:

  • Fewer unnecessary changes in diffs
  • Fewer rewrites due to overcomplication
  • Clarifying questions before implementation
  • Clean, minimal PRs without drive-by refactoring

Trade-offs

These guidelines bias toward caution over speed. For trivial tasks, use judgment - the goal is reducing costly mistakes on non-trivial work, not slowing down simple changes.

Integration

Designed to merge with project-specific instructions. Can be combined with existing CLAUDE.md files or used as foundation for team coding standards.

See also

Knowledge Compilation

page dédiée →

The process of systematically transforming raw source documents into structured, cross-referenced knowledge artifacts that persist between interactions. Core innovation of the llm-wiki-pattern that contrasts sharply with query-time retrieval approaches in traditional RAG systems.

Compilation vs Retrieval

Traditional RAG: Fragments documents into chunks, retrieves relevant pieces at query time, and re-synthesizes knowledge for each interaction. Knowledge is rediscovered from scratch repeatedly.

Knowledge Compilation: Pre-processes sources into integrated wiki structures where cross-references exist, contradictions are flagged, and synthesis reflects accumulated reading. Knowledge is compiled once and maintained continuously.

Compilation Process

Initial Integration

When ingesting new sources, the LLM:

  1. Reads and extracts key information from raw documents
  2. Integrates across existing pages - typically touching 10-15 wiki pages per source
  3. Creates new entity/concept pages for previously uncovered topics
  4. Updates cross-references to maintain knowledge graph connectivity
  5. Notes contradictions where new information challenges existing claims
  6. Strengthens synthesis by incorporating supporting evidence

Incremental Refinement

Each compilation cycle builds on previous work:

  • Concept pages deepen with additional sources and perspectives
  • Entity profiles expand with new activities, relationships, and attributes
  • Cross-references multiply as connections between topics emerge
  • Synthesis evolves reflecting cumulative understanding rather than isolated insights

Architectural Benefits

Persistent Structure: Knowledge exists in organized form between sessions, eliminating need to rebuild understanding from raw sources.

Cumulative Intelligence: Each new source makes the entire knowledge base more valuable through integration rather than simple addition.

Query Efficiency: Questions answered by synthesizing from pre-structured content rather than assembling fragments in real-time.

Maintenance Automation: LLMs handle tedious cross-referencing and consistency work that causes humans to abandon personal wikis.

Implementation Patterns

Three-Layer Architecture: Raw sources remain immutable while compiled wiki layer evolves continuously under LLM management guided by schema configuration.

Batch Integration: Single sources can update multiple concept areas simultaneously, creating natural knowledge clustering and relationship discovery.

Version Evolution: Compiled knowledge improves over time as more sources provide additional perspectives and corrections to initial understanding.

Distinction from Document Storage

Unlike traditional document management systems where each source exists in isolation, knowledge compilation creates semantic integration where information from multiple sources synthesizes into coherent, cross-referenced understanding that compounds in value.

The compiled knowledge base becomes greater than the sum of its sources through systematic integration rather than simple accumulation.

See also

LLM Wiki Pattern

page dédiée →

A paradigm for building personal knowledge bases where LLMs incrementally build and maintain persistent wikis rather than retrieving from raw documents at query time. Developed by andrej-karpathy as an alternative to traditional RAG systems that rediscover knowledge from scratch on every interaction.

Core Innovation

Persistent vs Ephemeral Knowledge: Most people's experience with LLMs and documents follows the RAG pattern - upload files, retrieve relevant chunks at query time, generate answers. This works but requires rediscovering knowledge from scratch on every question. The LLM Wiki Pattern instead creates persistent, compounding artifacts where knowledge is compiled once and kept current.

Key Difference: The wiki sits between you and raw sources. When adding new sources, the LLM doesn't just index for later retrieval - it reads, extracts key information, and integrates into existing wiki structure. Updates entity pages, revises summaries, flags contradictions, strengthens synthesis. Cross-references already exist. Contradictions already flagged. Synthesis already reflects everything read.

Three-Layer Architecture

  1. Raw Sources: Immutable curated collection (articles, papers, images, data files). LLM reads but never modifies. Source of truth.

  2. The Wiki: Directory of LLM-generated markdown files. Summaries, entity pages, concept pages, comparisons, synthesis. LLM owns this layer entirely - creates, updates, maintains cross-references, ensures consistency.

  3. The Schema: Configuration document (CLAUDE.md, AGENTS.md) defining wiki structure, conventions, workflows. Co-evolved between human and LLM over time as patterns emerge.

Core Operations

Ingest Workflow

Drop new source into raw collection. LLM reads source, discusses takeaways, writes summary page, updates index, updates relevant entity/concept pages across wiki, appends log entry. Single source might touch 10-15 wiki pages. Can be done one-at-a-time with supervision or batch-processed.

Query Workflow

Ask questions against wiki. LLM searches relevant pages, reads them, synthesizes answer with citations. Answers can take multiple forms - markdown pages, comparison tables, slide decks (marp-integration), charts, canvas. Critical insight: Good answers get filed back as new wiki pages. Explorations compound in knowledge base like ingested sources.

Lint Workflow

Periodic health-checking. Look for contradictions between pages, stale claims superseded by newer sources, orphan pages with no inbound links, important concepts mentioned but lacking pages, missing cross-references, data gaps. LLM suggests new questions to investigate and sources to find.

Index and Logging

  • index.md: Content-oriented catalog of all wiki pages with links, summaries, metadata. Organized by category. Updated on every ingest. LLM reads index first to find relevant pages for queries. Works well at moderate scale (~100 sources, hundreds of pages).

  • log.md: Chronological append-only record of ingests, queries, lint passes. Parseable with consistent prefixes (## [2026-04-02] ingest | Article Title). Provides timeline of wiki evolution.

Optional Tooling

qmd-search engine recommended for scaling beyond index-based navigation. Local search for markdown with hybrid BM25/vector search and LLM re-ranking. Available as CLI tool and MCP server.

Implementation Recommendations

Obsidian Integration

Recommended interface: obsidian-integration where human browses in Obsidian while LLM makes live edits to markdown files. "Obsidian is the IDE; the LLM is the programmer; the wiki is the codebase."

Useful Obsidian Features:

  • Web Clipper: Browser extension converting articles to markdown
  • Image handling: Download attachments locally, bind to hotkey (Ctrl+Shift+D)
  • Graph view: Visualize wiki connections, identify hubs and orphans
  • marp-integration: Generate slide decks from wiki content
  • dataview-plugin: Query page frontmatter for dynamic tables/lists
  • Git integration: Version history, branching, collaboration

Use Cases

Personal: Goals, health, psychology tracking. File journal entries, articles, podcast notes into structured self-picture over time.

Research: Deep topic exploration over weeks/months. Build comprehensive wiki with evolving thesis.

Reading Companion: File chapters as you go. Build pages for characters, themes, plot threads. Like fan wikis (Tolkien Gateway) but personal with LLM maintenance.

Business/Team: Internal wiki fed by Slack threads, meeting transcripts, project docs, customer calls. Humans review updates.

Other: Competitive analysis, due diligence, trip planning, course notes, hobby deep-dives.

Historical Foundation

Explicitly references vannevar-bush's memex-vision (1945) as spiritual predecessor. Bush envisioned personal, curated knowledge store with associative-trails between documents. His vision was private, actively curated, with connections between documents as valuable as documents themselves.

Key Insight: Bush couldn't solve the maintenance problem. LLMs handle that through bookkeeping-automation. Humans abandon wikis because maintenance burden grows faster than value. LLMs don't get bored, don't forget cross-references, can touch 15 files in one pass. Maintenance cost approaches zero.

Division of Labor

Human Role: Curate sources, direct analysis, ask good questions, think about meaning.

LLM Role: Summarizing, cross-referencing, filing, bookkeeping that makes knowledge base useful over time.

Design Philosophy

Intentionally abstract specification describing the pattern, not specific implementation. Directory structure, schema conventions, page formats, tooling all depend on domain, preferences, LLM choice. Everything modular and optional - pick what's useful. Designed to be shared with LLM agents for collaborative instantiation.

See also

Loop Stacking

page dédiée →

The art and science of designing nested autonomous loops that operate without human intervention, representing a fundamental paradigm shift from manual prompting to systematic leverage amplification. Core thesis: "the entire game of the next century is to be able to stack loops as effectively as possible."

Core Philosophy

Loop stacking emerged from the recognition that human researchers and developers become bottlenecks in AI-assisted workflows. Instead of manually prompting each step, practitioners design autonomous systems that can:

  • Remove themselves as the bottleneck in iterative processes
  • Maximize token throughput without human intervention
  • Scale leverage through systematic orchestration
  • Operate continuously without manual oversight

The Salty Lesson

Parallel to Rich Sutton's "Bitter Lesson" for models, the Salty Lesson for agents states:

Don't fix things yourself, as you have done historically. Instead focus on systems that scale with more agents, like goals and orchestration.

The implication: those who master loop stacking will outcompete those who remain in manual prompting paradigms.

Practical Approaches

Current Implementations

  • Steipete's Approach: "You shouldn't be prompting coding agents anymore. You should be designing loops that prompt your agents."
  • Boris's Method: "I don't prompt Claude anymore. I write loops, the loops do the work."
  • andrej-karpathy's Autoresearch: Complete autonomy in research workflows, removing human researchers from the loop entirely

Loop Hierarchy Design

Modern loop stacking requires understanding when to:

  • Go DOWN a loop: When things go wrong (for reliability)
  • Go UP a loop: As models improve (for leverage)

Applications

Automated Research

  • recursive-si using rapid iteration loops for optimization benchmarks
  • microsoft-arbor implementing persistent hypothesis-tree refinement
  • Continuous hypothesis generation and testing without human oversight

Development Workflows

  • Autonomous code generation and review cycles
  • Self-improving system optimization
  • Automated debugging and refactoring loops

Data Processing

  • macrodata-labs' robotics data pipeline automation
  • goodfire's predictive debugging loops
  • weaviate-engram's memory maintenance cycles

Technical Challenges

Reliability vs Leverage

Early loop implementations require careful balance between:

  • Autonomous operation (leverage maximization)
  • Error handling and human fallbacks (reliability assurance)
  • Loop termination conditions and escape hatches

Orchestration Complexity

Successful loop stacking demands sophisticated:

  • Goal management across nested loops
  • State synchronization between autonomous processes
  • Resource allocation and priority management

Future Implications

Loop stacking represents the foundational paradigm for the next phase of AI development, where the primary competitive advantage shifts from prompt engineering to autonomous system design. Organizations and individuals who master this transition will achieve orders of magnitude improvements in leverage and productivity.

See also

Memex Vision

page dédiée →

vannevar-bush's 1945 vision of a personal knowledge management system that would allow individuals to store, retrieve, and create associative-trails between documents and information. A foundational concept for modern knowledge management and the direct inspiration behind andrej-karpathy's llm-wiki-pattern and other AI-powered information systems.

Historical Context

Introduced in Bush's essay "As We May Think" (1945), the Memex was conceived as a mechanical device that would supplement human memory by allowing rapid consultation of personal collections of documents. Bush envisioned a desk-sized device with screens, keyboards, and mechanical storage systems.

Core Vision

Private curation - Personal knowledge stores actively maintained by individuals rather than shared databases

Associative trails - Connections between documents as valuable as the documents themselves. Users could create and follow paths linking related information across their collection.

Mechanical augmentation - Technology amplifying human cognitive capabilities rather than replacing human judgment

Active maintenance - Knowledge bases requiring continuous organization and cross-referencing to remain valuable

The Maintenance Problem

Bush's vision was prescient but incomplete. He understood the value of connected, curated knowledge but couldn't solve the fundamental challenge: who does the maintenance work?

Creating associative trails, maintaining cross-references, keeping summaries current, noting contradictions - this bookkeeping is essential but tedious. Humans consistently abandon personal knowledge systems because maintenance burden grows faster than perceived value.

Modern Resolution

The llm-wiki-pattern directly addresses Bush's unsolved maintenance problem. LLMs excel at the systematic bookkeeping that humans find burdensome:

  • Updating cross-references across multiple pages
  • Maintaining consistency as information accumulates
  • Noting contradictions between sources
  • Creating and strengthening associative connections

This allows Bush's original vision to be realized: private, actively curated knowledge stores where connections between documents are as valuable as documents themselves.

Influence on Contemporary Systems

The Memex vision influenced:

  • Hypertext systems and the World Wide Web
  • Personal knowledge management tools (Roam, Obsidian, Notion)
  • Information retrieval research
  • Modern AI-powered knowledge systems

However, most implementations lost Bush's emphasis on private curation and became either:

  • Public databases (Wikipedia, web)
  • Simple storage without maintained associations (file systems)
  • Tools requiring manual maintenance (personal wikis that get abandoned)

Key Insights for LLM Wiki Pattern

Connection value - The relationships between pieces of information often more valuable than isolated documents

Curation over scale - Personal, curated collections outperform generic large databases for individual needs

Maintenance as bottleneck - Technical capability less important than solving the maintenance burden problem

Human-machine collaboration - Best results combine human judgment (curation, direction) with machine capability (systematic maintenance)

See also

Obsidian Integration

page dédiée →

The recommended interface pattern for llm-wiki-pattern implementations where humans browse and explore wiki content in obsidian while LLMs act as programmers making live edits to the underlying markdown files. Creates seamless collaboration between human intuition and LLM maintenance.

Integration Model

Collaborative Workflow:

  • LLM agent open on one side, Obsidian open on the other
  • LLM makes edits based on conversation and source ingestion
  • Human browses results in real time - following links, checking graph view, reading updated pages
  • Metaphor: "Obsidian is the IDE; the LLM is the programmer; the wiki is the codebase"

This creates a natural division of labor where humans excel at exploration and pattern recognition while LLMs handle systematic maintenance and content generation.

Essential Obsidian Features

Graph View:

  • Best way to visualize wiki structure and relationships
  • Identifies hubs (highly connected pages) and orphans (isolated pages)
  • Reveals emerging themes and knowledge clusters
  • Essential for understanding wiki evolution over time

Real-time Markdown Editing:

  • See LLM changes appear instantly in Obsidian
  • Navigate between updated pages immediately
  • Visual feedback on cross-reference creation and updates

Wikilink Support:

  • Native support for internal links between pages
  • Automatic backlinking and reference tracking
  • Foundation for the associative trail navigation

Obsidian Web Clipper:

  • Browser extension that converts web articles to markdown
  • Essential for quickly adding sources to raw collection
  • Preserves formatting and structure for LLM processing

Dataview Plugin:

  • Runs queries over page frontmatter (YAML metadata)
  • Generates dynamic tables and lists from wiki content
  • Useful when LLM adds structured metadata (tags, dates, source counts)

Marp Plugin:

  • Markdown-based slide deck format support
  • Enables generating presentations directly from wiki content
  • LLM can create slide decks as query output format

Attachment Management:

  • Set "Attachment folder path" to fixed directory (e.g., raw/assets/)
  • Bind "Download attachments for current file" to hotkey (e.g., Ctrl+Shift+D)
  • Downloads all images locally for LLM analysis and stable references

Image Handling Workflow

Process:

  1. Clip article with Obsidian Web Clipper
  2. Hit download hotkey to store all images locally
  3. LLM reads text first, then examines referenced images separately
  4. Slightly clunky but works well for multimodal content integration

Benefits:

  • Stable image references (URLs don't break)
  • LLM can analyze images directly
  • Complete offline functionality

Version Control Integration

Since the wiki is just a git repo of markdown files:

  • Full version history of all changes
  • Branch different analysis directions
  • Collaborate with others through standard git workflows
  • Track LLM edits over time for quality assessment

Index-Driven Exploration:

  • Start with index.md for wiki overview
  • Follow category-based organization
  • Jump to specific pages via search or links

Graph-Driven Discovery:

  • Use graph view to find connection patterns
  • Identify knowledge gaps (missing connections)
  • Explore topic clusters and emerging themes

Log-Based Timeline:

  • Review log.md for chronological wiki evolution
  • Understand how knowledge accumulated over time
  • Identify periods of high activity or major updates

Real-time Collaboration Benefits

Immediate Feedback:

  • See LLM understanding reflected in page updates
  • Catch misinterpretations or missed connections quickly
  • Guide LLM emphasis through real-time review

Enhanced Exploration:

  • Follow interesting connections as they're created
  • Discover unexpected relationships through graph visualization
  • Build on LLM insights through guided browsing

Quality Assurance:

  • Verify cross-references are accurate and meaningful
  • Ensure page updates maintain consistency
  • Check that new content integrates well with existing knowledge

Scaling Considerations

As wikis grow beyond ~100 pages:

  • Graph view becomes dense but still valuable for cluster identification
  • Search becomes more important than browsing
  • Consider complementary tools like qmd for advanced search capabilities

The Obsidian interface remains effective even at scale due to its flexible navigation options and powerful visualization capabilities.

See also

Persistent Learning

page dédiée →

A paradigm in AI systems where knowledge accumulates and compounds over time rather than being re-derived from scratch for each interaction. Contrasts with stateless approaches where each query starts fresh. Core concept underlying andrej-karpathy's llm-wiki-pattern and compounding-artifacts.

Core Concept

Traditional systems like RAG rediscover knowledge from scratch on every question - finding relevant chunks, piecing together fragments, synthesizing answers from raw sources. Nothing accumulates. Ask a subtle question requiring synthesis across five documents, and the LLM repeats the same discovery process every time.

Persistent learning systems instead build and maintain accumulated knowledge structures that compound over interactions. The synthesis work is done once and then kept current, not repeated. Each new source strengthens the existing structure rather than existing in isolation.

Implementation in LLM Wiki Pattern

The llm-wiki-pattern implements persistent learning through:

Knowledge compilation - Raw sources transformed into structured, cross-referenced wiki pages that persist between sessions

Incremental integration - New information updates existing pages rather than creating isolated summaries

Maintained synthesis - Cross-references, contradictions, and connections continuously maintained as knowledge base evolves

Compounding exploration - Good query answers filed back as wiki pages, making explorations persistent rather than ephemeral

Benefits Over Stateless Systems

  • Accumulated insights: Connections and contradictions already identified
  • Refined understanding: Multiple sources integrated into coherent picture
  • Efficient querying: Pre-compiled knowledge vs. real-time fragment assembly
  • Progressive deepening: Each interaction builds on previous work
  • Preserved discoveries: Valuable insights don't disappear into chat history

Historical Context

Concept relates to vannevar-bush's memex-vision - personal knowledge stores that accumulate and cross-reference over time. Also parallels human learning where new information integrates with existing knowledge structures rather than existing in isolation.

Implementation Challenges

Requires systematic maintenance that humans find tedious but LLMs handle naturally:

  • Updating cross-references across multiple pages
  • Noting contradictions between old and new information
  • Maintaining consistency as knowledge base grows
  • Organizing and categorizing accumulated knowledge

See also

RAG Alternative

page dédiée →

The llm-wiki-pattern represents a fundamental alternative to traditional Retrieval-Augmented Generation (RAG) approaches. Instead of retrieving and synthesizing from raw documents on each query, knowledge is compiled once into structured wikis and maintained persistently.

Traditional RAG Limitations

Repeated Rediscovery: RAG systems upload document collections, retrieve relevant chunks at query time, and generate answers from fragments. The LLM rediscovers knowledge from scratch on every question. No accumulation or learning occurs.

Fragment Synthesis: Complex questions requiring synthesis of multiple sources force the LLM to find and piece together relevant fragments repeatedly. Nothing is built up or persists between queries.

Shallow Integration: Sources exist in isolation. Cross-references, contradictions, and deeper synthesis must be discovered anew for each interaction.

Wiki Pattern Advantages

Knowledge Compilation: Information is processed once during ingestion, integrated into existing understanding, cross-referenced with related content, and maintained as living knowledge base.

Persistent Synthesis: Complex analysis is performed once and stored. Cross-references are pre-computed, contradictions are pre-identified, synthesis reflects everything previously ingested.

Compounding Value: Each new source strengthens the entire knowledge base rather than existing in isolation. Value grows exponentially rather than linearly.

Architectural Comparison

Approach Knowledge Storage Query Processing Maintenance Value Growth
RAG Raw documents + embeddings Retrieve fragments → synthesize None Linear
Wiki Pattern Structured, cross-referenced pages Read relevant pages → reference Automated by LLM Exponential

Implementation Trade-offs

RAG Advantages:

  • Simpler initial setup
  • No maintenance overhead
  • Works well for straightforward Q&A
  • Established tooling ecosystem

Wiki Pattern Advantages:

  • Knowledge compounds over time
  • Complex synthesis performed once
  • Rich cross-referencing and navigation
  • Handles contradictions systematically
  • Scales to deeper analysis

When to Use Each Approach

RAG Appropriate For:

  • Simple document search and retrieval
  • Static document collections
  • Minimal ongoing engagement
  • Straightforward Q&A scenarios

Wiki Pattern Appropriate For:

  • Long-term knowledge building projects
  • Research requiring synthesis across sources
  • Personal knowledge management
  • Complex domain understanding
  • Ongoing learning and exploration

Hybrid Possibilities

Complementary Usage: Wiki pattern for core knowledge base with RAG for supplementary document retrieval. Wiki handles synthesis and cross-referencing while RAG provides access to broader document collections.

Migration Path: Start with RAG for document exploration, migrate valuable synthesis to wiki format for persistent reference and further development.

Technical Requirements

Wiki Pattern Demands:

  • LLM capable of multi-document reasoning
  • File system access for wiki maintenance
  • Schema-driven workflow execution
  • Structured output generation (markdown, YAML)

Integration Complexity: Higher initial setup cost but lower ongoing maintenance burden compared to RAG systems that require continuous re-processing.

See also

RAG Alternatives

page dédiée →

Approaches to knowledge management and information retrieval that move beyond traditional Retrieval-Augmented Generation (RAG) systems. Most notably exemplified by andrej-karpathy's llm-wiki-pattern which treats knowledge as compounding-artifacts rather than static document collections.

Traditional RAG Limitations

Standard RAG Approach:

  • Upload collection of files
  • LLM retrieves relevant chunks at query time
  • Generates answers from fragments
  • Rediscovers knowledge from scratch on every question
  • No accumulation or synthesis between queries

Problems:

  • Subtle questions requiring synthesis across multiple documents must be solved repeatedly
  • No building up of understanding over time
  • Connections between sources not maintained
  • Context limited to what can be retrieved in single query

LLM Wiki Pattern Alternative

Core Difference: Instead of retrieving from raw documents at query time, the LLM incrementally builds and maintains a persistent wiki that sits between user and sources.

Process:

  • New sources integrate into existing wiki structure
  • Updates entity pages, revises summaries, notes contradictions
  • Cross-references already established
  • Synthesis reflects all previous learning
  • Knowledge compounds with each addition

Key Advantages

Persistent Knowledge: Cross-references exist, contradictions flagged, synthesis current. Wiki keeps getting richer with every source and question.

Maintenance Automation: LLMs handle tedious bookkeeping - updating cross-references, keeping summaries current, maintaining consistency. Humans focus on curation and questions.

Compound Learning: Good answers become new wiki pages. Explorations compound in knowledge base just like ingested sources.

Other Alternative Approaches

Knowledge Graphs: Structured representation of entities and relationships, but typically requires significant manual curation or complex extraction pipelines.

Embedding-Based Memory Systems: Vector representations of experiences/documents that can be retrieved by similarity, but lack explicit structure and cross-referencing.

Agent Memory Architectures: Various approaches to giving AI systems persistent memory, though most focus on conversation history rather than structured knowledge building.

Implementation Considerations

Moving beyond RAG requires:

  • Schema Design: Clear workflows and conventions for knowledge maintenance
  • Integration Workflows: Systematic processes for incorporating new information
  • Quality Control: Methods for ensuring accuracy and consistency
  • Navigation Systems: Tools for exploring and searching the knowledge base

Applications

RAG alternatives particularly valuable for:

  • Research Synthesis: Building understanding across many sources over time
  • Personal Learning: Accumulating knowledge in specific domains
  • Business Intelligence: Maintaining current understanding of competitive landscape
  • Domain Expertise: Building deep, interconnected knowledge in specialized areas

See also

RAG vs Wiki Compilation

page dédiée →

Fundamental architectural comparison between traditional Retrieval-Augmented Generation (RAG) systems and andrej-karpathy's llm-wiki-pattern approach to knowledge management. Represents a shift from ephemeral retrieval to persistent knowledge compilation.

Core Philosophical Difference

Traditional RAG: Knowledge is rediscovered from scratch on every query Wiki Compilation: Knowledge is compiled once and maintained persistently

This difference has profound implications for knowledge accumulation, query efficiency, and system sophistication over time.

Operational Comparison

Traditional RAG Workflow

  1. User asks question
  2. System retrieves relevant document chunks
  3. LLM pieces together fragments at query time
  4. Generates answer from scratch
  5. Answer disappears into chat history

Wiki Compilation Workflow

  1. Sources ingested incrementally into persistent wiki
  2. Knowledge integrated with existing understanding
  3. User asks question
  4. System searches pre-compiled wiki pages
  5. LLM synthesizes from maintained knowledge
  6. Good answers filed back as new wiki pages

Architectural Trade-offs

RAG Advantages

  • Simplicity: Straightforward to implement and understand
  • Source Fidelity: Directly references original document chunks
  • Real-time: Can incorporate new documents immediately
  • Stateless: No complex state management required

RAG Limitations

  • Redundant Work: Same analysis repeated for similar queries
  • Fragment Dependency: Quality depends on chunk retrieval accuracy
  • No Accumulation: Insights don't compound over time
  • Context Loss: Subtle connections require re-discovery each time

Wiki Compilation Advantages

  • Knowledge Compounding: Understanding builds and strengthens over time
  • Sophisticated Synthesis: Cross-references and contradictions pre-analyzed
  • Query Efficiency: Answers draw from compiled knowledge, not raw chunks
  • Persistent Learning: System gets smarter with each source and query

Wiki Compilation Limitations

  • Complexity: Requires sophisticated maintenance workflows
  • Delayed Integration: New sources need processing before availability
  • State Management: Must maintain consistency across evolving knowledge base
  • Schema Evolution: System architecture must adapt as understanding grows

Scaling Characteristics

RAG: Performance degrades with document volume due to retrieval complexity Wiki Compilation: Performance improves with volume as knowledge density increases

Use Case Suitability

RAG Optimal For

  • Document search and basic Q&A
  • Large, relatively static document collections
  • Simple factual retrieval tasks
  • Systems requiring immediate source transparency

Wiki Compilation Optimal For

  • Research and analysis over time
  • Complex synthesis across multiple sources
  • Personal knowledge management
  • Domains requiring evolving understanding

Knowledge Quality Evolution

RAG: Knowledge quality remains constant - system doesn't learn Wiki Compilation: Knowledge quality improves through:

  • Cross-source validation and contradiction resolution
  • Incremental refinement of understanding
  • Connection discovery between previously isolated concepts
  • Synthesis sophistication increases with experience

Implementation Complexity

RAG: Vector databases, embedding models, retrieval tuning Wiki Compilation: Schema design, maintenance workflows, consistency checking

Future Hybrid Approaches

Emerging systems may combine both approaches:

  • Wiki compilation for core knowledge domains
  • RAG fallback for novel or peripheral queries
  • Automatic promotion from RAG results to wiki compilation

See also

RAG vs Wiki Pattern

page dédiée →

Fundamental comparison between traditional Retrieval-Augmented Generation (RAG) systems and andrej-karpathy's llm-wiki-pattern, highlighting different approaches to knowledge management and query answering.

Traditional RAG Approach

Process Flow

  1. Upload collection of files
  2. Chunk and embed documents
  3. At query time: retrieve relevant chunks
  4. Generate answer from retrieved fragments
  5. Repeat process for each new query

Characteristics

  • Stateless: Each query starts fresh
  • Rediscovery: Knowledge reconstructed every time
  • Fragment-based: Works with document chunks, not integrated knowledge
  • No accumulation: Understanding doesn't build over time
  • Limited synthesis: Difficult to connect insights across multiple documents

Examples

  • NotebookLM
  • ChatGPT file uploads
  • Most commercial RAG systems
  • Document Q&A tools

LLM Wiki Pattern Approach

Process Flow

  1. Ingest sources into persistent wiki structure
  2. LLM builds and maintains cross-referenced knowledge base
  3. At query time: synthesize from existing wiki pages
  4. File valuable answers back as new wiki content
  5. Knowledge compounds with each interaction

Characteristics

  • Stateful: Persistent knowledge accumulation
  • Pre-synthesized: Understanding built incrementally over time
  • Integration-based: Works with structured, cross-referenced content
  • Compounding: Each addition makes the whole more valuable
  • Rich synthesis: Connections already established across sources

Key Differences

Aspect Traditional RAG Wiki Pattern
Knowledge State Stateless retrieval Persistent accumulation
Processing Time Query-time discovery Ingest-time integration
Cross-reference Ad-hoc during query Pre-established and maintained
Contradictions Discovered per query Flagged and tracked systematically
Synthesis Quality Limited by retrieval Rich, pre-built connections
Maintenance None (static chunks) Automated wiki maintenance

When to Use Each

RAG Works Well For

  • Simple document Q&A
  • One-off queries against large document sets
  • When you don't need knowledge to compound
  • Rapid deployment without setup overhead
  • Documents that rarely need cross-referencing

Wiki Pattern Works Well For

  • Long-term knowledge building
  • Complex synthesis across multiple sources
  • Research that builds over time
  • When contradictions and evolution matter
  • Domains where connections between concepts are valuable

Hybrid Approaches

Some systems might combine both patterns:

  • Wiki pattern for core, frequently-accessed knowledge
  • RAG for supplementary document collections
  • Different layers for different types of content

Performance Implications

RAG

  • Pros: Simple setup, works immediately, scales to large document sets
  • Cons: Repeated processing overhead, limited synthesis depth, no knowledge accumulation

Wiki Pattern

  • Pros: Rich synthesis, compounding value, deep cross-references, maintained consistency
  • Cons: Higher setup cost, requires ongoing LLM maintenance, more complex architecture

See also

Role Division

page dédiée →

The systematic allocation of responsibilities between humans and LLMs in the llm-wiki-pattern, optimizing for each party's comparative advantages. Solves the knowledge management maintenance problem through cognitive specialization rather than trying to make humans better at tedious tasks.

Human Responsibilities

High-Level Cognitive Tasks

  • Source Curation: Decide what information is worth including
  • Analysis Direction: Guide what aspects to emphasize or explore
  • Question Formation: Ask the right questions to drive useful synthesis
  • Meaning Making: Think about what the accumulated knowledge means
  • Strategic Decisions: Determine knowledge base scope and priorities

Why Humans Excel Here

  • Contextual Judgment: Understanding relevance and quality in complex domains
  • Creative Connections: Seeing non-obvious relationships and implications
  • Value Alignment: Ensuring knowledge base serves intended purposes
  • Intuitive Filtering: Recognizing what's worth investigating further

LLM Responsibilities

Systematic Maintenance Tasks

  • Cross-Referencing: Creating and maintaining links between related concepts
  • Consistency Management: Ensuring information coherence across pages
  • Contradiction Detection: Flagging when new sources conflict with existing content
  • Integration Work: Updating multiple pages when new sources arrive
  • Organizational Maintenance: Filing, categorizing, and indexing systematically

Why LLMs Excel Here

  • No Boredom: Don't get tired of repetitive maintenance work
  • Perfect Consistency: Apply rules uniformly across entire knowledge base
  • Batch Processing: Can update many pages simultaneously
  • Pattern Recognition: Identify maintenance needs systematically
  • Attention to Detail: Don't forget to update cross-references or categorization

The Division Principle

Humans do what requires judgment and creativity
LLMs do everything else

This leverages comparative advantages rather than trying to overcome weaknesses. Humans don't need to become better at tedious maintenance; LLMs handle that entirely.

Historical Context

Traditional knowledge management fails because it asks humans to do both:

  • High-value work: Thinking, analyzing, connecting
  • Maintenance work: Filing, cross-referencing, updating

The maintenance burden eventually overwhelms the value creation, leading to abandonment. Role division solves this by removing maintenance burden from humans entirely.

Implementation in Practice

Human Workflow

  1. Find interesting sources
  2. Add to raw collection
  3. Guide LLM on what to emphasize during ingest
  4. Ask questions that drive useful synthesis
  5. Review and approve major changes
  6. Think about what accumulated knowledge means

LLM Workflow

  1. Process sources systematically
  2. Update all affected wiki pages
  3. Maintain cross-references and

Schema Coevolution

page dédiée →

The collaborative development process between human and LLM for evolving the configuration layer of llm-wiki-pattern systems. The schema document (CLAUDE.md, AGENTS.md, etc.) grows and adapts based on actual usage patterns, domain needs, and discovered workflows.

Core Concept

Unlike static system configurations, the schema in an LLM wiki system evolves with use. As humans and LLMs work together, they discover:

  • Effective workflows for the specific domain
  • Useful page formats and structures
  • Domain-specific conventions and tags
  • Optimal ingest and query patterns

These discoveries get documented in the schema, making the LLM a more effective collaborator over time.

Evolution Process

Initial Schema

  • Basic structure and conventions
  • Generic workflows from the pattern
  • Minimal domain-specific customization

Usage-Driven Refinement

  • Workflow optimization: Discovering efficient ingest patterns
  • Convention standardization: Establishing consistent tagging and formatting
  • Domain adaptation: Adding field-specific page types and structures
  • Tool integration: Incorporating discovered utilities and scripts

Continuous Improvement

  • Regular schema updates based on what works
  • Documentation of effective practices
  • Removal of unused conventions
  • Addition of new capabilities

Human-LLM Collaboration

Human Contributions

  • Domain expertise: Understanding field-specific needs
  • Workflow preferences: Preferred interaction patterns
  • Quality standards: Defining acceptable outputs
  • Strategic direction: Long-term knowledge goals

LLM Contributions

  • Pattern recognition: Identifying recurring structures
  • Consistency maintenance: Ensuring schema adherence
  • Workflow execution: Following documented procedures
  • Improvement suggestions: Proposing optimizations

Schema Components That Evolve

Page Types

  • Standard templates for different content types
  • Domain-specific entity categories
  • Specialized analysis formats (comparisons, timelines, etc.)

Tagging Systems

  • Hierarchical tag structures for the domain
  • Consistent naming conventions
  • Cross-reference patterns

Workflows

  • Detailed ingest procedures
  • Query and synthesis patterns
  • Maintenance and lint operations
  • Quality control checkpoints

Tool Integration

  • CLI utilities and their usage patterns
  • Search and navigation tools
  • Export and visualization formats
  • Integration with external systems

Benefits

Adaptive Systems

Schema coevolution enables knowledge systems that improve with use rather than becoming stale or rigid.

Domain Optimization

Over time, the system becomes specifically tuned to the user's field and preferences rather than remaining generic.

Reduced Friction

Well-evolved schemas reduce cognitive overhead by codifying effective practices and eliminating decision fatigue.

Knowledge Transfer

The schema serves as documentation of effective knowledge management practices that can be shared or adapted.

Example Evolution

Initial: "Create entity pages for people mentioned"
Evolved: "Create entity pages with standardized sections: Background, Key Ideas, Publications, Influence, Cross-References. Tag with domain (ai-researcher, entrepreneur, academic) and confidence level."

Challenges

Over-Specification

Schemas can become too rigid, constraining useful variation and experimentation.

Version Management

Managing schema changes while maintaining consistency across existing wiki content.

Complexity Growth

Balancing comprehensive documentation with usability for both human and LLM.

Implementation Patterns

Versioned Schemas

Track schema evolution with version control to understand what changes improve effectiveness.

Modular Structure

Organize schema into sections (workflows, conventions, formats) that can evolve independently.

Example-Driven Documentation

Include concrete examples in schema documentation to clarify abstract conventions.

See also

Schema-Driven Workflows

page dédiée →

Configuration approach where LLM behavior is governed by explicit schema documents that define structure, conventions, and operational workflows. Central to andrej-karpathy's llm-wiki-pattern for creating disciplined wiki maintainers rather than generic chatbots.

Core Concept

Configuration as Documentation: Schema document (e.g., CLAUDE.md, AGENTS.md, SCHEMA.md) serves as comprehensive instruction manual for LLM behavior. Defines how wiki is structured, what conventions to follow, and what workflows to execute for different operations.

Disciplined Automation: Transforms general-purpose LLM into specialized wiki maintainer through explicit operational instructions. Schema constrains and focuses LLM behavior for consistent, reliable knowledge management.

Schema Components

Structural Definitions:

  • Directory organization and file naming conventions
  • Page format templates and YAML frontmatter requirements
  • Cross-referencing patterns and linking conventions
  • Category taxonomies and tagging standards

Operational Workflows:

  • Ingest procedures (read → summarize → integrate → cross-reference)
  • Query handling (search index → read pages → synthesize → optionally file)
  • Maintenance routines (lint for contradictions, orphans, missing links)
  • Logging formats and indexing standards

Quality Standards:

  • Confidence levels and sourcing requirements
  • Cross-reference density expectations
  • Contradiction handling procedures
  • Update propagation rules

Workflow Examples

Ingest Workflow (from schema):

  1. Read source completely and extract key information
  2. Create summary page in wiki/sources/ with metadata
  3. Update 10-15 relevant existing pages across wiki
  4. Maintain cross-references and flag contradictions
  5. Update index.md and append to log.md
  6. Generate flashcards if applicable

Lint Workflow (from schema):

  1. Scan for contradictions between pages
  2. Identify orphan pages with no inbound links
  3. Find mentioned-but-missing concepts needing pages
  4. Check for stale content superseded by newer sources
  5. Suggest new questions and sources to investigate

Co-Evolution Pattern

Human-LLM Collaboration: Schema evolves through usage as human and LLM discover what works for specific domain. Initial framework adapts to actual needs and preferences over time.

Domain Adaptation: General pattern customizes to specific use cases - research wikis have different needs than business intelligence or personal knowledge management.

Workflow Refinement: Operational procedures improve based on experience with what produces highest-quality knowledge accumulation.

Technical Implementation

Schema as Context: LLM reads schema document at start of every session to understand current operational parameters and workflow requirements.

Parseable Formats: Log entries and metadata use consistent formats enabling programmatic analysis and tooling development.

Version Control: Schema documents tracked in git alongside wiki content, providing evolution history and rollback capability.

Benefits

Consistency: Ensures uniform quality and structure across all wiki operations regardless of session or time period.

Reliability: Reduces variability in LLM behavior through explicit instruction rather than relying on implicit understanding.

Maintainability: Schema serves as documentation for wiki structure and operational procedures, enabling debugging and improvement.

Scalability: Well-defined workflows allow wiki to grow systematically without degrading quality or consistency.

Configuration Examples

Directory Structure Rules:

wiki/concepts/     # Technical concepts and methods
wiki/entities/     # People, organizations, tools  
wiki/sources/      # Individual source summaries
raw/articles/      # Immutable source documents

Page Format Standards:

---
title: Page Title
category: entities|concepts|projects|sources
created: YYYY-MM-DD  
updated: YYYY-MM-DD
tags: [lowercase-hyphenated-tags]
sources: [list-of-raw-source-paths]
confidence: high|medium|low
---

See also

Three Layer Architecture

page dédiée →

The foundational architectural pattern underlying andrej-karpathy's llm-wiki-pattern that separates concerns across three distinct layers: raw sources, wiki content, and schema configuration. This separation enables clear ownership models, version control, and collaborative development between humans and LLMs.

Layer Definitions

Raw Sources Layer:

  • Curated collection of source documents (articles, papers, images, data files)
  • Immutable - LLM reads from them but never modifies them
  • Serves as the authoritative source of truth
  • Examples: raw/articles/, raw/papers/, raw/transcripts/

Wiki Layer:

  • Directory of LLM-generated markdown files
  • Contains summaries, entity pages, concept pages, comparisons, syntheses
  • LLM owns this layer entirely - creates, updates, maintains cross-references
  • Clear division: humans read it, LLMs write it
  • Examples: wiki/concepts/, wiki/entities/, wiki/sources/

Schema Layer:

  • Configuration document (CLAUDE.md, AGENTS.md, SCHEMA.md, etc.)
  • Defines wiki structure, conventions, and operational workflows
  • Co-evolved between human and LLM based on actual usage patterns
  • Makes the LLM a disciplined wiki maintainer rather than generic chatbot
  • Contains ingest workflows, query patterns, maintenance procedures

Ownership Model

The architecture establishes clear ownership and modification rights:

Human Responsibilities:

  • Curate and add sources to raw layer
  • Modify schema based on evolving needs
  • Browse and explore wiki content
  • Direct analysis and ask questions

LLM Responsibilities:

  • Generate and maintain all wiki content
  • Follow schema conventions and workflows
  • Update multiple pages per source integration
  • Maintain cross-references and consistency

Benefits of Separation

Version Control: Each layer can be versioned independently, allowing rollback of wiki changes without affecting sources or schema evolution.

Clear Interfaces: Well-defined boundaries prevent scope creep and maintain system integrity. Sources remain pristine, wiki stays automatically maintained.

Collaborative Development: Humans and LLMs can work simultaneously without conflicts - humans curate sources and refine schema while LLMs maintain wiki content.

Debugging and Maintenance: Issues can be isolated to specific layers, making troubleshooting more systematic.

Implementation Patterns

Directory Structure:

knowledge-base/
├── raw/              # Sources layer (human-curated)
│   ├── articles/
│   ├── papers/
│   └── transcripts/
├── wiki/             # Content layer (LLM-maintained)
│   ├── concepts/
│   ├── entities/
│   └── sources/
└── SCHEMA.md         # Configuration layer (co-evolved)

Navigation Files:

  • index.md - Content catalog (wiki layer)
  • log.md - Operation history (wiki layer)
  • Both maintained by LLM according to schema specifications

Schema Evolution

The schema layer is unique in being co-evolved rather than owned by either human or LLM. As usage patterns emerge and domain needs evolve, both parties contribute to schema refinement:

  • Human adds new domains or changes priorities
  • LLM suggests workflow improvements based on operational experience
  • Both parties iterate on conventions and standards

This evolutionary approach ensures the system adapts to actual usage rather than theoretical ideals.

Scaling Considerations

The architecture scales naturally:

  • Raw sources can grow indefinitely without affecting other layers
  • Wiki layer scales through improved navigation and search tools
  • Schema layer remains stable once mature, requiring only periodic refinement

See also

Wiki Maintenance Automation

page dédiée →

The systematic automation of knowledge base maintenance tasks using LLMs, addressing the core problem that causes humans to abandon personal wikis: maintenance burden grows faster than value. Central to andrej-karpathy's llm-wiki-pattern.

The Maintenance Problem

Why Humans Abandon Wikis:

  • Updating cross-references becomes tedious
  • Keeping summaries current requires re-reading
  • Noting contradictions between sources is time-consuming
  • Maintaining consistency across dozens of pages is overwhelming
  • Value doesn't justify the increasing effort required

The LLM Solution:

  • Don't get bored with repetitive tasks
  • Don't forget to update cross-references
  • Can touch 15 files in one pass
  • Maintain consistency without fatigue
  • Cost of maintenance approaches zero

Automated Maintenance Tasks

Cross-Reference Management:

  • Automatically create links between related pages
  • Update existing links when page titles change
  • Identify missing connections between related concepts
  • Maintain bidirectional linking consistency

Content Synchronization:

  • Update summary pages when source pages change
  • Propagate corrections across multiple related pages
  • Keep entity information consistent across references
  • Maintain version consistency across the knowledge base

Contradiction Detection:

  • Flag when new sources contradict existing claims
  • Highlight areas where information needs reconciliation
  • Track confidence levels of conflicting statements
  • Suggest resolution strategies for contradictions

Quality Assurance:

  • Identify orphan pages with no inbound links
  • Find concepts mentioned but lacking dedicated pages
  • Check for broken or outdated references
  • Verify citation accuracy and completeness

Implementation Patterns

Ingest-Time Maintenance:

  • Single source addition triggers updates to 10-15 related pages
  • Automatic cross-referencing during content integration
  • Real-time contradiction flagging and resolution
  • Index updates and consistency checks

Periodic Lint Operations:

  • Scheduled health checks across entire knowledge base
  • Systematic identification of maintenance needs
  • Batch correction of consistency issues
  • Performance optimization and cleanup

Query-Driven Updates:

  • Update pages based on insights from user questions
  • File good answers as new wiki pages
  • Strengthen cross-references based on query patterns
  • Evolve structure based on usage patterns

Workflow Integration

Schema-Driven: Maintenance workflows defined in schema document and consistently applied across all operations.

Logging: All maintenance actions recorded in chronological log for transparency and debugging.

Human Oversight: Maintenance automated but humans retain control over curation, priorities, and strategic decisions.

Tools and Infrastructure

Index Management: Automated updating of content catalogs and navigation aids.

Search Integration: Maintenance of search indices and query optimization.

Version Control: Git integration for tracking changes and enabling rollback.

Quality Metrics: Automated assessment of knowledge base health and completeness.

Success Metrics

Maintenance Burden: Time humans spend on bookkeeping approaches zero.

Knowledge Coherence: Cross-references current, contradictions flagged, summaries accurate.

Usage Patterns: Knowledge base becomes more valuable over time rather than degrading.

Growth Sustainability: Adding new sources strengthens rather than fragmenting the knowledge base.

See also