Concepts — vue longue
retour à la listeToutes les pages concaténées sur un seul document, pour un Ctrl-F direct.
Associative Trails
page dédiée →Persistent pathways through information that connect related documents, concepts, and ideas based on meaningful relationships rather than hierarchical organization. Core concept from vannevar-bush's memex-vision and foundation for modern knowledge management approaches including the llm-wiki-pattern.
Historical Origin
Introduced by vannevar-bush in his 1945 essay "As We May Think," associative trails were envisioned as a fundamental departure from traditional filing systems:
Traditional Filing: Hierarchical organization forces artificial single-category placement of documents.
Associative Trails: Documents connected by meaningful relationships that mirror how human minds work through association and connection.
Core Concept
Trail Building: Users create persistent paths through information, linking documents, passages, and ideas. These trails themselves become valuable intellectual artifacts.
Multiple Connections: Any document can participate in multiple trails, reflecting the reality that information has multiple valid organizational structures.
Trail Persistence: Unlike temporary search results, trails remain as permanent parts of the knowledge structure, available for future navigation and extension.
Modern Implementation
In contemporary knowledge management systems, associative trails manifest as:
Wiki Cross-References: wikilinks that connect related pages based on content relationships rather than folder hierarchy.
Tagged Relationships: Content organized by meaningful tags and relationships rather than rigid categories.
Backlink Networks: Automatic discovery of connections between related content through bidirectional linking.
LLM Wiki Pattern Integration
andrej-karpathy's llm-wiki-pattern implements associative trails through automated cross-referencing:
Automatic Trail Creation: LLM identifies meaningful connections during content ingestion and creates appropriate links.
Trail Maintenance: Cross-references kept current as new information is added, ensuring trails remain valid and valuable.
Emergent Navigation: Multiple pathways through information emerge naturally from content relationships rather than imposed structure.
Value Proposition
Reflects Human Cognition: Mirrors how minds actually work through association and connection rather than rigid hierarchies.
Multiple Access Paths: Same information accessible through various conceptual approaches depending on context and need.
Discovery Enabling: Well-constructed trails reveal unexpected connections and enable serendipitous discovery.
Intellectual Tools: Trails become reusable intellectual artifacts that can be shared, extended, and refined over time.
Implementation Challenges
Manual Maintenance: Traditional associative trail systems require significant human effort to keep connections current and meaningful.
Consistency: Ensuring trail quality and coherence across large knowledge bases becomes overwhelming for individual maintainers.
Discovery: Finding relevant trails when needed requires sophisticated search and navigation capabilities.
LLM Solution
The llm-wiki-pattern addresses traditional challenges through automation:
- Automated identification of meaningful connections
- Consistent maintenance of cross-references
- Dynamic updating as new information arrives
- Quality control across entire knowledge base
See also
- memex-vision
- vannevar-bush
- llm-wiki-pattern
- cross-referencing
- [[knowledge-management-
Bookkeeping Automation
page dédiée →The automation of tedious knowledge management tasks that typically cause humans to abandon personal wikis. Core insight behind andrej-karpathy's llm-wiki-pattern - LLMs excel at the systematic maintenance work that humans find burdensome but that's essential for knowledge base health.
The Abandonment Problem
Human Pattern: Start enthusiastic personal wiki → maintenance burden grows → abandon due to tedium Root Cause: "The maintenance burden grows faster than the value"
Tedious Tasks That Cause Abandonment
- Updating cross-references when adding new information
- Keeping summaries current as understanding evolves
- Noting when new data contradicts old claims
- Maintaining consistency across dozens of pages
- Filing and organizing information systematically
- Creating and maintaining category structures
LLM Advantages for Bookkeeping
No Boredom: LLMs don't get tired of repetitive maintenance tasks
Perfect Memory: Never forget to update a cross-reference or flag a contradiction
Batch Processing: Can touch 15 files in one pass during integration
Systematic Approach: Consistently apply rules and conventions across entire knowledge base
Pattern Recognition: Identify maintenance needs humans might miss
Automated Bookkeeping Tasks
During ingest-workflow
- Update entity pages when new sources mention existing entities
- Create new entity pages for newly mentioned people/organizations/tools
- Revise concept summaries with new information
- Add cross-references between related topics
- Flag contradictions with existing content
- Update index and categorization
During query-workflow
- File valuable answers as new wiki pages
- Update existing pages with insights from queries
- Create new cross-references discovered through questioning
During lint-workflow
- Find and fix broken internal links
- Identify orphan pages needing better integration
- Suggest missing cross-references
- Flag outdated claims for human review
Cost-Benefit Transformation
Traditional Wiki:
- High human maintenance cost
- Maintenance burden compounds with scale
- Value growth plateaus while cost accelerates
LLM Wiki Pattern:
- Near-zero maintenance cost (automated)
- Value compounds while cost stays flat
- Sustainable knowledge accumulation
Human-LLM Role Division
Human Responsibilities:
- Curate sources (what to include)
- Direct analysis (what to emphasize)
- Ask good questions (what to explore)
- Think about meaning (what it all means)
LLM Responsibilities:
- All bookkeeping and maintenance
- Cross-referencing and consistency
- Systematic organization and filing
- Routine updates and corrections
Implementation Strategy
The key is documenting bookkeeping requirements in the schema layer so LLMs can execute them consistently. This includes:
- Cross-reference conventions
- Page update triggers
- Consistency checking rules
- Integration workflows
- Quality standards
See also
- llm-wiki-pattern
- wiki-maintenance-automation
- andrej-karpathy
- three-layer-architecture
- ingest-workflow
- lint-workflow
Compounding Artifacts
page dédiée →Knowledge artifacts that grow in value and sophistication over time through incremental addition and integration of new information, rather than through simple accumulation. The key principle underlying andrej-karpathy's llm-wiki-pattern and modern approaches to persistent-learning.
Core Concept
Value Multiplication: Each new source doesn't just add isolated information - it enhances the entire knowledge base by creating cross-references, resolving contradictions, strengthening synthesis, and revealing new connections between existing concepts.
Persistent Structure: Unlike query-time retrieval systems, compounding artifacts maintain their organized structure between interactions. The cross-references are already there, contradictions already flagged, synthesis already complete and continuously updated.
Semantic Integration: Information from multiple sources synthesizes into coherent understanding that becomes greater than the sum of individual sources through systematic integration rather than simple accumulation.
Compounding Mechanisms
Cross-Reference Network
Each new source creates links to existing entities and concepts, building a knowledge graph where value emerges from connections as much as content.
Contradiction Resolution
New information challenges existing claims, forcing refinement of understanding and highlighting areas requiring deeper investigation.
Synthesis Evolution
Overarching themes and analyses evolve continuously as new perspectives integrate with existing frameworks.
Context Enrichment
Entity profiles and concept definitions deepen with each mention across sources, building comprehensive understanding over time.
Implementation in LLM Wikis
Incremental Integration: Single sources typically update 10-15 existing wiki pages while creating new concept/entity pages, maximizing integration density.
Automated Maintenance: LLMs handle the tedious cross-referencing and consistency work that causes humans to abandon knowledge bases, enabling continuous compounding.
Query Amplification: Valuable query answers get filed back as new wiki pages, turning exploration into knowledge base expansion.
Contrast with Document Storage
Traditional Approach: Each document exists in isolation, requiring reassembly of knowledge for each new question or analysis.
Compounding Approach: Knowledge integrates continuously, with each addition making the entire base more valuable through enhanced connections and refined understanding.
Examples
Personal Research Wiki: Papers and articles integrate into evolving concept pages with cross-referenced relationships and emerging thesis development.
Book Companion Wiki: Characters, themes, and plot threads build incrementally while reading, creating rich companion resource comparable to community-built fan wikis.
Business Intelligence Wiki: Customer calls, meeting transcripts, and market reports synthesize into continuously updated competitive analysis and strategic insight.
Historical Foundation
Builds on vannevar-bush's memex-vision where the connections between documents become as valuable as the documents themselves. Modern LLM automation solves Bush's unsolved maintenance problem.
See also
- llm-wiki-pattern - Framework for building compounding knowledge artifacts
- persistent-learning - Paradigm for accumulating vs. rediscovering knowledge
- knowledge-compilation - Process of transforming sources into integrated artifacts
- incremental-knowledge-building - Methodology for systematic information integration
Human-LLM Division of Labor
page dédiée →The strategic allocation of responsibilities between humans and LLMs in llm-wiki-pattern systems, optimizing each party's strengths while solving the traditional wiki abandonment problem. Core insight: humans excel at curation and synthesis; LLMs excel at maintenance and bookkeeping.
Task Allocation
Human Responsibilities
- Source curation: Selecting valuable documents to ingest
- Direction setting: Guiding analysis emphasis and priorities
- Question asking: Driving exploration through strategic queries
- Meaning synthesis: Understanding implications and significance
- Quality oversight: Reviewing summaries and checking updates
- Schema evolution: Adapting system configuration based on needs
LLM Responsibilities
- Content maintenance: Updating cross-references, keeping summaries current
- Consistency management: Noting contradictions, maintaining coherence
- Bookkeeping automation: Filing, indexing, logging operations
- Cross-referencing: Building and maintaining wikilinks networks
- Structure creation: Generating new pages, organizing content
- Workflow execution: Following schema-defined operational procedures
Solving the Abandonment Problem
Traditional Wiki Failure Pattern
- Initial enthusiasm: Humans start with high motivation
- Growing burden: Maintenance tasks accumulate faster than value
- Cognitive overhead: Cross-referencing becomes mentally taxing
- Inevitable abandonment: Effort required exceeds perceived benefit
LLM Solution
- Zero maintenance fatigue: LLMs don't experience tedium or boredom
- Consistent execution: Never forget to update cross-references
- Parallel processing: Can touch 15 files in one pass without cognitive load
- Near-zero cost: Maintenance burden becomes negligible
Cognitive Complementarity
Human Cognitive Strengths
- Contextual judgment: Understanding significance and relevance
- Creative synthesis: Making novel connections and insights
- Domain expertise: Applying specialized knowledge and intuition
- Strategic thinking: Long-term planning and goal-oriented exploration
LLM Cognitive Strengths
- Systematic processing: Consistent application of rules and procedures
- Pattern recognition: Identifying structural relationships across content
- Parallel attention: Managing multiple interconnected updates simultaneously
- Infinite patience: Performing repetitive tasks without degradation
Operational Boundaries
Human Decision Points
- Source selection: What documents deserve ingestion?
- Emphasis guidance: What aspects need highlighting?
- Quality gates: Are summaries accurate and useful?
- Exploration direction: What questions should drive further investigation?
LLM Execution Points
- Content integration: How to incorporate new information?
- Link maintenance: Which pages need cross-reference updates?
- Consistency checks: Where do contradictions need flagging?
- Structure organization: How to categorize and file content?
Interface Design
Human-LLM Interaction Model
- LLM agent open on one side of screen
- Obsidian (or wiki browser) open on other side
- Real-time collaboration: Human monitors LLM edits live
- Immediate feedback: Human can guide and correct during operation
Communication Patterns
- Explicit instructions: Human provides clear direction for emphasis
- Progress reporting: LLM describes what updates are being made
- Quality confirmation: Human reviews and approves significant changes
- Schema discussion: Collaborative evolution of system configuration
Benefits of Clear Division
Efficiency Optimization
- Leverage strengths: Each party focuses on optimal tasks
- Minimize waste: Avoid humans doing tedious work, LLMs making judgment calls
- Sustainable workflow: Maintenance burden doesn't grow with scale
Quality Assurance
- Human oversight: Strategic decisions remain under human control
- LLM consistency: Mechanical tasks executed reliably
- Complementary validation: Different cognitive approaches catch different errors
Implementation Considerations
Trust Building
- Gradual automation: Start with supervised workflows, increase autonomy
- Transparency: LLM reports all changes and reasoning
- Reversibility: Git versioning enables rollback of problematic updates
Workflow Evolution
- Usage-driven refinement: Division of labor adapts based on experience
- Domain customization: Different fields may require different task allocations
- Tool integration: Technical capabilities influence responsibility boundaries
See also
- llm-wiki-pattern - Framework implementing this division
- wiki-maintenance-automation - Automated tasks LLMs handle
- bookkeeping-automation - Specific maintenance functions
- schema-coevolution - Collaborative system configuration process
Knowledge Compilation
page dédiée →The process of systematically transforming raw source documents into structured, cross-referenced knowledge artifacts that persist between interactions. Core innovation of the llm-wiki-pattern that contrasts sharply with query-time retrieval approaches in traditional RAG systems.
Compilation vs Retrieval
Traditional RAG: Fragments documents into chunks, retrieves relevant pieces at query time, and re-synthesizes knowledge for each interaction. Knowledge is rediscovered from scratch repeatedly.
Knowledge Compilation: Pre-processes sources into integrated wiki structures where cross-references exist, contradictions are flagged, and synthesis reflects accumulated reading. Knowledge is compiled once and maintained continuously.
Compilation Process
Initial Integration
When ingesting new sources, the LLM:
- Reads and extracts key information from raw documents
- Integrates across existing pages - typically touching 10-15 wiki pages per source
- Creates new entity/concept pages for previously uncovered topics
- Updates cross-references to maintain knowledge graph connectivity
- Notes contradictions where new information challenges existing claims
- Strengthens synthesis by incorporating supporting evidence
Incremental Refinement
Each compilation cycle builds on previous work:
- Concept pages deepen with additional sources and perspectives
- Entity profiles expand with new activities, relationships, and attributes
- Cross-references multiply as connections between topics emerge
- Synthesis evolves reflecting cumulative understanding rather than isolated insights
Architectural Benefits
Persistent Structure: Knowledge exists in organized form between sessions, eliminating need to rebuild understanding from raw sources.
Cumulative Intelligence: Each new source makes the entire knowledge base more valuable through integration rather than simple addition.
Query Efficiency: Questions answered by synthesizing from pre-structured content rather than assembling fragments in real-time.
Maintenance Automation: LLMs handle tedious cross-referencing and consistency work that causes humans to abandon personal wikis.
Implementation Patterns
Three-Layer Architecture: Raw sources remain immutable while compiled wiki layer evolves continuously under LLM management guided by schema configuration.
Batch Integration: Single sources can update multiple concept areas simultaneously, creating natural knowledge clustering and relationship discovery.
Version Evolution: Compiled knowledge improves over time as more sources provide additional perspectives and corrections to initial understanding.
Distinction from Document Storage
Unlike traditional document management systems where each source exists in isolation, knowledge compilation creates semantic integration where information from multiple sources synthesizes into coherent, cross-referenced understanding that compounds in value.
The compiled knowledge base becomes greater than the sum of its sources through systematic integration rather than simple accumulation.
See also
- llm-wiki-pattern - Overall framework for persistent knowledge management
- persistent-learning - Paradigm for accumulating rather than rediscovering knowledge
- incremental-knowledge-building - Methodology for systematic information integration
- compounding-artifacts - Knowledge artifacts that grow in value over time
LLM Wiki Pattern
page dédiée →A paradigm for building personal knowledge bases where LLMs incrementally build and maintain persistent wikis rather than retrieving from raw documents at query time. Developed by andrej-karpathy as an alternative to traditional RAG systems that rediscover knowledge from scratch on every interaction.
Core Innovation
Persistent vs Ephemeral Knowledge: Most people's experience with LLMs and documents follows the RAG pattern - upload files, retrieve relevant chunks at query time, generate answers. This works but requires rediscovering knowledge from scratch on every question. The LLM Wiki Pattern instead creates persistent, compounding artifacts where knowledge is compiled once and kept current.
Key Difference: The wiki sits between you and raw sources. When adding new sources, the LLM doesn't just index for later retrieval - it reads, extracts key information, and integrates into existing wiki structure. Updates entity pages, revises summaries, flags contradictions, strengthens synthesis. Cross-references already exist. Contradictions already flagged. Synthesis already reflects everything read.
Three-Layer Architecture
-
Raw Sources: Immutable curated collection (articles, papers, images, data files). LLM reads but never modifies. Source of truth.
-
The Wiki: Directory of LLM-generated markdown files. Summaries, entity pages, concept pages, comparisons, synthesis. LLM owns this layer entirely - creates, updates, maintains cross-references, ensures consistency.
-
The Schema: Configuration document (CLAUDE.md, AGENTS.md) defining wiki structure, conventions, workflows. Co-evolved between human and LLM over time as patterns emerge.
Core Operations
Ingest Workflow
Drop new source into raw collection. LLM reads source, discusses takeaways, writes summary page, updates index, updates relevant entity/concept pages across wiki, appends log entry. Single source might touch 10-15 wiki pages. Can be done one-at-a-time with supervision or batch-processed.
Query Workflow
Ask questions against wiki. LLM searches relevant pages, reads them, synthesizes answer with citations. Answers can take multiple forms - markdown pages, comparison tables, slide decks (marp-integration), charts, canvas. Critical insight: Good answers get filed back as new wiki pages. Explorations compound in knowledge base like ingested sources.
Lint Workflow
Periodic health-checking. Look for contradictions between pages, stale claims superseded by newer sources, orphan pages with no inbound links, important concepts mentioned but lacking pages, missing cross-references, data gaps. LLM suggests new questions to investigate and sources to find.
Navigation Infrastructure
Index and Logging
-
index.md: Content-oriented catalog of all wiki pages with links, summaries, metadata. Organized by category. Updated on every ingest. LLM reads index first to find relevant pages for queries. Works well at moderate scale (~100 sources, hundreds of pages).
-
log.md: Chronological append-only record of ingests, queries, lint passes. Parseable with consistent prefixes (
## [2026-04-02] ingest | Article Title). Provides timeline of wiki evolution.
Optional Tooling
qmd-search engine recommended for scaling beyond index-based navigation. Local search for markdown with hybrid BM25/vector search and LLM re-ranking. Available as CLI tool and MCP server.
Implementation Recommendations
Obsidian Integration
Recommended interface: obsidian-integration where human browses in Obsidian while LLM makes live edits to markdown files. "Obsidian is the IDE; the LLM is the programmer; the wiki is the codebase."
Useful Obsidian Features:
- Web Clipper: Browser extension converting articles to markdown
- Image handling: Download attachments locally, bind to hotkey (Ctrl+Shift+D)
- Graph view: Visualize wiki connections, identify hubs and orphans
- marp-integration: Generate slide decks from wiki content
- dataview-plugin: Query page frontmatter for dynamic tables/lists
- Git integration: Version history, branching, collaboration
Use Cases
Personal: Goals, health, psychology tracking. File journal entries, articles, podcast notes into structured self-picture over time.
Research: Deep topic exploration over weeks/months. Build comprehensive wiki with evolving thesis.
Reading Companion: File chapters as you go. Build pages for characters, themes, plot threads. Like fan wikis (Tolkien Gateway) but personal with LLM maintenance.
Business/Team: Internal wiki fed by Slack threads, meeting transcripts, project docs, customer calls. Humans review updates.
Other: Competitive analysis, due diligence, trip planning, course notes, hobby deep-dives.
Historical Foundation
Explicitly references vannevar-bush's memex-vision (1945) as spiritual predecessor. Bush envisioned personal, curated knowledge store with associative-trails between documents. His vision was private, actively curated, with connections between documents as valuable as documents themselves.
Key Insight: Bush couldn't solve the maintenance problem. LLMs handle that through bookkeeping-automation. Humans abandon wikis because maintenance burden grows faster than value. LLMs don't get bored, don't forget cross-references, can touch 15 files in one pass. Maintenance cost approaches zero.
Division of Labor
Human Role: Curate sources, direct analysis, ask good questions, think about meaning.
LLM Role: Summarizing, cross-referencing, filing, bookkeeping that makes knowledge base useful over time.
Design Philosophy
Intentionally abstract specification describing the pattern, not specific implementation. Directory structure, schema conventions, page formats, tooling all depend on domain, preferences, LLM choice. Everything modular and optional - pick what's useful. Designed to be shared with LLM agents for collaborative instantiation.
See also
Memex Vision
page dédiée →vannevar-bush's 1945 vision of a personal knowledge management system that would allow individuals to store, retrieve, and create associative-trails between documents and information. A foundational concept for modern knowledge management and the direct inspiration behind andrej-karpathy's llm-wiki-pattern and other AI-powered information systems.
Historical Context
Introduced in Bush's essay "As We May Think" (1945), the Memex was conceived as a mechanical device that would supplement human memory by allowing rapid consultation of personal collections of documents. Bush envisioned a desk-sized device with screens, keyboards, and mechanical storage systems.
Core Vision
Private curation - Personal knowledge stores actively maintained by individuals rather than shared databases
Associative trails - Connections between documents as valuable as the documents themselves. Users could create and follow paths linking related information across their collection.
Mechanical augmentation - Technology amplifying human cognitive capabilities rather than replacing human judgment
Active maintenance - Knowledge bases requiring continuous organization and cross-referencing to remain valuable
The Maintenance Problem
Bush's vision was prescient but incomplete. He understood the value of connected, curated knowledge but couldn't solve the fundamental challenge: who does the maintenance work?
Creating associative trails, maintaining cross-references, keeping summaries current, noting contradictions - this bookkeeping is essential but tedious. Humans consistently abandon personal knowledge systems because maintenance burden grows faster than perceived value.
Modern Resolution
The llm-wiki-pattern directly addresses Bush's unsolved maintenance problem. LLMs excel at the systematic bookkeeping that humans find burdensome:
- Updating cross-references across multiple pages
- Maintaining consistency as information accumulates
- Noting contradictions between sources
- Creating and strengthening associative connections
This allows Bush's original vision to be realized: private, actively curated knowledge stores where connections between documents are as valuable as documents themselves.
Influence on Contemporary Systems
The Memex vision influenced:
- Hypertext systems and the World Wide Web
- Personal knowledge management tools (Roam, Obsidian, Notion)
- Information retrieval research
- Modern AI-powered knowledge systems
However, most implementations lost Bush's emphasis on private curation and became either:
- Public databases (Wikipedia, web)
- Simple storage without maintained associations (file systems)
- Tools requiring manual maintenance (personal wikis that get abandoned)
Key Insights for LLM Wiki Pattern
Connection value - The relationships between pieces of information often more valuable than isolated documents
Curation over scale - Personal, curated collections outperform generic large databases for individual needs
Maintenance as bottleneck - Technical capability less important than solving the maintenance burden problem
Human-machine collaboration - Best results combine human judgment (curation, direction) with machine capability (systematic maintenance)
See also
Obsidian Integration
page dédiée →The recommended interface pattern for llm-wiki-pattern implementations where humans browse and explore wiki content in obsidian while LLMs act as programmers making live edits to the underlying markdown files. Creates seamless collaboration between human intuition and LLM maintenance.
Integration Model
Collaborative Workflow:
- LLM agent open on one side, Obsidian open on the other
- LLM makes edits based on conversation and source ingestion
- Human browses results in real time - following links, checking graph view, reading updated pages
- Metaphor: "Obsidian is the IDE; the LLM is the programmer; the wiki is the codebase"
This creates a natural division of labor where humans excel at exploration and pattern recognition while LLMs handle systematic maintenance and content generation.
Essential Obsidian Features
Graph View:
- Best way to visualize wiki structure and relationships
- Identifies hubs (highly connected pages) and orphans (isolated pages)
- Reveals emerging themes and knowledge clusters
- Essential for understanding wiki evolution over time
Real-time Markdown Editing:
- See LLM changes appear instantly in Obsidian
- Navigate between updated pages immediately
- Visual feedback on cross-reference creation and updates
Wikilink Support:
- Native support for
internal linksbetween pages - Automatic backlinking and reference tracking
- Foundation for the associative trail navigation
Recommended Plugins and Tools
Obsidian Web Clipper:
- Browser extension that converts web articles to markdown
- Essential for quickly adding sources to raw collection
- Preserves formatting and structure for LLM processing
Dataview Plugin:
- Runs queries over page frontmatter (YAML metadata)
- Generates dynamic tables and lists from wiki content
- Useful when LLM adds structured metadata (tags, dates, source counts)
Marp Plugin:
- Markdown-based slide deck format support
- Enables generating presentations directly from wiki content
- LLM can create slide decks as query output format
Attachment Management:
- Set "Attachment folder path" to fixed directory (e.g.,
raw/assets/) - Bind "Download attachments for current file" to hotkey (e.g., Ctrl+Shift+D)
- Downloads all images locally for LLM analysis and stable references
Image Handling Workflow
Process:
- Clip article with Obsidian Web Clipper
- Hit download hotkey to store all images locally
- LLM reads text first, then examines referenced images separately
- Slightly clunky but works well for multimodal content integration
Benefits:
- Stable image references (URLs don't break)
- LLM can analyze images directly
- Complete offline functionality
Version Control Integration
Since the wiki is just a git repo of markdown files:
- Full version history of all changes
- Branch different analysis directions
- Collaborate with others through standard git workflows
- Track LLM edits over time for quality assessment
Navigation Patterns
Index-Driven Exploration:
- Start with
index.mdfor wiki overview - Follow category-based organization
- Jump to specific pages via search or links
Graph-Driven Discovery:
- Use graph view to find connection patterns
- Identify knowledge gaps (missing connections)
- Explore topic clusters and emerging themes
Log-Based Timeline:
- Review
log.mdfor chronological wiki evolution - Understand how knowledge accumulated over time
- Identify periods of high activity or major updates
Real-time Collaboration Benefits
Immediate Feedback:
- See LLM understanding reflected in page updates
- Catch misinterpretations or missed connections quickly
- Guide LLM emphasis through real-time review
Enhanced Exploration:
- Follow interesting connections as they're created
- Discover unexpected relationships through graph visualization
- Build on LLM insights through guided browsing
Quality Assurance:
- Verify cross-references are accurate and meaningful
- Ensure page updates maintain consistency
- Check that new content integrates well with existing knowledge
Scaling Considerations
As wikis grow beyond ~100 pages:
- Graph view becomes dense but still valuable for cluster identification
- Search becomes more important than browsing
- Consider complementary tools like qmd for advanced search capabilities
The Obsidian interface remains effective even at scale due to its flexible navigation options and powerful visualization capabilities.
See also
- llm-wiki-pattern - Overall framework this integration supports
- three-layer-architecture - How wiki layer interfaces with other components
- associative-trails - Navigation patterns enabled by wikilinks
- graph-view-analysis - Advanced techniques for wiki structure analysis
Persistent Learning
page dédiée →A paradigm in AI systems where knowledge accumulates and compounds over time rather than being re-derived from scratch for each interaction. Contrasts with stateless approaches where each query starts fresh. Core concept underlying andrej-karpathy's llm-wiki-pattern and compounding-artifacts.
Core Concept
Traditional systems like RAG rediscover knowledge from scratch on every question - finding relevant chunks, piecing together fragments, synthesizing answers from raw sources. Nothing accumulates. Ask a subtle question requiring synthesis across five documents, and the LLM repeats the same discovery process every time.
Persistent learning systems instead build and maintain accumulated knowledge structures that compound over interactions. The synthesis work is done once and then kept current, not repeated. Each new source strengthens the existing structure rather than existing in isolation.
Implementation in LLM Wiki Pattern
The llm-wiki-pattern implements persistent learning through:
Knowledge compilation - Raw sources transformed into structured, cross-referenced wiki pages that persist between sessions
Incremental integration - New information updates existing pages rather than creating isolated summaries
Maintained synthesis - Cross-references, contradictions, and connections continuously maintained as knowledge base evolves
Compounding exploration - Good query answers filed back as wiki pages, making explorations persistent rather than ephemeral
Benefits Over Stateless Systems
- Accumulated insights: Connections and contradictions already identified
- Refined understanding: Multiple sources integrated into coherent picture
- Efficient querying: Pre-compiled knowledge vs. real-time fragment assembly
- Progressive deepening: Each interaction builds on previous work
- Preserved discoveries: Valuable insights don't disappear into chat history
Historical Context
Concept relates to vannevar-bush's memex-vision - personal knowledge stores that accumulate and cross-reference over time. Also parallels human learning where new information integrates with existing knowledge structures rather than existing in isolation.
Implementation Challenges
Requires systematic maintenance that humans find tedious but LLMs handle naturally:
- Updating cross-references across multiple pages
- Noting contradictions between old and new information
- Maintaining consistency as knowledge base grows
- Organizing and categorizing accumulated knowledge
See also
RAG Alternative
page dédiée →The llm-wiki-pattern represents a fundamental alternative to traditional Retrieval-Augmented Generation (RAG) approaches. Instead of retrieving and synthesizing from raw documents on each query, knowledge is compiled once into structured wikis and maintained persistently.
Traditional RAG Limitations
Repeated Rediscovery: RAG systems upload document collections, retrieve relevant chunks at query time, and generate answers from fragments. The LLM rediscovers knowledge from scratch on every question. No accumulation or learning occurs.
Fragment Synthesis: Complex questions requiring synthesis of multiple sources force the LLM to find and piece together relevant fragments repeatedly. Nothing is built up or persists between queries.
Shallow Integration: Sources exist in isolation. Cross-references, contradictions, and deeper synthesis must be discovered anew for each interaction.
Wiki Pattern Advantages
Knowledge Compilation: Information is processed once during ingestion, integrated into existing understanding, cross-referenced with related content, and maintained as living knowledge base.
Persistent Synthesis: Complex analysis is performed once and stored. Cross-references are pre-computed, contradictions are pre-identified, synthesis reflects everything previously ingested.
Compounding Value: Each new source strengthens the entire knowledge base rather than existing in isolation. Value grows exponentially rather than linearly.
Architectural Comparison
| Approach | Knowledge Storage | Query Processing | Maintenance | Value Growth |
|---|---|---|---|---|
| RAG | Raw documents + embeddings | Retrieve fragments → synthesize | None | Linear |
| Wiki Pattern | Structured, cross-referenced pages | Read relevant pages → reference | Automated by LLM | Exponential |
Implementation Trade-offs
RAG Advantages:
- Simpler initial setup
- No maintenance overhead
- Works well for straightforward Q&A
- Established tooling ecosystem
Wiki Pattern Advantages:
- Knowledge compounds over time
- Complex synthesis performed once
- Rich cross-referencing and navigation
- Handles contradictions systematically
- Scales to deeper analysis
When to Use Each Approach
RAG Appropriate For:
- Simple document search and retrieval
- Static document collections
- Minimal ongoing engagement
- Straightforward Q&A scenarios
Wiki Pattern Appropriate For:
- Long-term knowledge building projects
- Research requiring synthesis across sources
- Personal knowledge management
- Complex domain understanding
- Ongoing learning and exploration
Hybrid Possibilities
Complementary Usage: Wiki pattern for core knowledge base with RAG for supplementary document retrieval. Wiki handles synthesis and cross-referencing while RAG provides access to broader document collections.
Migration Path: Start with RAG for document exploration, migrate valuable synthesis to wiki format for persistent reference and further development.
Technical Requirements
Wiki Pattern Demands:
- LLM capable of multi-document reasoning
- File system access for wiki maintenance
- Schema-driven workflow execution
- Structured output generation (markdown, YAML)
Integration Complexity: Higher initial setup cost but lower ongoing maintenance burden compared to RAG systems that require continuous re-processing.
See also
RAG Alternative Approaches
page dédiée →Methods for working with large document collections that go beyond traditional Retrieval-Augmented Generation patterns. Most notably exemplified by andrej-karpathy's llm-wiki-pattern, which creates persistent, pre-synthesized knowledge structures rather than retrieving raw chunks at query time.
Traditional RAG Limitations
Rediscovery Problem: RAG systems retrieve relevant chunks and generate answers from scratch on every query. No knowledge accumulates—subtle questions requiring synthesis across multiple documents must reconstruct understanding each time.
Fragment-Based Understanding: Working with retrieved chunks rather than integrated knowledge, limiting the ability to develop sophisticated cross-source synthesis.
Stateless Operation: Each query starts fresh with no building upon previous insights or discoveries.
Pre-Synthesis Approach
The llm-wiki-pattern represents a fundamental alternative where:
Knowledge Integration: Instead of retrieving fragments, LLMs incrementally build and maintain structured knowledge representations that integrate information across sources.
Persistent Artifacts: Cross-references, contradictions, and syntheses are identified once and maintained rather than rediscovered on each query.
Compounding Intelligence: Good questions and discoveries get filed back into the knowledge base, creating compounding-artifacts that become smarter over time.
Architectural Differences
Traditional RAG
Sources → Embedding → Vector DB → Retrieval → LLM Generation
Wiki Pattern Alternative
Sources → LLM Integration → Persistent Wiki → Query → Pre-Synthesized Answers
Key Advantages
Deep Synthesis: Can develop sophisticated understanding that builds across multiple sources and interactions.
Context Preservation: Maintains full context of how understanding developed rather than working with isolated fragments.
Knowledge Evolution: The system gets smarter over time rather than remaining static.
Rich Cross-Referencing: Connections between concepts are discovered and maintained automatically.
Implementation Patterns
Incremental Integration
New sources update existing entity pages, revise concept summaries, and note contradictions rather than simply being added to a retrieval corpus.
Pre-Compiled Knowledge
Answers draw from knowledge that has already been synthesized and cross-referenced rather than being generated from raw retrieval.
Active Maintenance
Systems actively identify gaps, contradictions, and opportunities for deeper synthesis rather than passively serving queries.
Use Cases
Particularly effective for:
- Long-term research projects requiring deep synthesis
- Personal knowledge development over months/years
- Complex domains where understanding evolves with exposure
- Business intelligence requiring integrated analysis
- Any context where knowledge should compound rather than remain fragmented
Trade-offs
Computational Investment: Requires upfront processing to build integrated knowledge structures rather than simple indexing.
Maintenance Complexity: Systems must actively maintain consistency and handle contradictions rather than serving static retrievals.
Domain Specificity: Requires careful schema design and workflow definition for specific knowledge domains.
Future Directions
The success of wiki-pattern approaches suggests broader possibilities for AI systems that build persistent, evolving knowledge representations rather than operating in stateless fashion.
See also
RAG Alternative Architectures
page dédiée →Architectural patterns that move beyond traditional Retrieval-Augmented Generation (RAG) to create persistent, compounding knowledge systems. Most notably exemplified by the llm-wiki-pattern, these approaches prioritize knowledge building over document retrieval.
Traditional RAG Limitations
Stateless Rediscovery
Standard RAG workflow:
- User asks question
- System retrieves relevant document chunks
- LLM generates answer from retrieved context
- Process repeats from scratch for each query
Core problem: "the LLM is rediscovering knowledge from scratch on every question. There's no accumulation. Ask a subtle question that requires synthesizing five documents, and the LLM has to find and piece together the relevant fragments every time. Nothing is built up."
Retrieval Quality Bottlenecks
- Limited by chunk similarity matching
- Cannot build complex multi-source arguments
- No persistent understanding of document relationships
- Insights lost after each interaction
No Knowledge Compounding
Each query provides discrete value without strengthening the system's overall understanding or capability.
Alternative: Persistent Wiki Architecture
Knowledge Compilation Pattern
Instead of retrieve-at-query-time:
- Compile knowledge once during ingestion
- Maintain persistent synthesis across sources
- Query against structured knowledge rather than raw documents
- Compound insights through each interaction
Three-Layer Implementation
Following llm-wiki-pattern:
Raw Sources
- Immutable document storage
- Source of truth for all knowledge
- Curated by human expertise
Wiki Knowledge Layer
- LLM-maintained structured pages
- Cross-referenced entities and concepts
- Continuously updated synthesis
- Persistent relationship mapping
Schema/Convention Layer
- Workflow definitions for maintenance
- Quality standards and formats
- Cross-referencing conventions
- Update and integration protocols
Architectural Advantages
Pre-Computed Relationships
- Cross-references already established
- Contradictions already identified
- Synthesis already reflects all sources
- Complex connections preserved
Incremental Enhancement
Each new source:
- Strengthens existing understanding
- Creates new conceptual connections
- Updates relationship networks
- Compounds rather than just adds
Query Efficiency
Answers leverage:
- Pre-existing synthesis work
- Established cross-reference networks
- Previously identified patterns
- Accumulated domain expertise
Implementation Patterns
Wiki-Based Knowledge Bases
- Structured markdown pages for entities/concepts
- Automated cross-referencing systems
- Index-based navigation at small scale
- Search integration as collections grow
Version-Controlled Knowledge
- Git-based storage for complete history
RAG Alternatives
page dédiée →Approaches to knowledge management and information retrieval that move beyond traditional Retrieval-Augmented Generation (RAG) systems. Most notably exemplified by andrej-karpathy's llm-wiki-pattern which treats knowledge as compounding-artifacts rather than static document collections.
Traditional RAG Limitations
Standard RAG Approach:
- Upload collection of files
- LLM retrieves relevant chunks at query time
- Generates answers from fragments
- Rediscovers knowledge from scratch on every question
- No accumulation or synthesis between queries
Problems:
- Subtle questions requiring synthesis across multiple documents must be solved repeatedly
- No building up of understanding over time
- Connections between sources not maintained
- Context limited to what can be retrieved in single query
LLM Wiki Pattern Alternative
Core Difference: Instead of retrieving from raw documents at query time, the LLM incrementally builds and maintains a persistent wiki that sits between user and sources.
Process:
- New sources integrate into existing wiki structure
- Updates entity pages, revises summaries, notes contradictions
- Cross-references already established
- Synthesis reflects all previous learning
- Knowledge compounds with each addition
Key Advantages
Persistent Knowledge: Cross-references exist, contradictions flagged, synthesis current. Wiki keeps getting richer with every source and question.
Maintenance Automation: LLMs handle tedious bookkeeping - updating cross-references, keeping summaries current, maintaining consistency. Humans focus on curation and questions.
Compound Learning: Good answers become new wiki pages. Explorations compound in knowledge base just like ingested sources.
Other Alternative Approaches
Knowledge Graphs: Structured representation of entities and relationships, but typically requires significant manual curation or complex extraction pipelines.
Embedding-Based Memory Systems: Vector representations of experiences/documents that can be retrieved by similarity, but lack explicit structure and cross-referencing.
Agent Memory Architectures: Various approaches to giving AI systems persistent memory, though most focus on conversation history rather than structured knowledge building.
Implementation Considerations
Moving beyond RAG requires:
- Schema Design: Clear workflows and conventions for knowledge maintenance
- Integration Workflows: Systematic processes for incorporating new information
- Quality Control: Methods for ensuring accuracy and consistency
- Navigation Systems: Tools for exploring and searching the knowledge base
Applications
RAG alternatives particularly valuable for:
- Research Synthesis: Building understanding across many sources over time
- Personal Learning: Accumulating knowledge in specific domains
- Business Intelligence: Maintaining current understanding of competitive landscape
- Domain Expertise: Building deep, interconnected knowledge in specialized areas
See also
RAG vs Wiki Compilation
page dédiée →Fundamental architectural comparison between traditional Retrieval-Augmented Generation (RAG) systems and andrej-karpathy's llm-wiki-pattern approach to knowledge management. Represents a shift from ephemeral retrieval to persistent knowledge compilation.
Core Philosophical Difference
Traditional RAG: Knowledge is rediscovered from scratch on every query Wiki Compilation: Knowledge is compiled once and maintained persistently
This difference has profound implications for knowledge accumulation, query efficiency, and system sophistication over time.
Operational Comparison
Traditional RAG Workflow
- User asks question
- System retrieves relevant document chunks
- LLM pieces together fragments at query time
- Generates answer from scratch
- Answer disappears into chat history
Wiki Compilation Workflow
- Sources ingested incrementally into persistent wiki
- Knowledge integrated with existing understanding
- User asks question
- System searches pre-compiled wiki pages
- LLM synthesizes from maintained knowledge
- Good answers filed back as new wiki pages
Architectural Trade-offs
RAG Advantages
- Simplicity: Straightforward to implement and understand
- Source Fidelity: Directly references original document chunks
- Real-time: Can incorporate new documents immediately
- Stateless: No complex state management required
RAG Limitations
- Redundant Work: Same analysis repeated for similar queries
- Fragment Dependency: Quality depends on chunk retrieval accuracy
- No Accumulation: Insights don't compound over time
- Context Loss: Subtle connections require re-discovery each time
Wiki Compilation Advantages
- Knowledge Compounding: Understanding builds and strengthens over time
- Sophisticated Synthesis: Cross-references and contradictions pre-analyzed
- Query Efficiency: Answers draw from compiled knowledge, not raw chunks
- Persistent Learning: System gets smarter with each source and query
Wiki Compilation Limitations
- Complexity: Requires sophisticated maintenance workflows
- Delayed Integration: New sources need processing before availability
- State Management: Must maintain consistency across evolving knowledge base
- Schema Evolution: System architecture must adapt as understanding grows
Scaling Characteristics
RAG: Performance degrades with document volume due to retrieval complexity Wiki Compilation: Performance improves with volume as knowledge density increases
Use Case Suitability
RAG Optimal For
- Document search and basic Q&A
- Large, relatively static document collections
- Simple factual retrieval tasks
- Systems requiring immediate source transparency
Wiki Compilation Optimal For
- Research and analysis over time
- Complex synthesis across multiple sources
- Personal knowledge management
- Domains requiring evolving understanding
Knowledge Quality Evolution
RAG: Knowledge quality remains constant - system doesn't learn Wiki Compilation: Knowledge quality improves through:
- Cross-source validation and contradiction resolution
- Incremental refinement of understanding
- Connection discovery between previously isolated concepts
- Synthesis sophistication increases with experience
Implementation Complexity
RAG: Vector databases, embedding models, retrieval tuning Wiki Compilation: Schema design, maintenance workflows, consistency checking
Future Hybrid Approaches
Emerging systems may combine both approaches:
- Wiki compilation for core knowledge domains
- RAG fallback for novel or peripheral queries
- Automatic promotion from RAG results to wiki compilation
See also
RAG vs Wiki Pattern
page dédiée →Fundamental comparison between traditional Retrieval-Augmented Generation (RAG) systems and andrej-karpathy's llm-wiki-pattern, highlighting different approaches to knowledge management and query answering.
Traditional RAG Approach
Process Flow
- Upload collection of files
- Chunk and embed documents
- At query time: retrieve relevant chunks
- Generate answer from retrieved fragments
- Repeat process for each new query
Characteristics
- Stateless: Each query starts fresh
- Rediscovery: Knowledge reconstructed every time
- Fragment-based: Works with document chunks, not integrated knowledge
- No accumulation: Understanding doesn't build over time
- Limited synthesis: Difficult to connect insights across multiple documents
Examples
- NotebookLM
- ChatGPT file uploads
- Most commercial RAG systems
- Document Q&A tools
LLM Wiki Pattern Approach
Process Flow
- Ingest sources into persistent wiki structure
- LLM builds and maintains cross-referenced knowledge base
- At query time: synthesize from existing wiki pages
- File valuable answers back as new wiki content
- Knowledge compounds with each interaction
Characteristics
- Stateful: Persistent knowledge accumulation
- Pre-synthesized: Understanding built incrementally over time
- Integration-based: Works with structured, cross-referenced content
- Compounding: Each addition makes the whole more valuable
- Rich synthesis: Connections already established across sources
Key Differences
| Aspect | Traditional RAG | Wiki Pattern |
|---|---|---|
| Knowledge State | Stateless retrieval | Persistent accumulation |
| Processing Time | Query-time discovery | Ingest-time integration |
| Cross-reference | Ad-hoc during query | Pre-established and maintained |
| Contradictions | Discovered per query | Flagged and tracked systematically |
| Synthesis Quality | Limited by retrieval | Rich, pre-built connections |
| Maintenance | None (static chunks) | Automated wiki maintenance |
When to Use Each
RAG Works Well For
- Simple document Q&A
- One-off queries against large document sets
- When you don't need knowledge to compound
- Rapid deployment without setup overhead
- Documents that rarely need cross-referencing
Wiki Pattern Works Well For
- Long-term knowledge building
- Complex synthesis across multiple sources
- Research that builds over time
- When contradictions and evolution matter
- Domains where connections between concepts are valuable
Hybrid Approaches
Some systems might combine both patterns:
- Wiki pattern for core, frequently-accessed knowledge
- RAG for supplementary document collections
- Different layers for different types of content
Performance Implications
RAG
- Pros: Simple setup, works immediately, scales to large document sets
- Cons: Repeated processing overhead, limited synthesis depth, no knowledge accumulation
Wiki Pattern
- Pros: Rich synthesis, compounding value, deep cross-references, maintained consistency
- Cons: Higher setup cost, requires ongoing LLM maintenance, more complex architecture
See also
Role Division
page dédiée →The systematic allocation of responsibilities between humans and LLMs in the llm-wiki-pattern, optimizing for each party's comparative advantages. Solves the knowledge management maintenance problem through cognitive specialization rather than trying to make humans better at tedious tasks.
Human Responsibilities
High-Level Cognitive Tasks
- Source Curation: Decide what information is worth including
- Analysis Direction: Guide what aspects to emphasize or explore
- Question Formation: Ask the right questions to drive useful synthesis
- Meaning Making: Think about what the accumulated knowledge means
- Strategic Decisions: Determine knowledge base scope and priorities
Why Humans Excel Here
- Contextual Judgment: Understanding relevance and quality in complex domains
- Creative Connections: Seeing non-obvious relationships and implications
- Value Alignment: Ensuring knowledge base serves intended purposes
- Intuitive Filtering: Recognizing what's worth investigating further
LLM Responsibilities
Systematic Maintenance Tasks
- Cross-Referencing: Creating and maintaining links between related concepts
- Consistency Management: Ensuring information coherence across pages
- Contradiction Detection: Flagging when new sources conflict with existing content
- Integration Work: Updating multiple pages when new sources arrive
- Organizational Maintenance: Filing, categorizing, and indexing systematically
Why LLMs Excel Here
- No Boredom: Don't get tired of repetitive maintenance work
- Perfect Consistency: Apply rules uniformly across entire knowledge base
- Batch Processing: Can update many pages simultaneously
- Pattern Recognition: Identify maintenance needs systematically
- Attention to Detail: Don't forget to update cross-references or categorization
The Division Principle
Humans do what requires judgment and creativity
LLMs do everything else
This leverages comparative advantages rather than trying to overcome weaknesses. Humans don't need to become better at tedious maintenance; LLMs handle that entirely.
Historical Context
Traditional knowledge management fails because it asks humans to do both:
- High-value work: Thinking, analyzing, connecting
- Maintenance work: Filing, cross-referencing, updating
The maintenance burden eventually overwhelms the value creation, leading to abandonment. Role division solves this by removing maintenance burden from humans entirely.
Implementation in Practice
Human Workflow
- Find interesting sources
- Add to raw collection
- Guide LLM on what to emphasize during ingest
- Ask questions that drive useful synthesis
- Review and approve major changes
- Think about what accumulated knowledge means
LLM Workflow
- Process sources systematically
- Update all affected wiki pages
- Maintain cross-references and
Schema Coevolution
page dédiée →The collaborative development process between human and LLM for evolving the configuration layer of llm-wiki-pattern systems. The schema document (CLAUDE.md, AGENTS.md, etc.) grows and adapts based on actual usage patterns, domain needs, and discovered workflows.
Core Concept
Unlike static system configurations, the schema in an LLM wiki system evolves with use. As humans and LLMs work together, they discover:
- Effective workflows for the specific domain
- Useful page formats and structures
- Domain-specific conventions and tags
- Optimal ingest and query patterns
These discoveries get documented in the schema, making the LLM a more effective collaborator over time.
Evolution Process
Initial Schema
- Basic structure and conventions
- Generic workflows from the pattern
- Minimal domain-specific customization
Usage-Driven Refinement
- Workflow optimization: Discovering efficient ingest patterns
- Convention standardization: Establishing consistent tagging and formatting
- Domain adaptation: Adding field-specific page types and structures
- Tool integration: Incorporating discovered utilities and scripts
Continuous Improvement
- Regular schema updates based on what works
- Documentation of effective practices
- Removal of unused conventions
- Addition of new capabilities
Human-LLM Collaboration
Human Contributions
- Domain expertise: Understanding field-specific needs
- Workflow preferences: Preferred interaction patterns
- Quality standards: Defining acceptable outputs
- Strategic direction: Long-term knowledge goals
LLM Contributions
- Pattern recognition: Identifying recurring structures
- Consistency maintenance: Ensuring schema adherence
- Workflow execution: Following documented procedures
- Improvement suggestions: Proposing optimizations
Schema Components That Evolve
Page Types
- Standard templates for different content types
- Domain-specific entity categories
- Specialized analysis formats (comparisons, timelines, etc.)
Tagging Systems
- Hierarchical tag structures for the domain
- Consistent naming conventions
- Cross-reference patterns
Workflows
- Detailed ingest procedures
- Query and synthesis patterns
- Maintenance and lint operations
- Quality control checkpoints
Tool Integration
- CLI utilities and their usage patterns
- Search and navigation tools
- Export and visualization formats
- Integration with external systems
Benefits
Adaptive Systems
Schema coevolution enables knowledge systems that improve with use rather than becoming stale or rigid.
Domain Optimization
Over time, the system becomes specifically tuned to the user's field and preferences rather than remaining generic.
Reduced Friction
Well-evolved schemas reduce cognitive overhead by codifying effective practices and eliminating decision fatigue.
Knowledge Transfer
The schema serves as documentation of effective knowledge management practices that can be shared or adapted.
Example Evolution
Initial: "Create entity pages for people mentioned"
Evolved: "Create entity pages with standardized sections: Background, Key Ideas, Publications, Influence, Cross-References. Tag with domain (ai-researcher, entrepreneur, academic) and confidence level."
Challenges
Over-Specification
Schemas can become too rigid, constraining useful variation and experimentation.
Version Management
Managing schema changes while maintaining consistency across existing wiki content.
Complexity Growth
Balancing comprehensive documentation with usability for both human and LLM.
Implementation Patterns
Versioned Schemas
Track schema evolution with version control to understand what changes improve effectiveness.
Modular Structure
Organize schema into sections (workflows, conventions, formats) that can evolve independently.
Example-Driven Documentation
Include concrete examples in schema documentation to clarify abstract conventions.
See also
- llm-wiki-pattern - System using schema coevolution
- three-layer-architecture - Schema as configuration layer
- adaptive-systems - Systems that improve with use
- workflow-optimization - Process improvement patterns
Schema-Driven Workflows
page dédiée →Configuration approach where LLM behavior is governed by explicit schema documents that define structure, conventions, and operational workflows. Central to andrej-karpathy's llm-wiki-pattern for creating disciplined wiki maintainers rather than generic chatbots.
Core Concept
Configuration as Documentation: Schema document (e.g., CLAUDE.md, AGENTS.md, SCHEMA.md) serves as comprehensive instruction manual for LLM behavior. Defines how wiki is structured, what conventions to follow, and what workflows to execute for different operations.
Disciplined Automation: Transforms general-purpose LLM into specialized wiki maintainer through explicit operational instructions. Schema constrains and focuses LLM behavior for consistent, reliable knowledge management.
Schema Components
Structural Definitions:
- Directory organization and file naming conventions
- Page format templates and YAML frontmatter requirements
- Cross-referencing patterns and linking conventions
- Category taxonomies and tagging standards
Operational Workflows:
- Ingest procedures (read → summarize → integrate → cross-reference)
- Query handling (search index → read pages → synthesize → optionally file)
- Maintenance routines (lint for contradictions, orphans, missing links)
- Logging formats and indexing standards
Quality Standards:
- Confidence levels and sourcing requirements
- Cross-reference density expectations
- Contradiction handling procedures
- Update propagation rules
Workflow Examples
Ingest Workflow (from schema):
- Read source completely and extract key information
- Create summary page in wiki/sources/ with metadata
- Update 10-15 relevant existing pages across wiki
- Maintain cross-references and flag contradictions
- Update index.md and append to log.md
- Generate flashcards if applicable
Lint Workflow (from schema):
- Scan for contradictions between pages
- Identify orphan pages with no inbound links
- Find mentioned-but-missing concepts needing pages
- Check for stale content superseded by newer sources
- Suggest new questions and sources to investigate
Co-Evolution Pattern
Human-LLM Collaboration: Schema evolves through usage as human and LLM discover what works for specific domain. Initial framework adapts to actual needs and preferences over time.
Domain Adaptation: General pattern customizes to specific use cases - research wikis have different needs than business intelligence or personal knowledge management.
Workflow Refinement: Operational procedures improve based on experience with what produces highest-quality knowledge accumulation.
Technical Implementation
Schema as Context: LLM reads schema document at start of every session to understand current operational parameters and workflow requirements.
Parseable Formats: Log entries and metadata use consistent formats enabling programmatic analysis and tooling development.
Version Control: Schema documents tracked in git alongside wiki content, providing evolution history and rollback capability.
Benefits
Consistency: Ensures uniform quality and structure across all wiki operations regardless of session or time period.
Reliability: Reduces variability in LLM behavior through explicit instruction rather than relying on implicit understanding.
Maintainability: Schema serves as documentation for wiki structure and operational procedures, enabling debugging and improvement.
Scalability: Well-defined workflows allow wiki to grow systematically without degrading quality or consistency.
Configuration Examples
Directory Structure Rules:
wiki/concepts/ # Technical concepts and methods
wiki/entities/ # People, organizations, tools
wiki/sources/ # Individual source summaries
raw/articles/ # Immutable source documents
Page Format Standards:
---
title: Page Title
category: entities|concepts|projects|sources
created: YYYY-MM-DD
updated: YYYY-MM-DD
tags: [lowercase-hyphenated-tags]
sources: [list-of-raw-source-paths]
confidence: high|medium|low
---
See also
- llm-wiki-pattern
- compounding-artifacts
- persistent-learning
- andrej-karpathy
Three Layer Architecture
page dédiée →The foundational architectural pattern underlying andrej-karpathy's llm-wiki-pattern that separates concerns across three distinct layers: raw sources, wiki content, and schema configuration. This separation enables clear ownership models, version control, and collaborative development between humans and LLMs.
Layer Definitions
Raw Sources Layer:
- Curated collection of source documents (articles, papers, images, data files)
- Immutable - LLM reads from them but never modifies them
- Serves as the authoritative source of truth
- Examples:
raw/articles/,raw/papers/,raw/transcripts/
Wiki Layer:
- Directory of LLM-generated markdown files
- Contains summaries, entity pages, concept pages, comparisons, syntheses
- LLM owns this layer entirely - creates, updates, maintains cross-references
- Clear division: humans read it, LLMs write it
- Examples:
wiki/concepts/,wiki/entities/,wiki/sources/
Schema Layer:
- Configuration document (CLAUDE.md, AGENTS.md, SCHEMA.md, etc.)
- Defines wiki structure, conventions, and operational workflows
- Co-evolved between human and LLM based on actual usage patterns
- Makes the LLM a disciplined wiki maintainer rather than generic chatbot
- Contains ingest workflows, query patterns, maintenance procedures
Ownership Model
The architecture establishes clear ownership and modification rights:
Human Responsibilities:
- Curate and add sources to raw layer
- Modify schema based on evolving needs
- Browse and explore wiki content
- Direct analysis and ask questions
LLM Responsibilities:
- Generate and maintain all wiki content
- Follow schema conventions and workflows
- Update multiple pages per source integration
- Maintain cross-references and consistency
Benefits of Separation
Version Control: Each layer can be versioned independently, allowing rollback of wiki changes without affecting sources or schema evolution.
Clear Interfaces: Well-defined boundaries prevent scope creep and maintain system integrity. Sources remain pristine, wiki stays automatically maintained.
Collaborative Development: Humans and LLMs can work simultaneously without conflicts - humans curate sources and refine schema while LLMs maintain wiki content.
Debugging and Maintenance: Issues can be isolated to specific layers, making troubleshooting more systematic.
Implementation Patterns
Directory Structure:
knowledge-base/
├── raw/ # Sources layer (human-curated)
│ ├── articles/
│ ├── papers/
│ └── transcripts/
├── wiki/ # Content layer (LLM-maintained)
│ ├── concepts/
│ ├── entities/
│ └── sources/
└── SCHEMA.md # Configuration layer (co-evolved)
Navigation Files:
index.md- Content catalog (wiki layer)log.md- Operation history (wiki layer)- Both maintained by LLM according to schema specifications
Schema Evolution
The schema layer is unique in being co-evolved rather than owned by either human or LLM. As usage patterns emerge and domain needs evolve, both parties contribute to schema refinement:
- Human adds new domains or changes priorities
- LLM suggests workflow improvements based on operational experience
- Both parties iterate on conventions and standards
This evolutionary approach ensures the system adapts to actual usage rather than theoretical ideals.
Scaling Considerations
The architecture scales naturally:
- Raw sources can grow indefinitely without affecting other layers
- Wiki layer scales through improved navigation and search tools
- Schema layer remains stable once mature, requiring only periodic refinement
See also
- llm-wiki-pattern - Overall framework using this architecture
- schema-coevolution - How configuration evolves over time
- persistent-learning - Why separation enables knowledge accumulation
- obsidian-integration - How humans interface with the wiki layer
Wiki Indexing Patterns
page dédiée →Systematic approaches to organizing and navigating knowledge bases, particularly in llm-wiki-pattern implementations. Two primary patterns serve different navigation needs: content-oriented catalogs and chronological logs.
Core Patterns
Content-Oriented Index (index.md)
Purpose: Catalog all wiki content organized by category and topic.
Structure:
- Links to all wiki pages with one-line summaries
- Organized by category (entities, concepts, sources, syntheses)
- Optional metadata (dates, source counts, confidence levels)
- Updated on every ingest operation
Navigation Model: LLM reads index first to identify relevant pages for queries, then drills into specific content. Works effectively at moderate scale (~100 sources, hundreds of pages) without requiring embedding-based infrastructure.
Example Structure:
## Entities (45 pages)
- openai - AI research company developing GPT models
- anthropic - AI safety company behind Claude
- andrej-karpathy - Former Tesla AI director, advocate of LLM wiki pattern
## Concepts (89 pages)
- [retrieval-augmented-generation](/concepts/retrieval-augmented-generation) - Combining retrieval with generation
- Vector Embeddings - Dense numerical representations of data
Chronological Log (log.md)
Purpose: Append-only record of all wiki operations and evolution.
Structure:
- Timestamped entries for ingests, queries, lint operations
- Consistent format enabling programmatic parsing
- Timeline of wiki evolution and recent activity
- Context for understanding system state
Navigation Model: Understand what's been done recently, track system evolution, provide context for LLM about recent operations.
Example Format:
## [2026-12-21 14:30] ingest | Karpathy LLM Wiki Pattern
- **Source**: raw/articles/karpathy-llm-wiki-pattern.md
- **Pages touched**: [llm-wiki-pattern](/concepts/llm-wiki-pattern), [compounding-artifacts](/concepts/compounding-artifacts), [persistent-learning](/concepts/persistent-learning)
- **Summary**: Ingested comprehensive blueprint for LLM-maintained knowledge bases
Design Principles
Complementary Functions
The two patterns serve different but complementary navigation needs:
- Index: "What knowledge exists and where is it?"
- Log: "How did this knowledge base evolve over time?"
Parseable Formats
Both use consistent formatting that enables programmatic access:
- Index supports category-based filtering and summary generation
- Log enables timeline analysis with simple Unix tools (
grep "^## \[" log.md | tail -5)
LLM-Maintained
Both files are automatically maintained by LLMs during normal operations, eliminating human maintenance burden while providing essential navigation capabilities.
Scaling Considerations
Moderate Scale Effectiveness
Content-oriented indexing works well up to hundreds of pages without requiring sophisticated search infrastructure. The human-readable format provides good overview while remaining LLM-navigable.
Search Enhancement
At larger scales, can be supplemented with search tools:
- Local search engines like qmd for hybrid BM25/vector search
- Full-text indexing for content that exceeds index-based navigation
- MCP servers for programmatic access to search capabilities
Hierarchical Organization
Index structure can evolve to use subcategories and nested organization as content volume grows:
### AI/ML Core (89 pages)
### RAG & Search (47 pages)
### Agent Systems (112 pages)
Implementation Patterns
Automatic Updates
Index updated during every ingest operation as LLM processes new sources and creates/updates wiki pages. Ensures index remains current with no human intervention.
Metadata Integration
Index can include metadata from page frontmatter:
- Creation/update dates
- Source counts indicating how well-researched topics are
- Confidence levels for content reliability
- Tag summaries for topic clustering
Cross-Reference Support
Index serves as hub for cross-reference discovery - LLM uses it to find related pages when updating content or answering queries requiring synthesis across topics.
Alternatives and Extensions
Graph-Based Navigation
Tools like Obsidian's graph view provide visual representation of page connections, complementing text-based index navigation.
Dynamic Queries
Obsidian's Dataview plugin can generate dynamic indexes based on page metadata, automatically categorizing content based on tags or other frontmatter fields.
Search Integration
Can be combined with full-text search engines for complex queries while maintaining human-readable overview structure.
See also
- llm-wiki-pattern
- Content Organization
- Knowledge Base Navigation
- Information Architecture
Wiki Maintenance Automation
page dédiée →The systematic automation of knowledge base maintenance tasks using LLMs, addressing the core problem that causes humans to abandon personal wikis: maintenance burden grows faster than value. Central to andrej-karpathy's llm-wiki-pattern.
The Maintenance Problem
Why Humans Abandon Wikis:
- Updating cross-references becomes tedious
- Keeping summaries current requires re-reading
- Noting contradictions between sources is time-consuming
- Maintaining consistency across dozens of pages is overwhelming
- Value doesn't justify the increasing effort required
The LLM Solution:
- Don't get bored with repetitive tasks
- Don't forget to update cross-references
- Can touch 15 files in one pass
- Maintain consistency without fatigue
- Cost of maintenance approaches zero
Automated Maintenance Tasks
Cross-Reference Management:
- Automatically create links between related pages
- Update existing links when page titles change
- Identify missing connections between related concepts
- Maintain bidirectional linking consistency
Content Synchronization:
- Update summary pages when source pages change
- Propagate corrections across multiple related pages
- Keep entity information consistent across references
- Maintain version consistency across the knowledge base
Contradiction Detection:
- Flag when new sources contradict existing claims
- Highlight areas where information needs reconciliation
- Track confidence levels of conflicting statements
- Suggest resolution strategies for contradictions
Quality Assurance:
- Identify orphan pages with no inbound links
- Find concepts mentioned but lacking dedicated pages
- Check for broken or outdated references
- Verify citation accuracy and completeness
Implementation Patterns
Ingest-Time Maintenance:
- Single source addition triggers updates to 10-15 related pages
- Automatic cross-referencing during content integration
- Real-time contradiction flagging and resolution
- Index updates and consistency checks
Periodic Lint Operations:
- Scheduled health checks across entire knowledge base
- Systematic identification of maintenance needs
- Batch correction of consistency issues
- Performance optimization and cleanup
Query-Driven Updates:
- Update pages based on insights from user questions
- File good answers as new wiki pages
- Strengthen cross-references based on query patterns
- Evolve structure based on usage patterns
Workflow Integration
Schema-Driven: Maintenance workflows defined in schema document and consistently applied across all operations.
Logging: All maintenance actions recorded in chronological log for transparency and debugging.
Human Oversight: Maintenance automated but humans retain control over curation, priorities, and strategic decisions.
Tools and Infrastructure
Index Management: Automated updating of content catalogs and navigation aids.
Search Integration: Maintenance of search indices and query optimization.
Version Control: Git integration for tracking changes and enabling rollback.
Quality Metrics: Automated assessment of knowledge base health and completeness.
Success Metrics
Maintenance Burden: Time humans spend on bookkeeping approaches zero.
Knowledge Coherence: Cross-references current, contradictions flagged, summaries accurate.
Usage Patterns: Knowledge base becomes more valuable over time rather than degrading.
Growth Sustainability: Adding new sources strengthens rather than fragmenting the knowledge base.