~/wiki

Concepts — vue longue

retour à la liste

Toutes les pages concaténées sur un seul document, pour un Ctrl-F direct.

Associative Trails

page dédiée →

Persistent pathways through information that connect related documents, concepts, and ideas based on meaningful relationships rather than hierarchical organization. Core concept from vannevar-bush's memex-vision and foundation for modern knowledge management approaches including the llm-wiki-pattern.

Historical Origin

Introduced by vannevar-bush in his 1945 essay "As We May Think," associative trails were envisioned as a fundamental departure from traditional filing systems:

Traditional Filing: Hierarchical organization forces artificial single-category placement of documents.

Associative Trails: Documents connected by meaningful relationships that mirror how human minds work through association and connection.

Core Concept

Trail Building: Users create persistent paths through information, linking documents, passages, and ideas. These trails themselves become valuable intellectual artifacts.

Multiple Connections: Any document can participate in multiple trails, reflecting the reality that information has multiple valid organizational structures.

Trail Persistence: Unlike temporary search results, trails remain as permanent parts of the knowledge structure, available for future navigation and extension.

Modern Implementation

In contemporary knowledge management systems, associative trails manifest as:

Wiki Cross-References: wikilinks that connect related pages based on content relationships rather than folder hierarchy.

Tagged Relationships: Content organized by meaningful tags and relationships rather than rigid categories.

Backlink Networks: Automatic discovery of connections between related content through bidirectional linking.

LLM Wiki Pattern Integration

andrej-karpathy's llm-wiki-pattern implements associative trails through automated cross-referencing:

Automatic Trail Creation: LLM identifies meaningful connections during content ingestion and creates appropriate links.

Trail Maintenance: Cross-references kept current as new information is added, ensuring trails remain valid and valuable.

Emergent Navigation: Multiple pathways through information emerge naturally from content relationships rather than imposed structure.

Value Proposition

Reflects Human Cognition: Mirrors how minds actually work through association and connection rather than rigid hierarchies.

Multiple Access Paths: Same information accessible through various conceptual approaches depending on context and need.

Discovery Enabling: Well-constructed trails reveal unexpected connections and enable serendipitous discovery.

Intellectual Tools: Trails become reusable intellectual artifacts that can be shared, extended, and refined over time.

Implementation Challenges

Manual Maintenance: Traditional associative trail systems require significant human effort to keep connections current and meaningful.

Consistency: Ensuring trail quality and coherence across large knowledge bases becomes overwhelming for individual maintainers.

Discovery: Finding relevant trails when needed requires sophisticated search and navigation capabilities.

LLM Solution

The llm-wiki-pattern addresses traditional challenges through automation:

  • Automated identification of meaningful connections
  • Consistent maintenance of cross-references
  • Dynamic updating as new information arrives
  • Quality control across entire knowledge base

See also

Automated Content Triage

page dédiée →

System for automatically evaluating and routing content based on quality and relevance criteria before ingestion into knowledge management systems. Essential for preventing information overload while maintaining high signal-to-noise ratios.

Core Methodology

Three-Tier Scoring System

  • HIGH: Direct ingestion - meets quality and relevance thresholds
  • MEDIUM: Manual review queue - potentially valuable but uncertain
  • LOW: Archive without processing - insufficient value or off-topic

Evaluation Criteria

Accept (HIGH):

  • Introduces new concepts, techniques, or patterns not yet documented
  • Contradicts or significantly nuances existing knowledge
  • Directly relates to active projects or core domains
  • High-quality source (papers, official docs, recognized authors)
  • Contains actionable patterns or architectural decisions

Queue for Review (MEDIUM):

  • Potentially relevant but uncertain value
  • Adjacent to core domains but not central
  • Interesting but shallow - might need deeper sources
  • Quality source but content overlap with existing knowledge

Reject (LOW):

  • Outside defined domains of interest
  • Too superficial (brief social media posts without substance)
  • Duplicate of already-ingested content
  • Promotional content disguised as technical content
  • Outdated information superseded in knowledge base

Implementation Patterns

LLM-Based Assessment

Modern triage systems use LLMs to evaluate content against structured criteria:

  • Compare against existing knowledge base index
  • Apply domain-specific relevance scoring
  • Assess content depth and actionability
  • Check for novelty and potential insights

Pre-filtering for Efficiency

For high-volume sources (like screenshots), implement lightweight pre-filtering:

  • Quick domain relevance check ("Is this tech/AI related?")
  • Source quality assessment before deep analysis
  • Token-efficient evaluation to reduce processing costs

Routing and Storage

Automated file organization based on triage results:

  • raw/sources/ for accepted content
  • raw/pending/ for manual review items
  • raw/discarded/ for rejected content (preserved but not processed)

Benefits

Quality Maintenance

  • Prevents dilution of knowledge base with low-value content
  • Maintains focus on core domains of interest
  • Reduces manual curation overhead
  • Improves search and discovery within knowledge base

Efficiency Optimization

  • Reduces LLM processing costs by filtering out irrelevant content
  • Prioritizes valuable content for immediate processing
  • Creates manageable review queues for borderline content
  • Scales with content volume without linear overhead increase

Advanced Features

Adaptive Thresholds

  • Adjust scoring criteria based on knowledge base growth
  • Lower acceptance thresholds for new topic areas
  • Raise standards for well-covered domains
  • Learn from manual review decisions to improve automation

Meta-Analysis

  • Track triage patterns to identify content trends
  • Monitor false positives/negatives in automated decisions
  • Generate insights about content source quality
  • Optimize ingestion pipelines based on triage data

Integration with Knowledge Systems

Pipeline Position

Triage typically occurs early in ingestion pipeline:

  1. Content source detection
  2. Basic format validation
  3. Automated triage and scoring
  4. Routing to appropriate processing path
  5. Full ingestion for accepted content

Feedback Loops

  • Manual review decisions inform future automatic scoring
  • User queries reveal gaps where rejected content might be valuable
  • Knowledge base usage patterns refine relevance criteria
  • Source quality tracking improves upstream filtering

See also

Complete Automation Stack

page dédiée →

Architecture pattern for knowledge management systems that eliminates manual content processing through coordinated automation scripts and intelligent orchestration. Enables autonomous operation of complex multi-channel content ingestion pipelines.

Core Components

1. Collection Scripts

Automated content gathering from multiple sources:

  • Project Sync: collect-projects.sh using rsync to sync markdown files from development directories
  • Screenshot Capture: process-screenshots.sh moving images from iCloud Drive to processing queues
  • Audio Processing: process-recordings.sh for transcribing conference talks and voice memos

2. Git-Based Change Detection

Uses git-diff-tracking to eliminate complex manifest systems:

# Detect exactly what changed since last ingest
git diff --name-status HEAD~1 raw/projects/

3. Daily Orchestration Agent

Centralized automation that runs all pipelines:

  • Executes collection scripts in sequence
  • Processes new content through LLM ingestion
  • Handles error recovery and logging
  • Commits changes and updates tracking

4. iOS Integration

Seamless mobile capture through ios-shortcuts-integration:

  • Share screenshots directly to iCloud Drive folder
  • Automatic sync to processing pipeline
  • Frictionless capture from any iOS app

Implementation Pattern

Directory Structure

automation-system/
├── scripts/
│   ├── collect-projects.sh    # Sync project files
│   ├── process-screenshots.sh # Handle images
│   ├── process-recordings.sh  # Audio transcription
│   └── daily-ingest.md       # Orchestration instructions
├── raw/                      # Staging area
│   ├── projects/            # Synced project files
│   ├── screenshots/inbox/   # Images awaiting processing
│   └── talks/               # Audio transcripts
└── processed/               # Final output location

Intelligent Exclusions

Critical for avoiding noise in large codebases:

# Exclude common noise directories
--exclude="node_modules" \
--exclude=".git" \
--exclude="dist" \
--exclude=".next" \
--exclude="venv"

Error Handling

Robust pipeline design with:

  • SIGPIPE handling for long-running processes
  • Proper exit codes and logging
  • Recovery mechanisms for partial failures
  • Idempotent operations for safe retries

Automation Benefits

1. Zero Manual Overhead

Once configured, the system operates autonomously:

  • New projects automatically detected and ingested
  • Screenshots processed without intervention
  • Conference recordings transcribed and integrated
  • Cross-references updated automatically

2. Compound Learning Effects

Automation enables sophisticated cross-modal integration:

  • Screenshots from social media link to related projects
  • Conference insights connect to technical documentation
  • Project learnings reference research papers automatically

3. Consistent Processing

Eliminates human inconsistency in:

  • Content categorization
  • Cross-reference creation
  • Quality assessment
  • Archive organization

4. Scalable Growth

System capacity grows with content volume:

  • Handles thousands of files efficiently
  • Incremental processing prevents performance degradation
  • Git-based tracking scales to large repositories

Real-World Results

From the epic brain-wiki implementation:

  • 4,060 markdown files automatically synchronized
  • 31 projects processed in first run
  • Zero manual intervention after initial setup
  • Daily operation with autonomous content updates

Architecture Principles

1. Pipeline Composition

Break complex workflows into composable scripts:

  • Each script handles one responsibility
  • Clear interfaces between components
  • Easy to test and debug individually

2. State Management

Use Git as the source of truth:

  • No complex manifest files
  • Built-in versioning and rollback
  • Distributed and reliable

3. Graceful Degradation

System continues operating with partial failures:

  • Individual pipeline failures don't break the whole system
  • Clear error reporting for manual intervention
  • Automatic retry mechanisms where appropriate

4. Human-in-the-Loop

Automation handles routine tasks, humans handle decisions:

  • Content quality assessment
  • New pipeline configuration
  • System monitoring and optimization

Implementation Considerations

Resource Management

  • Use rsync for efficient file synchronization
  • Implement rate limiting for API calls
  • Monitor disk usage and cleanup old files

Security

  • Protect API keys and credentials
  • Use private repositories for sensitive content
  • Implement access controls on automation scripts

Monitoring

  • Log all automation activities
  • Track success/failure rates
  • Monitor resource consumption
  • Set up alerts for system failures

Advanced Features

Conditional Processing

Smart content filtering to avoid redundant work:

# Only process files modified since last run
find raw/projects -name "*.md" -newer .last_ingest

Parallel Execution

For large content volumes:

  • Process multiple projects simultaneously
  • Batch OCR operations for efficiency
  • Async transcription of audio files

Quality Gates

Automated content validation:

  • Check for minimum content length
  • Validate markdown formatting
  • Verify cross-references resolve

See Also

Educational Content Preservation

page dédiée →

Strategic approach to maintaining long-term access to educational materials, particularly relevant for bootcamp alumni and students concerned about platform sustainability or policy changes. Involves creating independent systems to archive, export, and host educational content.

Preservation Motivations

Platform Sustainability Concerns

  • Educational platforms may discontinue services
  • Business model changes affecting content access
  • Lifetime access promises potentially unfulfilled
  • Need for continued learning resource availability

Knowledge Asset Protection

Educational content represents significant investment:

  • Time spent in intensive learning programs
  • Financial investment in bootcamp tuition
  • Accumulated knowledge and exercise solutions
  • Project templates and reference materials

Technical Preservation Strategies

Content Export and Migration

  • Structured Export: Maintaining original organization and metadata
  • Format Conversion: Converting platform-specific formats to standard formats (MDX, Markdown)
  • Asset Preservation: Downloading and archiving multimedia content
  • Relationship Mapping: Preserving links and dependencies between materials

Independent Hosting Solutions

  • Custom LMS Development: Building personal learning management systems
  • Static Site Generation: Creating browsable documentation sites
  • Local Server Deployment: Self-hosted solutions for content access
  • Cloud Storage Integration: Distributed backup strategies

Content Structure Preservation

Hierarchical Organization

Maintaining educational content structure:

Course/Track Level
├── Module Organization
│   ├── Lesson Sequence
│   ├── Exercise Materials
│   └── Project Templates
└── Assessment Components

Metadata Preservation

Critical information to maintain:

  • Learning objectives and outcomes
  • Prerequisites and dependencies
  • Difficulty levels and progression
  • Completion tracking data
  • Original publication dates

Implementation Approaches

Dual-Source Architecture

Supporting both preserved content and new material:

  • File-based Storage: Exported materials as local files
  • Database Integration: Dynamic content management capabilities
  • Hybrid Rendering: Seamless presentation of both content types

Progressive Enhancement

  • Offline Accessibility: Content available without internet connection
  • Search Capabilities: Full-text search across preserved materials
  • Interactive Features: Maintaining quiz and exercise functionality
  • Progress Tracking: Personal learning progress independent of original platform

Fair Use and Personal Backup

  • Personal archival for continued learning
  • Non-commercial use of educational materials
  • Respect for intellectual property rights
  • Compliance with platform terms of service

Knowledge Sharing Ethics

  • Appropriate attribution of original sources
  • Sharing methodologies rather than copyrighted content
  • Contributing to educational tool development
  • Supporting open-source educational initiatives

Benefits for Learners

Continued Access Assurance

  • Permanent access to educational investments
  • Independence from platform business decisions
  • Ability to reference materials throughout career
  • Foundation for continued learning and skill development

Enhanced Learning Experience

  • Personalized organization and note-taking
  • Custom search and discovery features
  • Integration with personal knowledge management systems
  • Ability to supplement with additional resources

Portfolio Development Value

Educational content preservation projects demonstrate:

  • Technical Proficiency: Full-stack development capabilities
  • Problem-Solving Skills: Addressing real-world sustainability concerns
  • Project Management: Planning and executing complex data migration
  • User Experience Design: Creating intuitive educational interfaces

See also

Git Diff Tracking

page dédiée →

Version control pattern for detecting and processing only changed content in knowledge management systems. Eliminates the need for complex manifest systems by leveraging Git's native change detection capabilities.

Core Concept

Instead of building custom tracking mechanisms, use Git diff to identify exactly what content has changed since the last processing run. This approach provides reliable change detection with minimal overhead and perfect accuracy.

Implementation Pattern

Basic Change Detection

# Identify all changed files since last successful ingest
git diff HEAD~1 --name-only raw/projects/ | while read file; do
    echo "Processing changed file: $file"
    # Process only this specific file
done

Incremental Project Synchronization

# Sync projects from development environment
rsync -av --include='*.md' --exclude='node_modules' ~/code/*/ raw/projects/

# Git shows exactly what changed
git add raw/projects/
if ! git diff --cached --quiet; then
    echo "New content detected, processing changes..."
    git diff --cached --name-only | process_changed_files
    git commit -m "Sync projects: $(date)"
fi

Architecture Benefits

Eliminates Manifest Complexity

Traditional approaches require maintaining separate tracking files:

  • Timestamp-based systems prone to clock skew
  • Hash-based manifests requiring custom logic
  • Database tracking adding system complexity

Git diff provides all tracking functionality with zero custom code.

Atomic Processing

Git commits create natural processing boundaries:

  • Each commit represents one complete ingestion cycle
  • Failed processing can be retried from last successful commit
  • Processing history provides complete audit trail

Distributed Reliability

Git's distributed nature provides automatic backup and synchronization:

  • Processing history preserved across machines
  • Remote repository serves as authoritative state
  • Merge conflicts reveal concurrent processing issues

Advanced Patterns

Selective Processing by File Type

# Process only markdown changes
git diff HEAD~1 --name-only --diff-filter=AM | grep '\.md$' | process_files

# Handle deletions separately  
git diff HEAD~1 --name-only --diff-filter=D | cleanup_deleted_files

Content-Aware Change Detection

# Get actual content changes, not just file list
git diff HEAD~1 raw/projects/ | while read line; do
    if $line =~ ^\+\+\+; then
        # New file or significant addition
        process_file_addition
    elif $line =~ ^\---; then
        # Deleted content
        process_file_removal  
    fi
done

Automated Commit Integration

# Combine sync and change detection
sync_projects() {
    rsync -av ~/code/*/ raw/projects/
    
    if git add raw/projects/ && ! git diff --cached --quiet; then
        # Process only changed content
        git diff --cached --name-only | process_incremental_changes
        
        # Commit after successful processing
        git commit -m "Auto-sync: $(date) - $(git diff --cached --numstat | wc -l) files"
        
        echo "✓ Processed $(git diff HEAD~1 --numstat | wc -l) changed files"
    else
        echo "No changes detected"
    fi
}

Error Handling

Processing Failure Recovery

# If processing fails, reset to last known good state
if ! process_changes; then
    echo "Processing failed, resetting to last commit"
    git reset --hard HEAD
    exit 1
fi

Partial Success Handling

# Process files individually to isolate failures
git diff HEAD~1 --name-only | while read file; do
    if process_single_file "$file"; then
        git add "$file"
        echo "✓ Processed $file"
    else
        echo "✗ Failed to process $file"
        # Continue with other files
    fi
done

# Commit whatever succeeded
if ! git diff --cached --quiet; then
    git commit -m "Partial sync: $(date)"
fi

Integration Patterns

Daily Automation Integration

# Daily agent workflow
cd /path/to/knowledge-base

# Sync latest content
./scripts/collect-projects.sh

# Process only changes
if ! git diff --quiet raw/projects/; then
    echo "Changes detected, starting processing..."
    git add raw/projects/
    changed_files=$(git diff --cached --name-only | wc -l)
    
    # Process incrementally
    process_git_changes
    
    # Update wiki index
    update_index_with_changes
    
    git commit -m "Daily sync: $changed_files files updated"
    echo "✓ Processed $changed_files changed files"
else
    echo "No changes since last sync"
fi

Multi-Repository Coordination

# Track changes across multiple source repositories
for repo in ~/code/*/; do
    repo_name=$(basename "$repo")
    
    # Sync and track per-repository changes

GitHub Private Repo Architecture

page dédiée →

Pattern for hosting knowledge management systems using GitHub private repositories as the central storage and collaboration layer, enabling version control, automated processing, and multi-tool integration while maintaining privacy and control.

Architecture Benefits

Version Control for Knowledge

Git provides natural versioning for wiki content evolution:

  • Track changes to concepts and understanding over time
  • Rollback incorrect updates or information
  • Branch for experimental content organization
  • Merge knowledge from multiple sources systematically

Tool Ecosystem Integration

GitHub serves as a neutral platform accessible by multiple tools:

  • Local editing through Obsidian, VSCode, or any markdown editor
  • Automated processing through GitHub Actions
  • API access for programmatic updates via wiki agents
  • Web interface for manual review and curation

Privacy and Control

Private repositories maintain confidentiality while enabling automation:

  • Sensitive project information remains private
  • API tokens control access without exposing content
  • Self-hosted processing maintains data sovereignty
  • Export/migration remains straightforward through Git

Repository Structure Pattern

llm-wiki/
  SCHEMA.md              # System conventions and workflows
  index.md               # Master catalog of all content
  log.md                 # Chronological operation history
  BUILDING.md            # Meta-construction documentation
  raw/                   # Immutable source documents
    screenshots/         # Mobile-captured images
    articles/           # Web content and newsletters
    transcripts/        # Audio processing results
    projects/           # Extracted project documentation
    conversations/      # Agent interaction exports
    papers/             # Research materials
    pending/            # Medium-quality sources awaiting review
    discarded/          # Low-quality sources with rejection reasons
  wiki/                 # LLM-maintained structured content
    entities/           # People, tools, organizations
    concepts/           # Technical knowledge and patterns
    projects/           # Implementation documentation
    sources/            # Individual source summaries
    syntheses/          # Cross-source analysis
    skills/             # Mastered techniques
    flashcards/         # Spaced repetition content
    digests/            # Periodic summaries
  tools/                # Automation and processing scripts

Automation Integration

Git Hooks

Trigger processing on content changes:

  • Pre-commit hooks validate markdown structure
  • Post-commit hooks trigger wiki agent processing
  • Push hooks initiate automated backups

GitHub Actions

Enable cloud-based processing workflows:

  • Automated content quality checking
  • Scheduled digest generation
  • Integration testing for wiki structure
  • Deployment to static site generators

Local Processing

Balance privacy with cloud convenience:

  • Sensitive processing remains local
  • Public API calls for general knowledge enhancement
  • Hybrid workflows combining local and cloud resources

Multi-Tool Compatibility

Obsidian: Native markdown and wikilink support for interactive browsing and manual editing.

VSCode/Cursor: Full development environment access for complex editing and automation development.

MCP Servers: Programmatic access for LLM tools and agents through standardized protocols.

Mobile Apps: Git clients enable basic editing and review from mobile devices.

Implementation Considerations

Authentication: SSH keys or personal access tokens for automated access without exposing credentials.

Backup Strategy: Git provides inherent backup, but consider additional external backups for critical knowledge bases.

Search Integration: GitHub's search capabilities combined with local indexing provide comprehensive content discovery.

Collaboration: Private repos support controlled sharing with team members while maintaining primary ownership.

See also

  • ai-engineering-wiki - Implementation example using this pattern
  • llm-wiki-pattern - Knowledge management approach
  • version-control-for-knowledge - Git workflows for content management
  • obsidian-github-integration - Tool-specific implementation patterns

Human-LLM Division of Labor

page dédiée →

The strategic allocation of responsibilities between humans and LLMs in llm-wiki-pattern systems, optimizing each party's strengths while solving the traditional wiki abandonment problem. Core insight: humans excel at curation and synthesis; LLMs excel at maintenance and bookkeeping.

Task Allocation

Human Responsibilities

  • Source curation: Selecting valuable documents to ingest
  • Direction setting: Guiding analysis emphasis and priorities
  • Question asking: Driving exploration through strategic queries
  • Meaning synthesis: Understanding implications and significance
  • Quality oversight: Reviewing summaries and checking updates
  • Schema evolution: Adapting system configuration based on needs

LLM Responsibilities

  • Content maintenance: Updating cross-references, keeping summaries current
  • Consistency management: Noting contradictions, maintaining coherence
  • Bookkeeping automation: Filing, indexing, logging operations
  • Cross-referencing: Building and maintaining wikilinks networks
  • Structure creation: Generating new pages, organizing content
  • Workflow execution: Following schema-defined operational procedures

Solving the Abandonment Problem

Traditional Wiki Failure Pattern

  • Initial enthusiasm: Humans start with high motivation
  • Growing burden: Maintenance tasks accumulate faster than value
  • Cognitive overhead: Cross-referencing becomes mentally taxing
  • Inevitable abandonment: Effort required exceeds perceived benefit

LLM Solution

  • Zero maintenance fatigue: LLMs don't experience tedium or boredom
  • Consistent execution: Never forget to update cross-references
  • Parallel processing: Can touch 15 files in one pass without cognitive load
  • Near-zero cost: Maintenance burden becomes negligible

Cognitive Complementarity

Human Cognitive Strengths

  • Contextual judgment: Understanding significance and relevance
  • Creative synthesis: Making novel connections and insights
  • Domain expertise: Applying specialized knowledge and intuition
  • Strategic thinking: Long-term planning and goal-oriented exploration

LLM Cognitive Strengths

  • Systematic processing: Consistent application of rules and procedures
  • Pattern recognition: Identifying structural relationships across content
  • Parallel attention: Managing multiple interconnected updates simultaneously
  • Infinite patience: Performing repetitive tasks without degradation

Operational Boundaries

Human Decision Points

  • Source selection: What documents deserve ingestion?
  • Emphasis guidance: What aspects need highlighting?
  • Quality gates: Are summaries accurate and useful?
  • Exploration direction: What questions should drive further investigation?

LLM Execution Points

  • Content integration: How to incorporate new information?
  • Link maintenance: Which pages need cross-reference updates?
  • Consistency checks: Where do contradictions need flagging?
  • Structure organization: How to categorize and file content?

Interface Design

Human-LLM Interaction Model

  • LLM agent open on one side of screen
  • Obsidian (or wiki browser) open on other side
  • Real-time collaboration: Human monitors LLM edits live
  • Immediate feedback: Human can guide and correct during operation

Communication Patterns

  • Explicit instructions: Human provides clear direction for emphasis
  • Progress reporting: LLM describes what updates are being made
  • Quality confirmation: Human reviews and approves significant changes
  • Schema discussion: Collaborative evolution of system configuration

Benefits of Clear Division

Efficiency Optimization

  • Leverage strengths: Each party focuses on optimal tasks
  • Minimize waste: Avoid humans doing tedious work, LLMs making judgment calls
  • Sustainable workflow: Maintenance burden doesn't grow with scale

Quality Assurance

  • Human oversight: Strategic decisions remain under human control
  • LLM consistency: Mechanical tasks executed reliably
  • Complementary validation: Different cognitive approaches catch different errors

Implementation Considerations

Trust Building

  • Gradual automation: Start with supervised workflows, increase autonomy
  • Transparency: LLM reports all changes and reasoning
  • Reversibility: Git versioning enables rollback of problematic updates

Workflow Evolution

  • Usage-driven refinement: Division of labor adapts based on experience
  • Domain customization: Different fields may require different task allocations
  • Tool integration: Technical capabilities influence responsibility boundaries

See also

Intelligent Content Triage

page dédiée →

Automated system for evaluating and routing incoming content based on relevance, quality, and potential value to knowledge base. Essential component of scalable knowledge management systems to prevent information overload while ensuring valuable content is captured and processed.

Triage Framework

Three-Tier Scoring System

HIGH (Auto-ingest):

  • Introduces concepts/techniques not yet in wiki
  • Contradicts or significantly nuances existing content
  • Directly relates to active projects
  • From high-quality sources (papers, official docs, recognized authors)
  • Contains actionable code patterns or architecture decisions

MEDIUM (Manual review queue):

  • Potentially relevant but uncertain value
  • Adjacent to current domains but not core
  • Interesting but shallow content that might need deeper sourcing
  • Routed to raw/pending/ folder for later evaluation

LOW (Archive/discard):

  • Not related to domains of interest
  • Too superficial (3-line social media posts without substance)
  • Duplicate of already-ingested content
  • Promotional content disguised as technical content
  • Outdated information already superseded in wiki

Implementation Architecture

Screenshot Filtering Pipeline

For iPhone social media screenshots - most common source requiring filtering:

  1. Quick Domain Check: "Is this tech/AI related?" - fast binary filter
  2. Content Analysis: OCR + vision model text extraction
  3. Relevance Scoring: Compare against existing wiki index
  4. Quality Assessment: Source credibility, depth, actionability
  5. Routing Decision: HIGH → ingest, MEDIUM → pending, LOW → discard

Automated Quality Gates

Source Credibility Factors:

  • Author reputation in AI/tech domains
  • Publication venue quality (arXiv, official docs, conferences)
  • Content depth and technical specificity
  • Presence of code examples or concrete implementations

Relevance Criteria:

  • Alignment with current project portfolio
  • Overlap assessment with existing wiki content
  • Introduction of genuinely new concepts vs. rehashed basics
  • Potential for compound learning and cross-referencing

Compound Learning Effect

The triage system improves over time through knowledge accumulation:

  • Basic Transformer articles scored LOW when wiki already contains 15 related pages
  • Novel techniques (Flash Attention, MoE architectures) scored HIGH
  • Scoring accuracy increases as wiki index grows more comprehensive
  • Domain understanding deepens, enabling better quality assessment

Real-World Performance

Live Implementation Results (June 2026):

  • Karpathy's LLM wiki document: HIGH score (correctly identified as authoritative, novel)
  • Automatic routing prevented manual review overhead
  • Generated 5 wiki pages + 5 flashcards from single high-quality source
  • Zero false positives in initial testing phase

Integration Patterns

Telegram Bot Integration

Direct scoring from mobile messaging:

  • Send URL: bot fetches, scores, routes appropriately
  • Send screenshot: vision model extracts text, scores content
  • Real-time feedback on triage decisions
  • Manual override capability for edge cases

Multi-Source Adaptation

Different scoring criteria per source type:

  • Academic papers: Novelty + technical depth
  • Project documentation: Direct applicability + completeness
  • Social media: Signal-to-noise ratio + source authority
  • Audio transcripts: Actionable insights + speaker expertise

Technical Implementation

LLM Prompt Structure:

# Compare against existing wiki index
# Apply domain relevance criteria from SCHEMA.md
# Score using established quality gates
# Return: HIGH/MEDIUM/LOW + reasoning

File System Routing:

  • raw/sources/ → HIGH content for immediate processing
  • raw/pending/ → MEDIUM content awaiting manual review
  • raw/discarded/ → LOW content archived with reasoning

This systematic approach enables high-volume content processing while maintaining knowledge base quality, essential for ai-engineering-wiki automation at scale.

See also

Knowledge Management Systems

page dédiée →

Systems designed to capture, organize, and make accessible the knowledge and expertise within an organization or for personal use. Range from simple note-taking tools to sophisticated AI-powered knowledge bases.

Evolution of Approaches

Traditional Approaches

  • File systems: Folders and documents
  • Wikis: Collaborative, hyperlinked documentation
  • Databases: Structured information storage
  • Content management: Document-centric systems

Modern AI-Enhanced Approaches

  • RAG systems: Retrieve relevant chunks at query time
  • llm-wiki-pattern: LLM-maintained persistent knowledge bases
  • Semantic search: Vector-based information retrieval
  • AI summarization: Automated content distillation

Key Challenges

The Maintenance Problem

Traditional knowledge management systems often fail because:

  1. High maintenance burden: Keeping information current requires constant human effort
  2. Cross-referencing overhead: Manually maintaining links between related concepts
  3. Inconsistency growth: Different pages contradict each other over time
  4. Abandonment: Systems become stale as maintenance burden exceeds value

The Discovery Problem

  • Finding relevant information in large knowledge bases
  • Understanding connections between concepts
  • Surfacing insights that span multiple sources

LLM-Enhanced Solutions

Automated Maintenance

LLMs can handle the tedious aspects of knowledge management:

  • Updating cross-references automatically
  • Identifying and flagging contradictions
  • Maintaining consistent formatting and structure
  • Generating summaries and overviews

Intelligent Organization

  • Automatic categorization and tagging
  • Dynamic relationship discovery
  • Content synthesis across sources
  • Gap identification and suggestion of new sources

Types of Knowledge Management

Personal Knowledge Management

  • Research notes and literature reviews
  • Learning from courses and books
  • Project documentation and lessons learned
  • Health, goals, and self-improvement tracking

Organizational Knowledge Management

  • Team documentation and best practices
  • Customer interaction histories
  • Technical knowledge bases
  • Institutional memory preservation

Tools and Platforms

Traditional Tools

  • Obsidian: Graph-based personal knowledge management
  • Notion: Block-based collaborative workspaces
  • Roam Research: Bi-directional linking for research
  • MediaWiki: Traditional collaborative wiki platform

AI-Enhanced Tools

  • NotebookLM: Google's AI-powered research assistant
  • ChatGPT with files: Document upload and querying
  • Custom LLM wiki systems: Using the LLM Wiki Pattern

Design Principles

For Sustainable Systems

  1. Low maintenance overhead: Automate bookkeeping tasks
  2. Clear human-AI division: Humans curate and direct; AI maintains
  3. Incremental value: Each addition makes the system more valuable
  4. Consistent structure: Standardized formats and conventions
  5. Version control: Track changes and maintain provenance

For Effective Discovery

  1. Multiple access patterns: Browse, search, and follow links
  2. Rich cross-referencing: Connect related concepts
  3. Hierarchical organization: Categories and subcategories
  4. Temporal tracking: Understand how knowledge evolved

Success Factors

  1. Clear purpose: Focused domain and use cases
  2. Consistent workflow: Standardized processes for adding content
  3. Regular use: Systems that aren't used regularly decay
  4. Appropriate tooling: Match tools to user preferences and technical skills
  5. Sustainable maintenance: Either automated or minimal manual overhead

See also

LLM Wiki Pattern

page dédiée →

A paradigm for building personal knowledge bases where LLMs incrementally build and maintain persistent wikis rather than retrieving from raw documents at query time. Developed by andrej-karpathy as an alternative to traditional RAG systems that rediscover knowledge from scratch on every interaction.

Core Innovation

Persistent vs Ephemeral Knowledge: Most people's experience with LLMs and documents follows the RAG pattern - upload files, retrieve relevant chunks at query time, generate answers. This works but requires rediscovering knowledge from scratch on every question. The LLM Wiki Pattern instead creates persistent, compounding artifacts where knowledge is compiled once and kept current.

Key Difference: The wiki sits between you and raw sources. When adding new sources, the LLM doesn't just index for later retrieval - it reads, extracts key information, and integrates into existing wiki structure. Updates entity pages, revises summaries, flags contradictions, strengthens synthesis. Cross-references already exist. Contradictions already flagged. Synthesis already reflects everything read.

Three-Layer Architecture

  1. Raw Sources: Immutable curated collection (articles, papers, images, data files). LLM reads but never modifies. Source of truth.

  2. The Wiki: Directory of LLM-generated markdown files. Summaries, entity pages, concept pages, comparisons, synthesis. LLM owns this layer entirely - creates, updates, maintains cross-references, ensures consistency.

  3. The Schema: Configuration document (CLAUDE.md, AGENTS.md) defining wiki structure, conventions, workflows. Co-evolved between human and LLM over time as patterns emerge.

Core Operations

Ingest Workflow

Drop new source into raw collection. LLM reads source, discusses takeaways, writes summary page, updates index, updates relevant entity/concept pages across wiki, appends log entry. Single source might touch 10-15 wiki pages. Can be done one-at-a-time with supervision or batch-processed.

Query Workflow

Ask questions against wiki. LLM searches relevant pages, reads them, synthesizes answer with citations. Answers can take multiple forms - markdown pages, comparison tables, slide decks (marp-integration), charts, canvas. Critical insight: Good answers get filed back as new wiki pages. Explorations compound in knowledge base like ingested sources.

Lint Workflow

Periodic health-checking. Look for contradictions between pages, stale claims superseded by newer sources, orphan pages with no inbound links, important concepts mentioned but lacking pages, missing cross-references, data gaps. LLM suggests new questions to investigate and sources to find.

Index and Logging

  • index.md: Content-oriented catalog of all wiki pages with links, summaries, metadata. Organized by category. Updated on every ingest. LLM reads index first to find relevant pages for queries. Works well at moderate scale (~100 sources, hundreds of pages).

  • log.md: Chronological append-only record of ingests, queries, lint passes. Parseable with consistent prefixes (## [2026-04-02] ingest | Article Title). Provides timeline of wiki evolution.

Optional Tooling

qmd-search engine recommended for scaling beyond index-based navigation. Local search for markdown with hybrid BM25/vector search and LLM re-ranking. Available as CLI tool and MCP server.

Implementation Recommendations

Obsidian Integration

Recommended interface: obsidian-integration where human browses in Obsidian while LLM makes live edits to markdown files. "Obsidian is the IDE; the LLM is the programmer; the wiki is the codebase."

Useful Obsidian Features:

  • Web Clipper: Browser extension converting articles to markdown
  • Image handling: Download attachments locally, bind to hotkey (Ctrl+Shift+D)
  • Graph view: Visualize wiki connections, identify hubs and orphans
  • marp-integration: Generate slide decks from wiki content
  • dataview-plugin: Query page frontmatter for dynamic tables/lists
  • Git integration: Version history, branching, collaboration

Use Cases

Personal: Goals, health, psychology tracking. File journal entries, articles, podcast notes into structured self-picture over time.

Research: Deep topic exploration over weeks/months. Build comprehensive wiki with evolving thesis.

Reading Companion: File chapters as you go. Build pages for characters, themes, plot threads. Like fan wikis (Tolkien Gateway) but personal with LLM maintenance.

Business/Team: Internal wiki fed by Slack threads, meeting transcripts, project docs, customer calls. Humans review updates.

Other: Competitive analysis, due diligence, trip planning, course notes, hobby deep-dives.

Historical Foundation

Explicitly references vannevar-bush's memex-vision (1945) as spiritual predecessor. Bush envisioned personal, curated knowledge store with associative-trails between documents. His vision was private, actively curated, with connections between documents as valuable as documents themselves.

Key Insight: Bush couldn't solve the maintenance problem. LLMs handle that through bookkeeping-automation. Humans abandon wikis because maintenance burden grows faster than value. LLMs don't get bored, don't forget cross-references, can touch 15 files in one pass. Maintenance cost approaches zero.

Division of Labor

Human Role: Curate sources, direct analysis, ask good questions, think about meaning.

LLM Role: Summarizing, cross-referencing, filing, bookkeeping that makes knowledge base useful over time.

Design Philosophy

Intentionally abstract specification describing the pattern, not specific implementation. Directory structure, schema conventions, page formats, tooling all depend on domain, preferences, LLM choice. Everything modular and optional - pick what's useful. Designed to be shared with LLM agents for collaborative instantiation.

See also

Memex Vision

page dédiée →

vannevar-bush's 1945 vision of a personal knowledge management system that would allow individuals to store, retrieve, and create associative-trails between documents and information. A foundational concept for modern knowledge management and the direct inspiration behind andrej-karpathy's llm-wiki-pattern and other AI-powered information systems.

Historical Context

Introduced in Bush's essay "As We May Think" (1945), the Memex was conceived as a mechanical device that would supplement human memory by allowing rapid consultation of personal collections of documents. Bush envisioned a desk-sized device with screens, keyboards, and mechanical storage systems.

Core Vision

Private curation - Personal knowledge stores actively maintained by individuals rather than shared databases

Associative trails - Connections between documents as valuable as the documents themselves. Users could create and follow paths linking related information across their collection.

Mechanical augmentation - Technology amplifying human cognitive capabilities rather than replacing human judgment

Active maintenance - Knowledge bases requiring continuous organization and cross-referencing to remain valuable

The Maintenance Problem

Bush's vision was prescient but incomplete. He understood the value of connected, curated knowledge but couldn't solve the fundamental challenge: who does the maintenance work?

Creating associative trails, maintaining cross-references, keeping summaries current, noting contradictions - this bookkeeping is essential but tedious. Humans consistently abandon personal knowledge systems because maintenance burden grows faster than perceived value.

Modern Resolution

The llm-wiki-pattern directly addresses Bush's unsolved maintenance problem. LLMs excel at the systematic bookkeeping that humans find burdensome:

  • Updating cross-references across multiple pages
  • Maintaining consistency as information accumulates
  • Noting contradictions between sources
  • Creating and strengthening associative connections

This allows Bush's original vision to be realized: private, actively curated knowledge stores where connections between documents are as valuable as documents themselves.

Influence on Contemporary Systems

The Memex vision influenced:

  • Hypertext systems and the World Wide Web
  • Personal knowledge management tools (Roam, Obsidian, Notion)
  • Information retrieval research
  • Modern AI-powered knowledge systems

However, most implementations lost Bush's emphasis on private curation and became either:

  • Public databases (Wikipedia, web)
  • Simple storage without maintained associations (file systems)
  • Tools requiring manual maintenance (personal wikis that get abandoned)

Key Insights for LLM Wiki Pattern

Connection value - The relationships between pieces of information often more valuable than isolated documents

Curation over scale - Personal, curated collections outperform generic large databases for individual needs

Maintenance as bottleneck - Technical capability less important than solving the maintenance burden problem

Human-machine collaboration - Best results combine human judgment (curation, direction) with machine capability (systematic maintenance)

See also

Multi-Source Ingestion

page dédiée →

Architecture pattern for knowledge management systems that automatically captures and processes content from diverse sources and formats. Essential for comprehensive knowledge accumulation without manual overhead, enabling systematic learning from all information channels.

Core Architecture

Source Categories

1. Development Projects

  • Automated sync from ~/code/*/ using rsync
  • Intelligent filtering (exclude node_modules, .git, build artifacts)
  • Git diff tracking to process only changed files
  • Captures READMEs, documentation, and learning artifacts

2. Visual Content

  • iOS Shortcuts integration for frictionless screenshot capture
  • iCloud Drive sync for automatic mobile-to-desktop transfer
  • LLM vision for OCR and semantic classification
  • Context-aware processing of social media content

3. Audio Content

  • Conference talks and meeting recordings
  • Voice memos and interview transcripts
  • Whisper-based transcription with MLX optimization
  • Automatic speaker identification and topic segmentation

4. Web Content

  • RSS feed aggregation from key sources
  • Manual article addition with URL parsing
  • Research paper ingestion from arXiv and academic sources
  • Blog posts and technical documentation

5. Conversational Content

  • AI agent conversations from Claude Code, Cursor IDE
  • Chat transcripts with technical discussions
  • Code review conversations and architectural decisions

Processing Pipeline

Raw Sources → Triage → Classification → Extraction → Integration
     ↓           ↓          ↓           ↓          ↓
  Diverse     Quality   Content-Type   Key Info   Wiki Pages
  Formats     Filter    Detection      Capture    & References

Triage System

Automated Quality Assessment:

  • Content length and depth analysis
  • Source credibility scoring
  • Relevance to existing knowledge base
  • Novelty detection against existing pages

Classification Outcomes:

  • High: Immediate processing and integration
  • Medium: Queue for manual review in raw/pending/
  • Low: Archive to raw/discarded/ with reasoning

Content Type Detection

Intelligent Format Handling:

  • Markdown parsing and structure analysis
  • Image OCR with context understanding
  • Audio transcription with speaker diarization
  • Code extraction and documentation linking

Implementation Patterns

Git-Based Change Detection

Revolutionary approach using git-diff-tracking to eliminate manifest complexity:

# Detect changes since last ingest
git diff --name-status HEAD~1 raw/projects/
# Process only modified files
find raw/ -newer .last_ingest -type f

iOS Shortcuts Integration

Seamless mobile capture through ios-shortcuts-integration:

iPhone Screenshot → Share → "Brain Wiki" Shortcut → iCloud Drive → Processing Queue

Batch Processing

Efficient handling of large content volumes:

  • Process screenshots in batches for OCR efficiency
  • Parallel transcription of multiple audio files
  • Async classification of web articles
  • Incremental project synchronization

Cross-Modal Synthesis

Automated Cross-Referencing

  • Screenshot insights link to related projects
  • Conference talks reference technical documentation
  • Project learnings connect to research papers
  • Agent conversations enhance concept pages

Compound Learning Effects

Multi-source integration creates knowledge compounds:

  • Visual content + code examples + academic papers = comprehensive understanding
  • Social media insights + project experience + formal documentation = practical wisdom

Real-World Implementation

Epic Brain Wiki Results

From the complete brain-wiki implementation:

Sources Processed:

  • 31 AI/ML projects (4,060 markdown files)
  • Screenshot pipeline with iOS Shortcuts integration
  • Audio transcription using MLX Whisper
  • Web content through RSS and manual addition
  • Agent conversations from Claude Code sessions

Automation Achieved:

  • Zero manual overhead for routine ingestion
  • Daily orchestration processing all sources
  • Intelligent filtering eliminating 90% of noise
  • Cross-modal integration creating compound insights

Measured Benefits

  • 10x faster knowledge capture compared to manual curation
  • 5x more cross-references discovered through automated analysis
  • 90% reduction in manual content processing time
  • Continuous operation without human intervention

Technical Implementation

Directory Structure

multi-source-system/
├── raw/                    # Immutable source storage
│   ├── projects/          # Development project files
│   ├── screenshots/       # Visual content capture
│   ├── talks/            # Audio transcriptions
│   ├── articles/         # Web content archive
│   ├── conversations/    # Agent chat logs
│   ├── pending/          # Medium-quality sources
│   └── discarded/        # Low-quality archive
├── scripts/
│   ├── collect-projects.sh    # Project synchronization
│   ├── process-screenshots.sh # Image handling
│   ├── process-recordings.sh  # Audio transcription
│   └── daily-ingest.md       # Orchestration guide
└── processed/             # Integrated wiki content

Automation Scripts

Project Collection:

# Sync all code projects
rsync -av --exclude="node_modules" --exclude=".git" \
      ~/code/ raw/projects/

Screenshot Processing:

# Move from iCloud to processing queue
mv ~/Library/Mobile\ Documents/com~apple~CloudDocs/brain-wiki-inbox/* \
   raw/screenshots/inbox/

Audio Transcription:

# Whisper transcription
mlx_whisper audio_file.mp3 --output-format txt

Quality Control

Automated Validation

  • Minimum content length thresholds
  • Duplicate detection across sources
  • Link validation and reference checking
  • Format consistency verification

Human Oversight

  • Weekly review of medium-quality sources
  • Manual curation of cross-references
  • System performance monitoring
  • Pipeline optimization decisions

Scaling Considerations

Performance Optimization

  • Incremental processing to handle growth
  • Batch operations for efficiency
  • Resource pooling for concurrent tasks
  • Cache layers for repeated operations

Storage Management

  • Compression for archived content
  • Automated cleanup of outdated sources
  • Backup strategies for critical content
  • Version control for all processed data

Advanced Features

Content Enrichment

  • Automatic tagging based on content analysis
  • Entity extraction and relationship mapping
  • Topic modeling for content clustering
  • Sentiment analysis for conversational content

Adaptive Processing

  • Learning from user feedback on content quality
  • Dynamic threshold adjustment for triage
  • Personalized relevance scoring
  • Context-aware classification improvement

Implementation Challenges

Content Quality Variability

  • Handling low-signal sources (screenshots of memes)
  • Dealing with incomplete or corrupted files
  • Managing different content formats and structures
  • Balancing automation with quality control

Technical Complexity

  • Coordinating multiple processing pipelines
  • Error handling across diverse input types
  • Resource management for intensive operations
  • Maintaining system reliability and uptime

Privacy and Security

  • Sensitive content identification and handling
  • Access control for different source types
  • Secure storage of personal information
  • Compliance with data protection requirements

See Also

RAG Alternative Approaches

page dédiée →

Methods for working with large document collections that go beyond traditional Retrieval-Augmented Generation patterns. Most notably exemplified by andrej-karpathy's llm-wiki-pattern, which creates persistent, pre-synthesized knowledge structures rather than retrieving raw chunks at query time.

Traditional RAG Limitations

Rediscovery Problem: RAG systems retrieve relevant chunks and generate answers from scratch on every query. No knowledge accumulates—subtle questions requiring synthesis across multiple documents must reconstruct understanding each time.

Fragment-Based Understanding: Working with retrieved chunks rather than integrated knowledge, limiting the ability to develop sophisticated cross-source synthesis.

Stateless Operation: Each query starts fresh with no building upon previous insights or discoveries.

Pre-Synthesis Approach

The llm-wiki-pattern represents a fundamental alternative where:

Knowledge Integration: Instead of retrieving fragments, LLMs incrementally build and maintain structured knowledge representations that integrate information across sources.

Persistent Artifacts: Cross-references, contradictions, and syntheses are identified once and maintained rather than rediscovered on each query.

Compounding Intelligence: Good questions and discoveries get filed back into the knowledge base, creating compounding-artifacts that become smarter over time.

Architectural Differences

Traditional RAG

Sources → Embedding → Vector DB → Retrieval → LLM Generation

Wiki Pattern Alternative

Sources → LLM Integration → Persistent Wiki → Query → Pre-Synthesized Answers

Key Advantages

Deep Synthesis: Can develop sophisticated understanding that builds across multiple sources and interactions.

Context Preservation: Maintains full context of how understanding developed rather than working with isolated fragments.

Knowledge Evolution: The system gets smarter over time rather than remaining static.

Rich Cross-Referencing: Connections between concepts are discovered and maintained automatically.

Implementation Patterns

Incremental Integration

New sources update existing entity pages, revise concept summaries, and note contradictions rather than simply being added to a retrieval corpus.

Pre-Compiled Knowledge

Answers draw from knowledge that has already been synthesized and cross-referenced rather than being generated from raw retrieval.

Active Maintenance

Systems actively identify gaps, contradictions, and opportunities for deeper synthesis rather than passively serving queries.

Use Cases

Particularly effective for:

  • Long-term research projects requiring deep synthesis
  • Personal knowledge development over months/years
  • Complex domains where understanding evolves with exposure
  • Business intelligence requiring integrated analysis
  • Any context where knowledge should compound rather than remain fragmented

Trade-offs

Computational Investment: Requires upfront processing to build integrated knowledge structures rather than simple indexing.

Maintenance Complexity: Systems must actively maintain consistency and handle contradictions rather than serving static retrievals.

Domain Specificity: Requires careful schema design and workflow definition for specific knowledge domains.

Future Directions

The success of wiki-pattern approaches suggests broader possibilities for AI systems that build persistent, evolving knowledge representations rather than operating in stateless fashion.

See also

RAG Alternatives

page dédiée →

Approaches to knowledge management and information retrieval that move beyond traditional Retrieval-Augmented Generation (RAG) systems. Most notably exemplified by andrej-karpathy's llm-wiki-pattern which treats knowledge as compounding-artifacts rather than static document collections.

Traditional RAG Limitations

Standard RAG Approach:

  • Upload collection of files
  • LLM retrieves relevant chunks at query time
  • Generates answers from fragments
  • Rediscovers knowledge from scratch on every question
  • No accumulation or synthesis between queries

Problems:

  • Subtle questions requiring synthesis across multiple documents must be solved repeatedly
  • No building up of understanding over time
  • Connections between sources not maintained
  • Context limited to what can be retrieved in single query

LLM Wiki Pattern Alternative

Core Difference: Instead of retrieving from raw documents at query time, the LLM incrementally builds and maintains a persistent wiki that sits between user and sources.

Process:

  • New sources integrate into existing wiki structure
  • Updates entity pages, revises summaries, notes contradictions
  • Cross-references already established
  • Synthesis reflects all previous learning
  • Knowledge compounds with each addition

Key Advantages

Persistent Knowledge: Cross-references exist, contradictions flagged, synthesis current. Wiki keeps getting richer with every source and question.

Maintenance Automation: LLMs handle tedious bookkeeping - updating cross-references, keeping summaries current, maintaining consistency. Humans focus on curation and questions.

Compound Learning: Good answers become new wiki pages. Explorations compound in knowledge base just like ingested sources.

Other Alternative Approaches

Knowledge Graphs: Structured representation of entities and relationships, but typically requires significant manual curation or complex extraction pipelines.

Embedding-Based Memory Systems: Vector representations of experiences/documents that can be retrieved by similarity, but lack explicit structure and cross-referencing.

Agent Memory Architectures: Various approaches to giving AI systems persistent memory, though most focus on conversation history rather than structured knowledge building.

Implementation Considerations

Moving beyond RAG requires:

  • Schema Design: Clear workflows and conventions for knowledge maintenance
  • Integration Workflows: Systematic processes for incorporating new information
  • Quality Control: Methods for ensuring accuracy and consistency
  • Navigation Systems: Tools for exploring and searching the knowledge base

Applications

RAG alternatives particularly valuable for:

  • Research Synthesis: Building understanding across many sources over time
  • Personal Learning: Accumulating knowledge in specific domains
  • Business Intelligence: Maintaining current understanding of competitive landscape
  • Domain Expertise: Building deep, interconnected knowledge in specialized areas

See also

Retrieval Augmented Generation (RAG)

page dédiée →

A technique that enhances language model responses by retrieving relevant information from external knowledge bases at query time. The retrieved context is then used to ground the model's generation, reducing hallucinations and enabling access to information beyond the training data.

Core Process

  1. Query Processing: User question is converted to embeddings
  2. Retrieval: Similar document chunks are found using vector search
  3. Context Formation: Retrieved chunks are assembled as context
  4. Generation: LLM generates response using retrieved context
  5. Response: Final answer combines retrieved facts with model knowledge

Architecture Components

Vector Store

  • Document embeddings stored in vector database (chroma, pinecone, weaviate)
  • Enables semantic similarity search
  • Supports hybrid search combining dense and sparse retrieval

Embedding Models

  • Convert text to high-dimensional vectors
  • Popular options: OpenAI Ada-002, sentence-transformers, Cohere
  • Critical for retrieval quality

Chunking Strategy

  • Documents split into manageable segments
  • Balance between context preservation and retrieval precision
  • Common approaches: fixed-size, semantic, hierarchical

Advanced Patterns

Multi-Query RAG

Generate multiple query variants to improve retrieval coverage:

def multi_query_rag(question):
    variants = llm.generate_query_variants(question)
    all_chunks = []
    for variant in variants:
        chunks = vector_store.search(variant)
        all_chunks.extend(chunks)
    return llm.generate(question, dedupe(all_chunks))

rag-fusion

Combines multiple retrieval strategies and reranks results for improved relevance.

conversational-rag

Maintains conversation history to enable multi-turn interactions with context awareness.

agentic-rag

Uses AI agents to orchestrate complex retrieval workflows, including tool use and multi-step reasoning.

Limitations

Knowledge Rediscovery Problem

RAG systems rediscover knowledge from scratch on every query. As andrej-karpathy notes in the llm-wiki-pattern, this prevents knowledge accumulation - subtle questions requiring synthesis across multiple documents must repeatedly piece together fragments without building persistent understanding.

Context Window Constraints

  • Limited by model's maximum context length
  • Must balance breadth vs. depth of retrieved information
  • May miss relevant information if not in top-k results

Retrieval Quality Issues

  • Embedding similarity doesn't always match semantic relevance
  • Struggles with concepts spanning multiple documents
  • May retrieve contradictory information without resolution

Alternative Approaches

llm-wiki-pattern

Instead of retrieving from raw documents, LLMs maintain persistent wikis that incrementally integrate knowledge over time. This creates compounding-artifacts where cross-references exist permanently and contradictions are pre-resolved.

knowledge-graphs

Structured representation of entities and relationships can provide more precise retrieval paths than vector similarity.

Use Cases

  • Document Q&A: Customer support, internal documentation
  • Research Assistance: Scientific literature review, fact-checking
  • Content Generation: Blog writing with factual grounding
  • Educational Tools: Personalized tutoring with curriculum materials

Implementation Considerations

Evaluation Metrics

  • Retrieval Accuracy: Precision/recall of relevant documents
  • Answer Quality: Factual correctness, completeness, relevance
  • Latency: End-to-end response time
  • Cost: Embedding generation and storage costs

Production Challenges

  • Keeping knowledge base current with new information
  • Handling contradictory or outdated information
  • Scaling retrieval performance with growing document corpus
  • Managing embedding model updates and reindexing

See also

Schema Coevolution

page dédiée →

The collaborative development process between human and LLM for evolving the configuration layer of llm-wiki-pattern systems. The schema document (CLAUDE.md, AGENTS.md, etc.) grows and adapts based on actual usage patterns, domain needs, and discovered workflows.

Core Concept

Unlike static system configurations, the schema in an LLM wiki system evolves with use. As humans and LLMs work together, they discover:

  • Effective workflows for the specific domain
  • Useful page formats and structures
  • Domain-specific conventions and tags
  • Optimal ingest and query patterns

These discoveries get documented in the schema, making the LLM a more effective collaborator over time.

Evolution Process

Initial Schema

  • Basic structure and conventions
  • Generic workflows from the pattern
  • Minimal domain-specific customization

Usage-Driven Refinement

  • Workflow optimization: Discovering efficient ingest patterns
  • Convention standardization: Establishing consistent tagging and formatting
  • Domain adaptation: Adding field-specific page types and structures
  • Tool integration: Incorporating discovered utilities and scripts

Continuous Improvement

  • Regular schema updates based on what works
  • Documentation of effective practices
  • Removal of unused conventions
  • Addition of new capabilities

Human-LLM Collaboration

Human Contributions

  • Domain expertise: Understanding field-specific needs
  • Workflow preferences: Preferred interaction patterns
  • Quality standards: Defining acceptable outputs
  • Strategic direction: Long-term knowledge goals

LLM Contributions

  • Pattern recognition: Identifying recurring structures
  • Consistency maintenance: Ensuring schema adherence
  • Workflow execution: Following documented procedures
  • Improvement suggestions: Proposing optimizations

Schema Components That Evolve

Page Types

  • Standard templates for different content types
  • Domain-specific entity categories
  • Specialized analysis formats (comparisons, timelines, etc.)

Tagging Systems

  • Hierarchical tag structures for the domain
  • Consistent naming conventions
  • Cross-reference patterns

Workflows

  • Detailed ingest procedures
  • Query and synthesis patterns
  • Maintenance and lint operations
  • Quality control checkpoints

Tool Integration

  • CLI utilities and their usage patterns
  • Search and navigation tools
  • Export and visualization formats
  • Integration with external systems

Benefits

Adaptive Systems

Schema coevolution enables knowledge systems that improve with use rather than becoming stale or rigid.

Domain Optimization

Over time, the system becomes specifically tuned to the user's field and preferences rather than remaining generic.

Reduced Friction

Well-evolved schemas reduce cognitive overhead by codifying effective practices and eliminating decision fatigue.

Knowledge Transfer

The schema serves as documentation of effective knowledge management practices that can be shared or adapted.

Example Evolution

Initial: "Create entity pages for people mentioned"
Evolved: "Create entity pages with standardized sections: Background, Key Ideas, Publications, Influence, Cross-References. Tag with domain (ai-researcher, entrepreneur, academic) and confidence level."

Challenges

Over-Specification

Schemas can become too rigid, constraining useful variation and experimentation.

Version Management

Managing schema changes while maintaining consistency across existing wiki content.

Complexity Growth

Balancing comprehensive documentation with usability for both human and LLM.

Implementation Patterns

Versioned Schemas

Track schema evolution with version control to understand what changes improve effectiveness.

Modular Structure

Organize schema into sections (workflows, conventions, formats) that can evolve independently.

Example-Driven Documentation

Include concrete examples in schema documentation to clarify abstract conventions.

See also

Wiki Indexing Patterns

page dédiée →

Systematic approaches to organizing and navigating knowledge bases, particularly in llm-wiki-pattern implementations. Two primary patterns serve different navigation needs: content-oriented catalogs and chronological logs.

Core Patterns

Content-Oriented Index (index.md)

Purpose: Catalog all wiki content organized by category and topic.

Structure:

  • Links to all wiki pages with one-line summaries
  • Organized by category (entities, concepts, sources, syntheses)
  • Optional metadata (dates, source counts, confidence levels)
  • Updated on every ingest operation

Navigation Model: LLM reads index first to identify relevant pages for queries, then drills into specific content. Works effectively at moderate scale (~100 sources, hundreds of pages) without requiring embedding-based infrastructure.

Example Structure:

## Entities (45 pages)
- openai - AI research company developing GPT models
- anthropic - AI safety company behind Claude
- andrej-karpathy - Former Tesla AI director, advocate of LLM wiki pattern

## Concepts (89 pages)  
- [retrieval-augmented-generation](/concepts/retrieval-augmented-generation) - Combining retrieval with generation
- Vector Embeddings - Dense numerical representations of data

Chronological Log (log.md)

Purpose: Append-only record of all wiki operations and evolution.

Structure:

  • Timestamped entries for ingests, queries, lint operations
  • Consistent format enabling programmatic parsing
  • Timeline of wiki evolution and recent activity
  • Context for understanding system state

Navigation Model: Understand what's been done recently, track system evolution, provide context for LLM about recent operations.

Example Format:

## [2026-12-21 14:30] ingest | Karpathy LLM Wiki Pattern
- **Source**: raw/articles/karpathy-llm-wiki-pattern.md
- **Pages touched**: [llm-wiki-pattern](/concepts/llm-wiki-pattern), [compounding-artifacts](/concepts/compounding-artifacts), [persistent-learning](/concepts/persistent-learning)
- **Summary**: Ingested comprehensive blueprint for LLM-maintained knowledge bases

Design Principles

Complementary Functions

The two patterns serve different but complementary navigation needs:

  • Index: "What knowledge exists and where is it?"
  • Log: "How did this knowledge base evolve over time?"

Parseable Formats

Both use consistent formatting that enables programmatic access:

  • Index supports category-based filtering and summary generation
  • Log enables timeline analysis with simple Unix tools (grep "^## \[" log.md | tail -5)

LLM-Maintained

Both files are automatically maintained by LLMs during normal operations, eliminating human maintenance burden while providing essential navigation capabilities.

Scaling Considerations

Moderate Scale Effectiveness

Content-oriented indexing works well up to hundreds of pages without requiring sophisticated search infrastructure. The human-readable format provides good overview while remaining LLM-navigable.

Search Enhancement

At larger scales, can be supplemented with search tools:

  • Local search engines like qmd for hybrid BM25/vector search
  • Full-text indexing for content that exceeds index-based navigation
  • MCP servers for programmatic access to search capabilities

Hierarchical Organization

Index structure can evolve to use subcategories and nested organization as content volume grows:

### AI/ML Core (89 pages)
### RAG & Search (47 pages)  
### Agent Systems (112 pages)

Implementation Patterns

Automatic Updates

Index updated during every ingest operation as LLM processes new sources and creates/updates wiki pages. Ensures index remains current with no human intervention.

Metadata Integration

Index can include metadata from page frontmatter:

  • Creation/update dates
  • Source counts indicating how well-researched topics are
  • Confidence levels for content reliability
  • Tag summaries for topic clustering

Cross-Reference Support

Index serves as hub for cross-reference discovery - LLM uses it to find related pages when updating content or answering queries requiring synthesis across topics.

Alternatives and Extensions

Graph-Based Navigation

Tools like Obsidian's graph view provide visual representation of page connections, complementing text-based index navigation.

Dynamic Queries

Obsidian's Dataview plugin can generate dynamic indexes based on page metadata, automatically categorizing content based on tags or other frontmatter fields.

Search Integration

Can be combined with full-text search engines for complex queries while maintaining human-readable overview structure.

See also