Concepts — vue longue
retour à la listeToutes les pages concaténées sur un seul document, pour un Ctrl-F direct.
Associative Trails
page dédiée →Persistent pathways through information that connect related documents, concepts, and ideas based on meaningful relationships rather than hierarchical organization. Core concept from vannevar-bush's memex-vision and foundation for modern knowledge management approaches including the llm-wiki-pattern.
Historical Origin
Introduced by vannevar-bush in his 1945 essay "As We May Think," associative trails were envisioned as a fundamental departure from traditional filing systems:
Traditional Filing: Hierarchical organization forces artificial single-category placement of documents.
Associative Trails: Documents connected by meaningful relationships that mirror how human minds work through association and connection.
Core Concept
Trail Building: Users create persistent paths through information, linking documents, passages, and ideas. These trails themselves become valuable intellectual artifacts.
Multiple Connections: Any document can participate in multiple trails, reflecting the reality that information has multiple valid organizational structures.
Trail Persistence: Unlike temporary search results, trails remain as permanent parts of the knowledge structure, available for future navigation and extension.
Modern Implementation
In contemporary knowledge management systems, associative trails manifest as:
Wiki Cross-References: wikilinks that connect related pages based on content relationships rather than folder hierarchy.
Tagged Relationships: Content organized by meaningful tags and relationships rather than rigid categories.
Backlink Networks: Automatic discovery of connections between related content through bidirectional linking.
LLM Wiki Pattern Integration
andrej-karpathy's llm-wiki-pattern implements associative trails through automated cross-referencing:
Automatic Trail Creation: LLM identifies meaningful connections during content ingestion and creates appropriate links.
Trail Maintenance: Cross-references kept current as new information is added, ensuring trails remain valid and valuable.
Emergent Navigation: Multiple pathways through information emerge naturally from content relationships rather than imposed structure.
Value Proposition
Reflects Human Cognition: Mirrors how minds actually work through association and connection rather than rigid hierarchies.
Multiple Access Paths: Same information accessible through various conceptual approaches depending on context and need.
Discovery Enabling: Well-constructed trails reveal unexpected connections and enable serendipitous discovery.
Intellectual Tools: Trails become reusable intellectual artifacts that can be shared, extended, and refined over time.
Implementation Challenges
Manual Maintenance: Traditional associative trail systems require significant human effort to keep connections current and meaningful.
Consistency: Ensuring trail quality and coherence across large knowledge bases becomes overwhelming for individual maintainers.
Discovery: Finding relevant trails when needed requires sophisticated search and navigation capabilities.
LLM Solution
The llm-wiki-pattern addresses traditional challenges through automation:
- Automated identification of meaningful connections
- Consistent maintenance of cross-references
- Dynamic updating as new information arrives
- Quality control across entire knowledge base
See also
- memex-vision
- vannevar-bush
- llm-wiki-pattern
- cross-referencing
- [[knowledge-management-
Automated Content Triage
page dédiée →System for automatically evaluating and routing content based on quality and relevance criteria before ingestion into knowledge management systems. Essential for preventing information overload while maintaining high signal-to-noise ratios.
Core Methodology
Three-Tier Scoring System
- HIGH: Direct ingestion - meets quality and relevance thresholds
- MEDIUM: Manual review queue - potentially valuable but uncertain
- LOW: Archive without processing - insufficient value or off-topic
Evaluation Criteria
Accept (HIGH):
- Introduces new concepts, techniques, or patterns not yet documented
- Contradicts or significantly nuances existing knowledge
- Directly relates to active projects or core domains
- High-quality source (papers, official docs, recognized authors)
- Contains actionable patterns or architectural decisions
Queue for Review (MEDIUM):
- Potentially relevant but uncertain value
- Adjacent to core domains but not central
- Interesting but shallow - might need deeper sources
- Quality source but content overlap with existing knowledge
Reject (LOW):
- Outside defined domains of interest
- Too superficial (brief social media posts without substance)
- Duplicate of already-ingested content
- Promotional content disguised as technical content
- Outdated information superseded in knowledge base
Implementation Patterns
LLM-Based Assessment
Modern triage systems use LLMs to evaluate content against structured criteria:
- Compare against existing knowledge base index
- Apply domain-specific relevance scoring
- Assess content depth and actionability
- Check for novelty and potential insights
Pre-filtering for Efficiency
For high-volume sources (like screenshots), implement lightweight pre-filtering:
- Quick domain relevance check ("Is this tech/AI related?")
- Source quality assessment before deep analysis
- Token-efficient evaluation to reduce processing costs
Routing and Storage
Automated file organization based on triage results:
raw/sources/for accepted contentraw/pending/for manual review itemsraw/discarded/for rejected content (preserved but not processed)
Benefits
Quality Maintenance
- Prevents dilution of knowledge base with low-value content
- Maintains focus on core domains of interest
- Reduces manual curation overhead
- Improves search and discovery within knowledge base
Efficiency Optimization
- Reduces LLM processing costs by filtering out irrelevant content
- Prioritizes valuable content for immediate processing
- Creates manageable review queues for borderline content
- Scales with content volume without linear overhead increase
Advanced Features
Adaptive Thresholds
- Adjust scoring criteria based on knowledge base growth
- Lower acceptance thresholds for new topic areas
- Raise standards for well-covered domains
- Learn from manual review decisions to improve automation
Meta-Analysis
- Track triage patterns to identify content trends
- Monitor false positives/negatives in automated decisions
- Generate insights about content source quality
- Optimize ingestion pipelines based on triage data
Integration with Knowledge Systems
Pipeline Position
Triage typically occurs early in ingestion pipeline:
- Content source detection
- Basic format validation
- Automated triage and scoring
- Routing to appropriate processing path
- Full ingestion for accepted content
Feedback Loops
- Manual review decisions inform future automatic scoring
- User queries reveal gaps where rejected content might be valuable
- Knowledge base usage patterns refine relevance criteria
- Source quality tracking improves upstream filtering
See also
Complete Automation Stack
page dédiée →Architecture pattern for knowledge management systems that eliminates manual content processing through coordinated automation scripts and intelligent orchestration. Enables autonomous operation of complex multi-channel content ingestion pipelines.
Core Components
1. Collection Scripts
Automated content gathering from multiple sources:
- Project Sync:
collect-projects.shusing rsync to sync markdown files from development directories - Screenshot Capture:
process-screenshots.shmoving images from iCloud Drive to processing queues - Audio Processing:
process-recordings.shfor transcribing conference talks and voice memos
2. Git-Based Change Detection
Uses git-diff-tracking to eliminate complex manifest systems:
# Detect exactly what changed since last ingest
git diff --name-status HEAD~1 raw/projects/
3. Daily Orchestration Agent
Centralized automation that runs all pipelines:
- Executes collection scripts in sequence
- Processes new content through LLM ingestion
- Handles error recovery and logging
- Commits changes and updates tracking
4. iOS Integration
Seamless mobile capture through ios-shortcuts-integration:
- Share screenshots directly to iCloud Drive folder
- Automatic sync to processing pipeline
- Frictionless capture from any iOS app
Implementation Pattern
Directory Structure
automation-system/
├── scripts/
│ ├── collect-projects.sh # Sync project files
│ ├── process-screenshots.sh # Handle images
│ ├── process-recordings.sh # Audio transcription
│ └── daily-ingest.md # Orchestration instructions
├── raw/ # Staging area
│ ├── projects/ # Synced project files
│ ├── screenshots/inbox/ # Images awaiting processing
│ └── talks/ # Audio transcripts
└── processed/ # Final output location
Intelligent Exclusions
Critical for avoiding noise in large codebases:
# Exclude common noise directories
--exclude="node_modules" \
--exclude=".git" \
--exclude="dist" \
--exclude=".next" \
--exclude="venv"
Error Handling
Robust pipeline design with:
- SIGPIPE handling for long-running processes
- Proper exit codes and logging
- Recovery mechanisms for partial failures
- Idempotent operations for safe retries
Automation Benefits
1. Zero Manual Overhead
Once configured, the system operates autonomously:
- New projects automatically detected and ingested
- Screenshots processed without intervention
- Conference recordings transcribed and integrated
- Cross-references updated automatically
2. Compound Learning Effects
Automation enables sophisticated cross-modal integration:
- Screenshots from social media link to related projects
- Conference insights connect to technical documentation
- Project learnings reference research papers automatically
3. Consistent Processing
Eliminates human inconsistency in:
- Content categorization
- Cross-reference creation
- Quality assessment
- Archive organization
4. Scalable Growth
System capacity grows with content volume:
- Handles thousands of files efficiently
- Incremental processing prevents performance degradation
- Git-based tracking scales to large repositories
Real-World Results
From the epic brain-wiki implementation:
- 4,060 markdown files automatically synchronized
- 31 projects processed in first run
- Zero manual intervention after initial setup
- Daily operation with autonomous content updates
Architecture Principles
1. Pipeline Composition
Break complex workflows into composable scripts:
- Each script handles one responsibility
- Clear interfaces between components
- Easy to test and debug individually
2. State Management
Use Git as the source of truth:
- No complex manifest files
- Built-in versioning and rollback
- Distributed and reliable
3. Graceful Degradation
System continues operating with partial failures:
- Individual pipeline failures don't break the whole system
- Clear error reporting for manual intervention
- Automatic retry mechanisms where appropriate
4. Human-in-the-Loop
Automation handles routine tasks, humans handle decisions:
- Content quality assessment
- New pipeline configuration
- System monitoring and optimization
Implementation Considerations
Resource Management
- Use rsync for efficient file synchronization
- Implement rate limiting for API calls
- Monitor disk usage and cleanup old files
Security
- Protect API keys and credentials
- Use private repositories for sensitive content
- Implement access controls on automation scripts
Monitoring
- Log all automation activities
- Track success/failure rates
- Monitor resource consumption
- Set up alerts for system failures
Advanced Features
Conditional Processing
Smart content filtering to avoid redundant work:
# Only process files modified since last run
find raw/projects -name "*.md" -newer .last_ingest
Parallel Execution
For large content volumes:
- Process multiple projects simultaneously
- Batch OCR operations for efficiency
- Async transcription of audio files
Quality Gates
Automated content validation:
- Check for minimum content length
- Validate markdown formatting
- Verify cross-references resolve
See Also
- multi-source-ingestion - Content pipeline architecture
- git-diff-tracking - Change detection methodology
- daily-automation-agents - Orchestration patterns
- ios-shortcuts-integration - Mobile capture workflows
- brain-wiki - Real-world implementation example
Educational Content Preservation
page dédiée →Strategic approach to maintaining long-term access to educational materials, particularly relevant for bootcamp alumni and students concerned about platform sustainability or policy changes. Involves creating independent systems to archive, export, and host educational content.
Preservation Motivations
Platform Sustainability Concerns
- Educational platforms may discontinue services
- Business model changes affecting content access
- Lifetime access promises potentially unfulfilled
- Need for continued learning resource availability
Knowledge Asset Protection
Educational content represents significant investment:
- Time spent in intensive learning programs
- Financial investment in bootcamp tuition
- Accumulated knowledge and exercise solutions
- Project templates and reference materials
Technical Preservation Strategies
Content Export and Migration
- Structured Export: Maintaining original organization and metadata
- Format Conversion: Converting platform-specific formats to standard formats (MDX, Markdown)
- Asset Preservation: Downloading and archiving multimedia content
- Relationship Mapping: Preserving links and dependencies between materials
Independent Hosting Solutions
- Custom LMS Development: Building personal learning management systems
- Static Site Generation: Creating browsable documentation sites
- Local Server Deployment: Self-hosted solutions for content access
- Cloud Storage Integration: Distributed backup strategies
Content Structure Preservation
Hierarchical Organization
Maintaining educational content structure:
Course/Track Level
├── Module Organization
│ ├── Lesson Sequence
│ ├── Exercise Materials
│ └── Project Templates
└── Assessment Components
Metadata Preservation
Critical information to maintain:
- Learning objectives and outcomes
- Prerequisites and dependencies
- Difficulty levels and progression
- Completion tracking data
- Original publication dates
Implementation Approaches
Dual-Source Architecture
Supporting both preserved content and new material:
- File-based Storage: Exported materials as local files
- Database Integration: Dynamic content management capabilities
- Hybrid Rendering: Seamless presentation of both content types
Progressive Enhancement
- Offline Accessibility: Content available without internet connection
- Search Capabilities: Full-text search across preserved materials
- Interactive Features: Maintaining quiz and exercise functionality
- Progress Tracking: Personal learning progress independent of original platform
Legal and Ethical Considerations
Fair Use and Personal Backup
- Personal archival for continued learning
- Non-commercial use of educational materials
- Respect for intellectual property rights
- Compliance with platform terms of service
Knowledge Sharing Ethics
- Appropriate attribution of original sources
- Sharing methodologies rather than copyrighted content
- Contributing to educational tool development
- Supporting open-source educational initiatives
Benefits for Learners
Continued Access Assurance
- Permanent access to educational investments
- Independence from platform business decisions
- Ability to reference materials throughout career
- Foundation for continued learning and skill development
Enhanced Learning Experience
- Personalized organization and note-taking
- Custom search and discovery features
- Integration with personal knowledge management systems
- Ability to supplement with additional resources
Portfolio Development Value
Educational content preservation projects demonstrate:
- Technical Proficiency: Full-stack development capabilities
- Problem-Solving Skills: Addressing real-world sustainability concerns
- Project Management: Planning and executing complex data migration
- User Experience Design: Creating intuitive educational interfaces
See also
- le-wagon
- lms-mvp-development
- knowledge-management-systems
- Content Migration Strategies
- Educational Technology
Git Diff Tracking
page dédiée →Version control pattern for detecting and processing only changed content in knowledge management systems. Eliminates the need for complex manifest systems by leveraging Git's native change detection capabilities.
Core Concept
Instead of building custom tracking mechanisms, use Git diff to identify exactly what content has changed since the last processing run. This approach provides reliable change detection with minimal overhead and perfect accuracy.
Implementation Pattern
Basic Change Detection
# Identify all changed files since last successful ingest
git diff HEAD~1 --name-only raw/projects/ | while read file; do
echo "Processing changed file: $file"
# Process only this specific file
done
Incremental Project Synchronization
# Sync projects from development environment
rsync -av --include='*.md' --exclude='node_modules' ~/code/*/ raw/projects/
# Git shows exactly what changed
git add raw/projects/
if ! git diff --cached --quiet; then
echo "New content detected, processing changes..."
git diff --cached --name-only | process_changed_files
git commit -m "Sync projects: $(date)"
fi
Architecture Benefits
Eliminates Manifest Complexity
Traditional approaches require maintaining separate tracking files:
- Timestamp-based systems prone to clock skew
- Hash-based manifests requiring custom logic
- Database tracking adding system complexity
Git diff provides all tracking functionality with zero custom code.
Atomic Processing
Git commits create natural processing boundaries:
- Each commit represents one complete ingestion cycle
- Failed processing can be retried from last successful commit
- Processing history provides complete audit trail
Distributed Reliability
Git's distributed nature provides automatic backup and synchronization:
- Processing history preserved across machines
- Remote repository serves as authoritative state
- Merge conflicts reveal concurrent processing issues
Advanced Patterns
Selective Processing by File Type
# Process only markdown changes
git diff HEAD~1 --name-only --diff-filter=AM | grep '\.md$' | process_files
# Handle deletions separately
git diff HEAD~1 --name-only --diff-filter=D | cleanup_deleted_files
Content-Aware Change Detection
# Get actual content changes, not just file list
git diff HEAD~1 raw/projects/ | while read line; do
if $line =~ ^\+\+\+; then
# New file or significant addition
process_file_addition
elif $line =~ ^\---; then
# Deleted content
process_file_removal
fi
done
Automated Commit Integration
# Combine sync and change detection
sync_projects() {
rsync -av ~/code/*/ raw/projects/
if git add raw/projects/ && ! git diff --cached --quiet; then
# Process only changed content
git diff --cached --name-only | process_incremental_changes
# Commit after successful processing
git commit -m "Auto-sync: $(date) - $(git diff --cached --numstat | wc -l) files"
echo "✓ Processed $(git diff HEAD~1 --numstat | wc -l) changed files"
else
echo "No changes detected"
fi
}
Error Handling
Processing Failure Recovery
# If processing fails, reset to last known good state
if ! process_changes; then
echo "Processing failed, resetting to last commit"
git reset --hard HEAD
exit 1
fi
Partial Success Handling
# Process files individually to isolate failures
git diff HEAD~1 --name-only | while read file; do
if process_single_file "$file"; then
git add "$file"
echo "✓ Processed $file"
else
echo "✗ Failed to process $file"
# Continue with other files
fi
done
# Commit whatever succeeded
if ! git diff --cached --quiet; then
git commit -m "Partial sync: $(date)"
fi
Integration Patterns
Daily Automation Integration
# Daily agent workflow
cd /path/to/knowledge-base
# Sync latest content
./scripts/collect-projects.sh
# Process only changes
if ! git diff --quiet raw/projects/; then
echo "Changes detected, starting processing..."
git add raw/projects/
changed_files=$(git diff --cached --name-only | wc -l)
# Process incrementally
process_git_changes
# Update wiki index
update_index_with_changes
git commit -m "Daily sync: $changed_files files updated"
echo "✓ Processed $changed_files changed files"
else
echo "No changes since last sync"
fi
Multi-Repository Coordination
# Track changes across multiple source repositories
for repo in ~/code/*/; do
repo_name=$(basename "$repo")
# Sync and track per-repository changes
GitHub Private Repo Architecture
page dédiée →Pattern for hosting knowledge management systems using GitHub private repositories as the central storage and collaboration layer, enabling version control, automated processing, and multi-tool integration while maintaining privacy and control.
Architecture Benefits
Version Control for Knowledge
Git provides natural versioning for wiki content evolution:
- Track changes to concepts and understanding over time
- Rollback incorrect updates or information
- Branch for experimental content organization
- Merge knowledge from multiple sources systematically
Tool Ecosystem Integration
GitHub serves as a neutral platform accessible by multiple tools:
- Local editing through Obsidian, VSCode, or any markdown editor
- Automated processing through GitHub Actions
- API access for programmatic updates via wiki agents
- Web interface for manual review and curation
Privacy and Control
Private repositories maintain confidentiality while enabling automation:
- Sensitive project information remains private
- API tokens control access without exposing content
- Self-hosted processing maintains data sovereignty
- Export/migration remains straightforward through Git
Repository Structure Pattern
llm-wiki/
SCHEMA.md # System conventions and workflows
index.md # Master catalog of all content
log.md # Chronological operation history
BUILDING.md # Meta-construction documentation
raw/ # Immutable source documents
screenshots/ # Mobile-captured images
articles/ # Web content and newsletters
transcripts/ # Audio processing results
projects/ # Extracted project documentation
conversations/ # Agent interaction exports
papers/ # Research materials
pending/ # Medium-quality sources awaiting review
discarded/ # Low-quality sources with rejection reasons
wiki/ # LLM-maintained structured content
entities/ # People, tools, organizations
concepts/ # Technical knowledge and patterns
projects/ # Implementation documentation
sources/ # Individual source summaries
syntheses/ # Cross-source analysis
skills/ # Mastered techniques
flashcards/ # Spaced repetition content
digests/ # Periodic summaries
tools/ # Automation and processing scripts
Automation Integration
Git Hooks
Trigger processing on content changes:
- Pre-commit hooks validate markdown structure
- Post-commit hooks trigger wiki agent processing
- Push hooks initiate automated backups
GitHub Actions
Enable cloud-based processing workflows:
- Automated content quality checking
- Scheduled digest generation
- Integration testing for wiki structure
- Deployment to static site generators
Local Processing
Balance privacy with cloud convenience:
- Sensitive processing remains local
- Public API calls for general knowledge enhancement
- Hybrid workflows combining local and cloud resources
Multi-Tool Compatibility
Obsidian: Native markdown and wikilink support for interactive browsing and manual editing.
VSCode/Cursor: Full development environment access for complex editing and automation development.
MCP Servers: Programmatic access for LLM tools and agents through standardized protocols.
Mobile Apps: Git clients enable basic editing and review from mobile devices.
Implementation Considerations
Authentication: SSH keys or personal access tokens for automated access without exposing credentials.
Backup Strategy: Git provides inherent backup, but consider additional external backups for critical knowledge bases.
Search Integration: GitHub's search capabilities combined with local indexing provide comprehensive content discovery.
Collaboration: Private repos support controlled sharing with team members while maintaining primary ownership.
See also
- ai-engineering-wiki - Implementation example using this pattern
- llm-wiki-pattern - Knowledge management approach
- version-control-for-knowledge - Git workflows for content management
- obsidian-github-integration - Tool-specific implementation patterns
Human-LLM Division of Labor
page dédiée →The strategic allocation of responsibilities between humans and LLMs in llm-wiki-pattern systems, optimizing each party's strengths while solving the traditional wiki abandonment problem. Core insight: humans excel at curation and synthesis; LLMs excel at maintenance and bookkeeping.
Task Allocation
Human Responsibilities
- Source curation: Selecting valuable documents to ingest
- Direction setting: Guiding analysis emphasis and priorities
- Question asking: Driving exploration through strategic queries
- Meaning synthesis: Understanding implications and significance
- Quality oversight: Reviewing summaries and checking updates
- Schema evolution: Adapting system configuration based on needs
LLM Responsibilities
- Content maintenance: Updating cross-references, keeping summaries current
- Consistency management: Noting contradictions, maintaining coherence
- Bookkeeping automation: Filing, indexing, logging operations
- Cross-referencing: Building and maintaining wikilinks networks
- Structure creation: Generating new pages, organizing content
- Workflow execution: Following schema-defined operational procedures
Solving the Abandonment Problem
Traditional Wiki Failure Pattern
- Initial enthusiasm: Humans start with high motivation
- Growing burden: Maintenance tasks accumulate faster than value
- Cognitive overhead: Cross-referencing becomes mentally taxing
- Inevitable abandonment: Effort required exceeds perceived benefit
LLM Solution
- Zero maintenance fatigue: LLMs don't experience tedium or boredom
- Consistent execution: Never forget to update cross-references
- Parallel processing: Can touch 15 files in one pass without cognitive load
- Near-zero cost: Maintenance burden becomes negligible
Cognitive Complementarity
Human Cognitive Strengths
- Contextual judgment: Understanding significance and relevance
- Creative synthesis: Making novel connections and insights
- Domain expertise: Applying specialized knowledge and intuition
- Strategic thinking: Long-term planning and goal-oriented exploration
LLM Cognitive Strengths
- Systematic processing: Consistent application of rules and procedures
- Pattern recognition: Identifying structural relationships across content
- Parallel attention: Managing multiple interconnected updates simultaneously
- Infinite patience: Performing repetitive tasks without degradation
Operational Boundaries
Human Decision Points
- Source selection: What documents deserve ingestion?
- Emphasis guidance: What aspects need highlighting?
- Quality gates: Are summaries accurate and useful?
- Exploration direction: What questions should drive further investigation?
LLM Execution Points
- Content integration: How to incorporate new information?
- Link maintenance: Which pages need cross-reference updates?
- Consistency checks: Where do contradictions need flagging?
- Structure organization: How to categorize and file content?
Interface Design
Human-LLM Interaction Model
- LLM agent open on one side of screen
- Obsidian (or wiki browser) open on other side
- Real-time collaboration: Human monitors LLM edits live
- Immediate feedback: Human can guide and correct during operation
Communication Patterns
- Explicit instructions: Human provides clear direction for emphasis
- Progress reporting: LLM describes what updates are being made
- Quality confirmation: Human reviews and approves significant changes
- Schema discussion: Collaborative evolution of system configuration
Benefits of Clear Division
Efficiency Optimization
- Leverage strengths: Each party focuses on optimal tasks
- Minimize waste: Avoid humans doing tedious work, LLMs making judgment calls
- Sustainable workflow: Maintenance burden doesn't grow with scale
Quality Assurance
- Human oversight: Strategic decisions remain under human control
- LLM consistency: Mechanical tasks executed reliably
- Complementary validation: Different cognitive approaches catch different errors
Implementation Considerations
Trust Building
- Gradual automation: Start with supervised workflows, increase autonomy
- Transparency: LLM reports all changes and reasoning
- Reversibility: Git versioning enables rollback of problematic updates
Workflow Evolution
- Usage-driven refinement: Division of labor adapts based on experience
- Domain customization: Different fields may require different task allocations
- Tool integration: Technical capabilities influence responsibility boundaries
See also
- llm-wiki-pattern - Framework implementing this division
- wiki-maintenance-automation - Automated tasks LLMs handle
- bookkeeping-automation - Specific maintenance functions
- schema-coevolution - Collaborative system configuration process
Intelligent Content Triage
page dédiée →Automated system for evaluating and routing incoming content based on relevance, quality, and potential value to knowledge base. Essential component of scalable knowledge management systems to prevent information overload while ensuring valuable content is captured and processed.
Triage Framework
Three-Tier Scoring System
HIGH (Auto-ingest):
- Introduces concepts/techniques not yet in wiki
- Contradicts or significantly nuances existing content
- Directly relates to active projects
- From high-quality sources (papers, official docs, recognized authors)
- Contains actionable code patterns or architecture decisions
MEDIUM (Manual review queue):
- Potentially relevant but uncertain value
- Adjacent to current domains but not core
- Interesting but shallow content that might need deeper sourcing
- Routed to
raw/pending/folder for later evaluation
LOW (Archive/discard):
- Not related to domains of interest
- Too superficial (3-line social media posts without substance)
- Duplicate of already-ingested content
- Promotional content disguised as technical content
- Outdated information already superseded in wiki
Implementation Architecture
Screenshot Filtering Pipeline
For iPhone social media screenshots - most common source requiring filtering:
- Quick Domain Check: "Is this tech/AI related?" - fast binary filter
- Content Analysis: OCR + vision model text extraction
- Relevance Scoring: Compare against existing wiki index
- Quality Assessment: Source credibility, depth, actionability
- Routing Decision: HIGH → ingest, MEDIUM → pending, LOW → discard
Automated Quality Gates
Source Credibility Factors:
- Author reputation in AI/tech domains
- Publication venue quality (arXiv, official docs, conferences)
- Content depth and technical specificity
- Presence of code examples or concrete implementations
Relevance Criteria:
- Alignment with current project portfolio
- Overlap assessment with existing wiki content
- Introduction of genuinely new concepts vs. rehashed basics
- Potential for compound learning and cross-referencing
Compound Learning Effect
The triage system improves over time through knowledge accumulation:
- Basic Transformer articles scored LOW when wiki already contains 15 related pages
- Novel techniques (Flash Attention, MoE architectures) scored HIGH
- Scoring accuracy increases as wiki index grows more comprehensive
- Domain understanding deepens, enabling better quality assessment
Real-World Performance
Live Implementation Results (June 2026):
- Karpathy's LLM wiki document: HIGH score (correctly identified as authoritative, novel)
- Automatic routing prevented manual review overhead
- Generated 5 wiki pages + 5 flashcards from single high-quality source
- Zero false positives in initial testing phase
Integration Patterns
Telegram Bot Integration
Direct scoring from mobile messaging:
- Send URL: bot fetches, scores, routes appropriately
- Send screenshot: vision model extracts text, scores content
- Real-time feedback on triage decisions
- Manual override capability for edge cases
Multi-Source Adaptation
Different scoring criteria per source type:
- Academic papers: Novelty + technical depth
- Project documentation: Direct applicability + completeness
- Social media: Signal-to-noise ratio + source authority
- Audio transcripts: Actionable insights + speaker expertise
Technical Implementation
LLM Prompt Structure:
# Compare against existing wiki index
# Apply domain relevance criteria from SCHEMA.md
# Score using established quality gates
# Return: HIGH/MEDIUM/LOW + reasoning
File System Routing:
raw/sources/→ HIGH content for immediate processingraw/pending/→ MEDIUM content awaiting manual reviewraw/discarded/→ LOW content archived with reasoning
This systematic approach enables high-volume content processing while maintaining knowledge base quality, essential for ai-engineering-wiki automation at scale.
See also
- multi-source-ingestion
- ai-engineering-wiki
- content-quality-assessment
- llm-wiki-pattern
Knowledge Management Systems
page dédiée →Systems designed to capture, organize, and make accessible the knowledge and expertise within an organization or for personal use. Range from simple note-taking tools to sophisticated AI-powered knowledge bases.
Evolution of Approaches
Traditional Approaches
- File systems: Folders and documents
- Wikis: Collaborative, hyperlinked documentation
- Databases: Structured information storage
- Content management: Document-centric systems
Modern AI-Enhanced Approaches
- RAG systems: Retrieve relevant chunks at query time
- llm-wiki-pattern: LLM-maintained persistent knowledge bases
- Semantic search: Vector-based information retrieval
- AI summarization: Automated content distillation
Key Challenges
The Maintenance Problem
Traditional knowledge management systems often fail because:
- High maintenance burden: Keeping information current requires constant human effort
- Cross-referencing overhead: Manually maintaining links between related concepts
- Inconsistency growth: Different pages contradict each other over time
- Abandonment: Systems become stale as maintenance burden exceeds value
The Discovery Problem
- Finding relevant information in large knowledge bases
- Understanding connections between concepts
- Surfacing insights that span multiple sources
LLM-Enhanced Solutions
Automated Maintenance
LLMs can handle the tedious aspects of knowledge management:
- Updating cross-references automatically
- Identifying and flagging contradictions
- Maintaining consistent formatting and structure
- Generating summaries and overviews
Intelligent Organization
- Automatic categorization and tagging
- Dynamic relationship discovery
- Content synthesis across sources
- Gap identification and suggestion of new sources
Types of Knowledge Management
Personal Knowledge Management
- Research notes and literature reviews
- Learning from courses and books
- Project documentation and lessons learned
- Health, goals, and self-improvement tracking
Organizational Knowledge Management
- Team documentation and best practices
- Customer interaction histories
- Technical knowledge bases
- Institutional memory preservation
Tools and Platforms
Traditional Tools
- Obsidian: Graph-based personal knowledge management
- Notion: Block-based collaborative workspaces
- Roam Research: Bi-directional linking for research
- MediaWiki: Traditional collaborative wiki platform
AI-Enhanced Tools
- NotebookLM: Google's AI-powered research assistant
- ChatGPT with files: Document upload and querying
- Custom LLM wiki systems: Using the LLM Wiki Pattern
Design Principles
For Sustainable Systems
- Low maintenance overhead: Automate bookkeeping tasks
- Clear human-AI division: Humans curate and direct; AI maintains
- Incremental value: Each addition makes the system more valuable
- Consistent structure: Standardized formats and conventions
- Version control: Track changes and maintain provenance
For Effective Discovery
- Multiple access patterns: Browse, search, and follow links
- Rich cross-referencing: Connect related concepts
- Hierarchical organization: Categories and subcategories
- Temporal tracking: Understand how knowledge evolved
Success Factors
- Clear purpose: Focused domain and use cases
- Consistent workflow: Standardized processes for adding content
- Regular use: Systems that aren't used regularly decay
- Appropriate tooling: Match tools to user preferences and technical skills
- Sustainable maintenance: Either automated or minimal manual overhead
See also
LLM Wiki Pattern
page dédiée →A paradigm for building personal knowledge bases where LLMs incrementally build and maintain persistent wikis rather than retrieving from raw documents at query time. Developed by andrej-karpathy as an alternative to traditional RAG systems that rediscover knowledge from scratch on every interaction.
Core Innovation
Persistent vs Ephemeral Knowledge: Most people's experience with LLMs and documents follows the RAG pattern - upload files, retrieve relevant chunks at query time, generate answers. This works but requires rediscovering knowledge from scratch on every question. The LLM Wiki Pattern instead creates persistent, compounding artifacts where knowledge is compiled once and kept current.
Key Difference: The wiki sits between you and raw sources. When adding new sources, the LLM doesn't just index for later retrieval - it reads, extracts key information, and integrates into existing wiki structure. Updates entity pages, revises summaries, flags contradictions, strengthens synthesis. Cross-references already exist. Contradictions already flagged. Synthesis already reflects everything read.
Three-Layer Architecture
-
Raw Sources: Immutable curated collection (articles, papers, images, data files). LLM reads but never modifies. Source of truth.
-
The Wiki: Directory of LLM-generated markdown files. Summaries, entity pages, concept pages, comparisons, synthesis. LLM owns this layer entirely - creates, updates, maintains cross-references, ensures consistency.
-
The Schema: Configuration document (CLAUDE.md, AGENTS.md) defining wiki structure, conventions, workflows. Co-evolved between human and LLM over time as patterns emerge.
Core Operations
Ingest Workflow
Drop new source into raw collection. LLM reads source, discusses takeaways, writes summary page, updates index, updates relevant entity/concept pages across wiki, appends log entry. Single source might touch 10-15 wiki pages. Can be done one-at-a-time with supervision or batch-processed.
Query Workflow
Ask questions against wiki. LLM searches relevant pages, reads them, synthesizes answer with citations. Answers can take multiple forms - markdown pages, comparison tables, slide decks (marp-integration), charts, canvas. Critical insight: Good answers get filed back as new wiki pages. Explorations compound in knowledge base like ingested sources.
Lint Workflow
Periodic health-checking. Look for contradictions between pages, stale claims superseded by newer sources, orphan pages with no inbound links, important concepts mentioned but lacking pages, missing cross-references, data gaps. LLM suggests new questions to investigate and sources to find.
Navigation Infrastructure
Index and Logging
-
index.md: Content-oriented catalog of all wiki pages with links, summaries, metadata. Organized by category. Updated on every ingest. LLM reads index first to find relevant pages for queries. Works well at moderate scale (~100 sources, hundreds of pages).
-
log.md: Chronological append-only record of ingests, queries, lint passes. Parseable with consistent prefixes (
## [2026-04-02] ingest | Article Title). Provides timeline of wiki evolution.
Optional Tooling
qmd-search engine recommended for scaling beyond index-based navigation. Local search for markdown with hybrid BM25/vector search and LLM re-ranking. Available as CLI tool and MCP server.
Implementation Recommendations
Obsidian Integration
Recommended interface: obsidian-integration where human browses in Obsidian while LLM makes live edits to markdown files. "Obsidian is the IDE; the LLM is the programmer; the wiki is the codebase."
Useful Obsidian Features:
- Web Clipper: Browser extension converting articles to markdown
- Image handling: Download attachments locally, bind to hotkey (Ctrl+Shift+D)
- Graph view: Visualize wiki connections, identify hubs and orphans
- marp-integration: Generate slide decks from wiki content
- dataview-plugin: Query page frontmatter for dynamic tables/lists
- Git integration: Version history, branching, collaboration
Use Cases
Personal: Goals, health, psychology tracking. File journal entries, articles, podcast notes into structured self-picture over time.
Research: Deep topic exploration over weeks/months. Build comprehensive wiki with evolving thesis.
Reading Companion: File chapters as you go. Build pages for characters, themes, plot threads. Like fan wikis (Tolkien Gateway) but personal with LLM maintenance.
Business/Team: Internal wiki fed by Slack threads, meeting transcripts, project docs, customer calls. Humans review updates.
Other: Competitive analysis, due diligence, trip planning, course notes, hobby deep-dives.
Historical Foundation
Explicitly references vannevar-bush's memex-vision (1945) as spiritual predecessor. Bush envisioned personal, curated knowledge store with associative-trails between documents. His vision was private, actively curated, with connections between documents as valuable as documents themselves.
Key Insight: Bush couldn't solve the maintenance problem. LLMs handle that through bookkeeping-automation. Humans abandon wikis because maintenance burden grows faster than value. LLMs don't get bored, don't forget cross-references, can touch 15 files in one pass. Maintenance cost approaches zero.
Division of Labor
Human Role: Curate sources, direct analysis, ask good questions, think about meaning.
LLM Role: Summarizing, cross-referencing, filing, bookkeeping that makes knowledge base useful over time.
Design Philosophy
Intentionally abstract specification describing the pattern, not specific implementation. Directory structure, schema conventions, page formats, tooling all depend on domain, preferences, LLM choice. Everything modular and optional - pick what's useful. Designed to be shared with LLM agents for collaborative instantiation.
See also
Memex Vision
page dédiée →vannevar-bush's 1945 vision of a personal knowledge management system that would allow individuals to store, retrieve, and create associative-trails between documents and information. A foundational concept for modern knowledge management and the direct inspiration behind andrej-karpathy's llm-wiki-pattern and other AI-powered information systems.
Historical Context
Introduced in Bush's essay "As We May Think" (1945), the Memex was conceived as a mechanical device that would supplement human memory by allowing rapid consultation of personal collections of documents. Bush envisioned a desk-sized device with screens, keyboards, and mechanical storage systems.
Core Vision
Private curation - Personal knowledge stores actively maintained by individuals rather than shared databases
Associative trails - Connections between documents as valuable as the documents themselves. Users could create and follow paths linking related information across their collection.
Mechanical augmentation - Technology amplifying human cognitive capabilities rather than replacing human judgment
Active maintenance - Knowledge bases requiring continuous organization and cross-referencing to remain valuable
The Maintenance Problem
Bush's vision was prescient but incomplete. He understood the value of connected, curated knowledge but couldn't solve the fundamental challenge: who does the maintenance work?
Creating associative trails, maintaining cross-references, keeping summaries current, noting contradictions - this bookkeeping is essential but tedious. Humans consistently abandon personal knowledge systems because maintenance burden grows faster than perceived value.
Modern Resolution
The llm-wiki-pattern directly addresses Bush's unsolved maintenance problem. LLMs excel at the systematic bookkeeping that humans find burdensome:
- Updating cross-references across multiple pages
- Maintaining consistency as information accumulates
- Noting contradictions between sources
- Creating and strengthening associative connections
This allows Bush's original vision to be realized: private, actively curated knowledge stores where connections between documents are as valuable as documents themselves.
Influence on Contemporary Systems
The Memex vision influenced:
- Hypertext systems and the World Wide Web
- Personal knowledge management tools (Roam, Obsidian, Notion)
- Information retrieval research
- Modern AI-powered knowledge systems
However, most implementations lost Bush's emphasis on private curation and became either:
- Public databases (Wikipedia, web)
- Simple storage without maintained associations (file systems)
- Tools requiring manual maintenance (personal wikis that get abandoned)
Key Insights for LLM Wiki Pattern
Connection value - The relationships between pieces of information often more valuable than isolated documents
Curation over scale - Personal, curated collections outperform generic large databases for individual needs
Maintenance as bottleneck - Technical capability less important than solving the maintenance burden problem
Human-machine collaboration - Best results combine human judgment (curation, direction) with machine capability (systematic maintenance)
See also
Multi-Source Ingestion
page dédiée →Architecture pattern for knowledge management systems that automatically captures and processes content from diverse sources and formats. Essential for comprehensive knowledge accumulation without manual overhead, enabling systematic learning from all information channels.
Core Architecture
Source Categories
1. Development Projects
- Automated sync from
~/code/*/using rsync - Intelligent filtering (exclude node_modules, .git, build artifacts)
- Git diff tracking to process only changed files
- Captures READMEs, documentation, and learning artifacts
2. Visual Content
- iOS Shortcuts integration for frictionless screenshot capture
- iCloud Drive sync for automatic mobile-to-desktop transfer
- LLM vision for OCR and semantic classification
- Context-aware processing of social media content
3. Audio Content
- Conference talks and meeting recordings
- Voice memos and interview transcripts
- Whisper-based transcription with MLX optimization
- Automatic speaker identification and topic segmentation
4. Web Content
- RSS feed aggregation from key sources
- Manual article addition with URL parsing
- Research paper ingestion from arXiv and academic sources
- Blog posts and technical documentation
5. Conversational Content
- AI agent conversations from Claude Code, Cursor IDE
- Chat transcripts with technical discussions
- Code review conversations and architectural decisions
Processing Pipeline
Raw Sources → Triage → Classification → Extraction → Integration
↓ ↓ ↓ ↓ ↓
Diverse Quality Content-Type Key Info Wiki Pages
Formats Filter Detection Capture & References
Triage System
Automated Quality Assessment:
- Content length and depth analysis
- Source credibility scoring
- Relevance to existing knowledge base
- Novelty detection against existing pages
Classification Outcomes:
- High: Immediate processing and integration
- Medium: Queue for manual review in
raw/pending/ - Low: Archive to
raw/discarded/with reasoning
Content Type Detection
Intelligent Format Handling:
- Markdown parsing and structure analysis
- Image OCR with context understanding
- Audio transcription with speaker diarization
- Code extraction and documentation linking
Implementation Patterns
Git-Based Change Detection
Revolutionary approach using git-diff-tracking to eliminate manifest complexity:
# Detect changes since last ingest
git diff --name-status HEAD~1 raw/projects/
# Process only modified files
find raw/ -newer .last_ingest -type f
iOS Shortcuts Integration
Seamless mobile capture through ios-shortcuts-integration:
iPhone Screenshot → Share → "Brain Wiki" Shortcut → iCloud Drive → Processing Queue
Batch Processing
Efficient handling of large content volumes:
- Process screenshots in batches for OCR efficiency
- Parallel transcription of multiple audio files
- Async classification of web articles
- Incremental project synchronization
Cross-Modal Synthesis
Automated Cross-Referencing
- Screenshot insights link to related projects
- Conference talks reference technical documentation
- Project learnings connect to research papers
- Agent conversations enhance concept pages
Compound Learning Effects
Multi-source integration creates knowledge compounds:
- Visual content + code examples + academic papers = comprehensive understanding
- Social media insights + project experience + formal documentation = practical wisdom
Real-World Implementation
Epic Brain Wiki Results
From the complete brain-wiki implementation:
Sources Processed:
- 31 AI/ML projects (4,060 markdown files)
- Screenshot pipeline with iOS Shortcuts integration
- Audio transcription using MLX Whisper
- Web content through RSS and manual addition
- Agent conversations from Claude Code sessions
Automation Achieved:
- Zero manual overhead for routine ingestion
- Daily orchestration processing all sources
- Intelligent filtering eliminating 90% of noise
- Cross-modal integration creating compound insights
Measured Benefits
- 10x faster knowledge capture compared to manual curation
- 5x more cross-references discovered through automated analysis
- 90% reduction in manual content processing time
- Continuous operation without human intervention
Technical Implementation
Directory Structure
multi-source-system/
├── raw/ # Immutable source storage
│ ├── projects/ # Development project files
│ ├── screenshots/ # Visual content capture
│ ├── talks/ # Audio transcriptions
│ ├── articles/ # Web content archive
│ ├── conversations/ # Agent chat logs
│ ├── pending/ # Medium-quality sources
│ └── discarded/ # Low-quality archive
├── scripts/
│ ├── collect-projects.sh # Project synchronization
│ ├── process-screenshots.sh # Image handling
│ ├── process-recordings.sh # Audio transcription
│ └── daily-ingest.md # Orchestration guide
└── processed/ # Integrated wiki content
Automation Scripts
Project Collection:
# Sync all code projects
rsync -av --exclude="node_modules" --exclude=".git" \
~/code/ raw/projects/
Screenshot Processing:
# Move from iCloud to processing queue
mv ~/Library/Mobile\ Documents/com~apple~CloudDocs/brain-wiki-inbox/* \
raw/screenshots/inbox/
Audio Transcription:
# Whisper transcription
mlx_whisper audio_file.mp3 --output-format txt
Quality Control
Automated Validation
- Minimum content length thresholds
- Duplicate detection across sources
- Link validation and reference checking
- Format consistency verification
Human Oversight
- Weekly review of medium-quality sources
- Manual curation of cross-references
- System performance monitoring
- Pipeline optimization decisions
Scaling Considerations
Performance Optimization
- Incremental processing to handle growth
- Batch operations for efficiency
- Resource pooling for concurrent tasks
- Cache layers for repeated operations
Storage Management
- Compression for archived content
- Automated cleanup of outdated sources
- Backup strategies for critical content
- Version control for all processed data
Advanced Features
Content Enrichment
- Automatic tagging based on content analysis
- Entity extraction and relationship mapping
- Topic modeling for content clustering
- Sentiment analysis for conversational content
Adaptive Processing
- Learning from user feedback on content quality
- Dynamic threshold adjustment for triage
- Personalized relevance scoring
- Context-aware classification improvement
Implementation Challenges
Content Quality Variability
- Handling low-signal sources (screenshots of memes)
- Dealing with incomplete or corrupted files
- Managing different content formats and structures
- Balancing automation with quality control
Technical Complexity
- Coordinating multiple processing pipelines
- Error handling across diverse input types
- Resource management for intensive operations
- Maintaining system reliability and uptime
Privacy and Security
- Sensitive content identification and handling
- Access control for different source types
- Secure storage of personal information
- Compliance with data protection requirements
See Also
- complete-automation-stack - Overall automation architecture
- git-diff-tracking - Change detection methodology
- ios-shortcuts-integration - Mobile capture workflows
- whisper-transcription-workflows - Audio processing
- daily-automation-agents - Orchestration system
- brain-wiki - Complete implementation example
RAG Alternative Approaches
page dédiée →Methods for working with large document collections that go beyond traditional Retrieval-Augmented Generation patterns. Most notably exemplified by andrej-karpathy's llm-wiki-pattern, which creates persistent, pre-synthesized knowledge structures rather than retrieving raw chunks at query time.
Traditional RAG Limitations
Rediscovery Problem: RAG systems retrieve relevant chunks and generate answers from scratch on every query. No knowledge accumulates—subtle questions requiring synthesis across multiple documents must reconstruct understanding each time.
Fragment-Based Understanding: Working with retrieved chunks rather than integrated knowledge, limiting the ability to develop sophisticated cross-source synthesis.
Stateless Operation: Each query starts fresh with no building upon previous insights or discoveries.
Pre-Synthesis Approach
The llm-wiki-pattern represents a fundamental alternative where:
Knowledge Integration: Instead of retrieving fragments, LLMs incrementally build and maintain structured knowledge representations that integrate information across sources.
Persistent Artifacts: Cross-references, contradictions, and syntheses are identified once and maintained rather than rediscovered on each query.
Compounding Intelligence: Good questions and discoveries get filed back into the knowledge base, creating compounding-artifacts that become smarter over time.
Architectural Differences
Traditional RAG
Sources → Embedding → Vector DB → Retrieval → LLM Generation
Wiki Pattern Alternative
Sources → LLM Integration → Persistent Wiki → Query → Pre-Synthesized Answers
Key Advantages
Deep Synthesis: Can develop sophisticated understanding that builds across multiple sources and interactions.
Context Preservation: Maintains full context of how understanding developed rather than working with isolated fragments.
Knowledge Evolution: The system gets smarter over time rather than remaining static.
Rich Cross-Referencing: Connections between concepts are discovered and maintained automatically.
Implementation Patterns
Incremental Integration
New sources update existing entity pages, revise concept summaries, and note contradictions rather than simply being added to a retrieval corpus.
Pre-Compiled Knowledge
Answers draw from knowledge that has already been synthesized and cross-referenced rather than being generated from raw retrieval.
Active Maintenance
Systems actively identify gaps, contradictions, and opportunities for deeper synthesis rather than passively serving queries.
Use Cases
Particularly effective for:
- Long-term research projects requiring deep synthesis
- Personal knowledge development over months/years
- Complex domains where understanding evolves with exposure
- Business intelligence requiring integrated analysis
- Any context where knowledge should compound rather than remain fragmented
Trade-offs
Computational Investment: Requires upfront processing to build integrated knowledge structures rather than simple indexing.
Maintenance Complexity: Systems must actively maintain consistency and handle contradictions rather than serving static retrievals.
Domain Specificity: Requires careful schema design and workflow definition for specific knowledge domains.
Future Directions
The success of wiki-pattern approaches suggests broader possibilities for AI systems that build persistent, evolving knowledge representations rather than operating in stateless fashion.
See also
RAG Alternatives
page dédiée →Approaches to knowledge management and information retrieval that move beyond traditional Retrieval-Augmented Generation (RAG) systems. Most notably exemplified by andrej-karpathy's llm-wiki-pattern which treats knowledge as compounding-artifacts rather than static document collections.
Traditional RAG Limitations
Standard RAG Approach:
- Upload collection of files
- LLM retrieves relevant chunks at query time
- Generates answers from fragments
- Rediscovers knowledge from scratch on every question
- No accumulation or synthesis between queries
Problems:
- Subtle questions requiring synthesis across multiple documents must be solved repeatedly
- No building up of understanding over time
- Connections between sources not maintained
- Context limited to what can be retrieved in single query
LLM Wiki Pattern Alternative
Core Difference: Instead of retrieving from raw documents at query time, the LLM incrementally builds and maintains a persistent wiki that sits between user and sources.
Process:
- New sources integrate into existing wiki structure
- Updates entity pages, revises summaries, notes contradictions
- Cross-references already established
- Synthesis reflects all previous learning
- Knowledge compounds with each addition
Key Advantages
Persistent Knowledge: Cross-references exist, contradictions flagged, synthesis current. Wiki keeps getting richer with every source and question.
Maintenance Automation: LLMs handle tedious bookkeeping - updating cross-references, keeping summaries current, maintaining consistency. Humans focus on curation and questions.
Compound Learning: Good answers become new wiki pages. Explorations compound in knowledge base just like ingested sources.
Other Alternative Approaches
Knowledge Graphs: Structured representation of entities and relationships, but typically requires significant manual curation or complex extraction pipelines.
Embedding-Based Memory Systems: Vector representations of experiences/documents that can be retrieved by similarity, but lack explicit structure and cross-referencing.
Agent Memory Architectures: Various approaches to giving AI systems persistent memory, though most focus on conversation history rather than structured knowledge building.
Implementation Considerations
Moving beyond RAG requires:
- Schema Design: Clear workflows and conventions for knowledge maintenance
- Integration Workflows: Systematic processes for incorporating new information
- Quality Control: Methods for ensuring accuracy and consistency
- Navigation Systems: Tools for exploring and searching the knowledge base
Applications
RAG alternatives particularly valuable for:
- Research Synthesis: Building understanding across many sources over time
- Personal Learning: Accumulating knowledge in specific domains
- Business Intelligence: Maintaining current understanding of competitive landscape
- Domain Expertise: Building deep, interconnected knowledge in specialized areas
See also
Retrieval Augmented Generation (RAG)
page dédiée →A technique that enhances language model responses by retrieving relevant information from external knowledge bases at query time. The retrieved context is then used to ground the model's generation, reducing hallucinations and enabling access to information beyond the training data.
Core Process
- Query Processing: User question is converted to embeddings
- Retrieval: Similar document chunks are found using vector search
- Context Formation: Retrieved chunks are assembled as context
- Generation: LLM generates response using retrieved context
- Response: Final answer combines retrieved facts with model knowledge
Architecture Components
Vector Store
- Document embeddings stored in vector database (chroma, pinecone, weaviate)
- Enables semantic similarity search
- Supports hybrid search combining dense and sparse retrieval
Embedding Models
- Convert text to high-dimensional vectors
- Popular options: OpenAI Ada-002, sentence-transformers, Cohere
- Critical for retrieval quality
Chunking Strategy
- Documents split into manageable segments
- Balance between context preservation and retrieval precision
- Common approaches: fixed-size, semantic, hierarchical
Advanced Patterns
Multi-Query RAG
Generate multiple query variants to improve retrieval coverage:
def multi_query_rag(question):
variants = llm.generate_query_variants(question)
all_chunks = []
for variant in variants:
chunks = vector_store.search(variant)
all_chunks.extend(chunks)
return llm.generate(question, dedupe(all_chunks))
rag-fusion
Combines multiple retrieval strategies and reranks results for improved relevance.
conversational-rag
Maintains conversation history to enable multi-turn interactions with context awareness.
agentic-rag
Uses AI agents to orchestrate complex retrieval workflows, including tool use and multi-step reasoning.
Limitations
Knowledge Rediscovery Problem
RAG systems rediscover knowledge from scratch on every query. As andrej-karpathy notes in the llm-wiki-pattern, this prevents knowledge accumulation - subtle questions requiring synthesis across multiple documents must repeatedly piece together fragments without building persistent understanding.
Context Window Constraints
- Limited by model's maximum context length
- Must balance breadth vs. depth of retrieved information
- May miss relevant information if not in top-k results
Retrieval Quality Issues
- Embedding similarity doesn't always match semantic relevance
- Struggles with concepts spanning multiple documents
- May retrieve contradictory information without resolution
Alternative Approaches
llm-wiki-pattern
Instead of retrieving from raw documents, LLMs maintain persistent wikis that incrementally integrate knowledge over time. This creates compounding-artifacts where cross-references exist permanently and contradictions are pre-resolved.
knowledge-graphs
Structured representation of entities and relationships can provide more precise retrieval paths than vector similarity.
Use Cases
- Document Q&A: Customer support, internal documentation
- Research Assistance: Scientific literature review, fact-checking
- Content Generation: Blog writing with factual grounding
- Educational Tools: Personalized tutoring with curriculum materials
Implementation Considerations
Evaluation Metrics
- Retrieval Accuracy: Precision/recall of relevant documents
- Answer Quality: Factual correctness, completeness, relevance
- Latency: End-to-end response time
- Cost: Embedding generation and storage costs
Production Challenges
- Keeping knowledge base current with new information
- Handling contradictory or outdated information
- Scaling retrieval performance with growing document corpus
- Managing embedding model updates and reindexing
See also
- vector-databases
- embedding-models
- agentic-rag
- conversational-rag
- llm-wiki-pattern
- knowledge-graphs
Schema Coevolution
page dédiée →The collaborative development process between human and LLM for evolving the configuration layer of llm-wiki-pattern systems. The schema document (CLAUDE.md, AGENTS.md, etc.) grows and adapts based on actual usage patterns, domain needs, and discovered workflows.
Core Concept
Unlike static system configurations, the schema in an LLM wiki system evolves with use. As humans and LLMs work together, they discover:
- Effective workflows for the specific domain
- Useful page formats and structures
- Domain-specific conventions and tags
- Optimal ingest and query patterns
These discoveries get documented in the schema, making the LLM a more effective collaborator over time.
Evolution Process
Initial Schema
- Basic structure and conventions
- Generic workflows from the pattern
- Minimal domain-specific customization
Usage-Driven Refinement
- Workflow optimization: Discovering efficient ingest patterns
- Convention standardization: Establishing consistent tagging and formatting
- Domain adaptation: Adding field-specific page types and structures
- Tool integration: Incorporating discovered utilities and scripts
Continuous Improvement
- Regular schema updates based on what works
- Documentation of effective practices
- Removal of unused conventions
- Addition of new capabilities
Human-LLM Collaboration
Human Contributions
- Domain expertise: Understanding field-specific needs
- Workflow preferences: Preferred interaction patterns
- Quality standards: Defining acceptable outputs
- Strategic direction: Long-term knowledge goals
LLM Contributions
- Pattern recognition: Identifying recurring structures
- Consistency maintenance: Ensuring schema adherence
- Workflow execution: Following documented procedures
- Improvement suggestions: Proposing optimizations
Schema Components That Evolve
Page Types
- Standard templates for different content types
- Domain-specific entity categories
- Specialized analysis formats (comparisons, timelines, etc.)
Tagging Systems
- Hierarchical tag structures for the domain
- Consistent naming conventions
- Cross-reference patterns
Workflows
- Detailed ingest procedures
- Query and synthesis patterns
- Maintenance and lint operations
- Quality control checkpoints
Tool Integration
- CLI utilities and their usage patterns
- Search and navigation tools
- Export and visualization formats
- Integration with external systems
Benefits
Adaptive Systems
Schema coevolution enables knowledge systems that improve with use rather than becoming stale or rigid.
Domain Optimization
Over time, the system becomes specifically tuned to the user's field and preferences rather than remaining generic.
Reduced Friction
Well-evolved schemas reduce cognitive overhead by codifying effective practices and eliminating decision fatigue.
Knowledge Transfer
The schema serves as documentation of effective knowledge management practices that can be shared or adapted.
Example Evolution
Initial: "Create entity pages for people mentioned"
Evolved: "Create entity pages with standardized sections: Background, Key Ideas, Publications, Influence, Cross-References. Tag with domain (ai-researcher, entrepreneur, academic) and confidence level."
Challenges
Over-Specification
Schemas can become too rigid, constraining useful variation and experimentation.
Version Management
Managing schema changes while maintaining consistency across existing wiki content.
Complexity Growth
Balancing comprehensive documentation with usability for both human and LLM.
Implementation Patterns
Versioned Schemas
Track schema evolution with version control to understand what changes improve effectiveness.
Modular Structure
Organize schema into sections (workflows, conventions, formats) that can evolve independently.
Example-Driven Documentation
Include concrete examples in schema documentation to clarify abstract conventions.
See also
- llm-wiki-pattern - System using schema coevolution
- three-layer-architecture - Schema as configuration layer
- adaptive-systems - Systems that improve with use
- workflow-optimization - Process improvement patterns
Wiki Indexing Patterns
page dédiée →Systematic approaches to organizing and navigating knowledge bases, particularly in llm-wiki-pattern implementations. Two primary patterns serve different navigation needs: content-oriented catalogs and chronological logs.
Core Patterns
Content-Oriented Index (index.md)
Purpose: Catalog all wiki content organized by category and topic.
Structure:
- Links to all wiki pages with one-line summaries
- Organized by category (entities, concepts, sources, syntheses)
- Optional metadata (dates, source counts, confidence levels)
- Updated on every ingest operation
Navigation Model: LLM reads index first to identify relevant pages for queries, then drills into specific content. Works effectively at moderate scale (~100 sources, hundreds of pages) without requiring embedding-based infrastructure.
Example Structure:
## Entities (45 pages)
- openai - AI research company developing GPT models
- anthropic - AI safety company behind Claude
- andrej-karpathy - Former Tesla AI director, advocate of LLM wiki pattern
## Concepts (89 pages)
- [retrieval-augmented-generation](/concepts/retrieval-augmented-generation) - Combining retrieval with generation
- Vector Embeddings - Dense numerical representations of data
Chronological Log (log.md)
Purpose: Append-only record of all wiki operations and evolution.
Structure:
- Timestamped entries for ingests, queries, lint operations
- Consistent format enabling programmatic parsing
- Timeline of wiki evolution and recent activity
- Context for understanding system state
Navigation Model: Understand what's been done recently, track system evolution, provide context for LLM about recent operations.
Example Format:
## [2026-12-21 14:30] ingest | Karpathy LLM Wiki Pattern
- **Source**: raw/articles/karpathy-llm-wiki-pattern.md
- **Pages touched**: [llm-wiki-pattern](/concepts/llm-wiki-pattern), [compounding-artifacts](/concepts/compounding-artifacts), [persistent-learning](/concepts/persistent-learning)
- **Summary**: Ingested comprehensive blueprint for LLM-maintained knowledge bases
Design Principles
Complementary Functions
The two patterns serve different but complementary navigation needs:
- Index: "What knowledge exists and where is it?"
- Log: "How did this knowledge base evolve over time?"
Parseable Formats
Both use consistent formatting that enables programmatic access:
- Index supports category-based filtering and summary generation
- Log enables timeline analysis with simple Unix tools (
grep "^## \[" log.md | tail -5)
LLM-Maintained
Both files are automatically maintained by LLMs during normal operations, eliminating human maintenance burden while providing essential navigation capabilities.
Scaling Considerations
Moderate Scale Effectiveness
Content-oriented indexing works well up to hundreds of pages without requiring sophisticated search infrastructure. The human-readable format provides good overview while remaining LLM-navigable.
Search Enhancement
At larger scales, can be supplemented with search tools:
- Local search engines like qmd for hybrid BM25/vector search
- Full-text indexing for content that exceeds index-based navigation
- MCP servers for programmatic access to search capabilities
Hierarchical Organization
Index structure can evolve to use subcategories and nested organization as content volume grows:
### AI/ML Core (89 pages)
### RAG & Search (47 pages)
### Agent Systems (112 pages)
Implementation Patterns
Automatic Updates
Index updated during every ingest operation as LLM processes new sources and creates/updates wiki pages. Ensures index remains current with no human intervention.
Metadata Integration
Index can include metadata from page frontmatter:
- Creation/update dates
- Source counts indicating how well-researched topics are
- Confidence levels for content reliability
- Tag summaries for topic clustering
Cross-Reference Support
Index serves as hub for cross-reference discovery - LLM uses it to find related pages when updating content or answering queries requiring synthesis across topics.
Alternatives and Extensions
Graph-Based Navigation
Tools like Obsidian's graph view provide visual representation of page connections, complementing text-based index navigation.
Dynamic Queries
Obsidian's Dataview plugin can generate dynamic indexes based on page metadata, automatically categorizing content based on tags or other frontmatter fields.
Search Integration
Can be combined with full-text search engines for complex queries while maintaining human-readable overview structure.
See also
- llm-wiki-pattern
- Content Organization
- Knowledge Base Navigation
- Information Architecture