Wiki Indexing Patterns
Systematic approaches to organizing and navigating knowledge bases, particularly in llm-wiki-pattern implementations. Two primary patterns serve different navigation needs: content-oriented catalogs and chronological logs.
Core Patterns
Content-Oriented Index (index.md)
Purpose: Catalog all wiki content organized by category and topic.
Structure:
- Links to all wiki pages with one-line summaries
- Organized by category (entities, concepts, sources, syntheses)
- Optional metadata (dates, source counts, confidence levels)
- Updated on every ingest operation
Navigation Model: LLM reads index first to identify relevant pages for queries, then drills into specific content. Works effectively at moderate scale (~100 sources, hundreds of pages) without requiring embedding-based infrastructure.
Example Structure:
## Entities (45 pages)
- openai - AI research company developing GPT models
- anthropic - AI safety company behind Claude
- andrej-karpathy - Former Tesla AI director, advocate of LLM wiki pattern
## Concepts (89 pages)
- [retrieval-augmented-generation](/concepts/retrieval-augmented-generation) - Combining retrieval with generation
- Vector Embeddings - Dense numerical representations of data
Chronological Log (log.md)
Purpose: Append-only record of all wiki operations and evolution.
Structure:
- Timestamped entries for ingests, queries, lint operations
- Consistent format enabling programmatic parsing
- Timeline of wiki evolution and recent activity
- Context for understanding system state
Navigation Model: Understand what's been done recently, track system evolution, provide context for LLM about recent operations.
Example Format:
## [2026-12-21 14:30] ingest | Karpathy LLM Wiki Pattern
- **Source**: raw/articles/karpathy-llm-wiki-pattern.md
- **Pages touched**: [llm-wiki-pattern](/concepts/llm-wiki-pattern), [compounding-artifacts](/concepts/compounding-artifacts), [persistent-learning](/concepts/persistent-learning)
- **Summary**: Ingested comprehensive blueprint for LLM-maintained knowledge bases
Design Principles
Complementary Functions
The two patterns serve different but complementary navigation needs:
- Index: "What knowledge exists and where is it?"
- Log: "How did this knowledge base evolve over time?"
Parseable Formats
Both use consistent formatting that enables programmatic access:
- Index supports category-based filtering and summary generation
- Log enables timeline analysis with simple Unix tools (
grep "^## \[" log.md | tail -5)
LLM-Maintained
Both files are automatically maintained by LLMs during normal operations, eliminating human maintenance burden while providing essential navigation capabilities.
Scaling Considerations
Moderate Scale Effectiveness
Content-oriented indexing works well up to hundreds of pages without requiring sophisticated search infrastructure. The human-readable format provides good overview while remaining LLM-navigable.
Search Enhancement
At larger scales, can be supplemented with search tools:
- Local search engines like qmd for hybrid BM25/vector search
- Full-text indexing for content that exceeds index-based navigation
- MCP servers for programmatic access to search capabilities
Hierarchical Organization
Index structure can evolve to use subcategories and nested organization as content volume grows:
### AI/ML Core (89 pages)
### RAG & Search (47 pages)
### Agent Systems (112 pages)
Implementation Patterns
Automatic Updates
Index updated during every ingest operation as LLM processes new sources and creates/updates wiki pages. Ensures index remains current with no human intervention.
Metadata Integration
Index can include metadata from page frontmatter:
- Creation/update dates
- Source counts indicating how well-researched topics are
- Confidence levels for content reliability
- Tag summaries for topic clustering
Cross-Reference Support
Index serves as hub for cross-reference discovery - LLM uses it to find related pages when updating content or answering queries requiring synthesis across topics.
Alternatives and Extensions
Graph-Based Navigation
Tools like Obsidian's graph view provide visual representation of page connections, complementing text-based index navigation.
Dynamic Queries
Obsidian's Dataview plugin can generate dynamic indexes based on page metadata, automatically categorizing content based on tags or other frontmatter fields.
Search Integration
Can be combined with full-text search engines for complex queries while maintaining human-readable overview structure.
See also
- llm-wiki-pattern
- Content Organization
- Knowledge Base Navigation
- Information Architecture