Intelligent Content Triage
Automated system for evaluating and routing incoming content based on relevance, quality, and potential value to knowledge base. Essential component of scalable knowledge management systems to prevent information overload while ensuring valuable content is captured and processed.
Triage Framework
Three-Tier Scoring System
HIGH (Auto-ingest):
- Introduces concepts/techniques not yet in wiki
- Contradicts or significantly nuances existing content
- Directly relates to active projects
- From high-quality sources (papers, official docs, recognized authors)
- Contains actionable code patterns or architecture decisions
MEDIUM (Manual review queue):
- Potentially relevant but uncertain value
- Adjacent to current domains but not core
- Interesting but shallow content that might need deeper sourcing
- Routed to
raw/pending/folder for later evaluation
LOW (Archive/discard):
- Not related to domains of interest
- Too superficial (3-line social media posts without substance)
- Duplicate of already-ingested content
- Promotional content disguised as technical content
- Outdated information already superseded in wiki
Implementation Architecture
Screenshot Filtering Pipeline
For iPhone social media screenshots - most common source requiring filtering:
- Quick Domain Check: "Is this tech/AI related?" - fast binary filter
- Content Analysis: OCR + vision model text extraction
- Relevance Scoring: Compare against existing wiki index
- Quality Assessment: Source credibility, depth, actionability
- Routing Decision: HIGH → ingest, MEDIUM → pending, LOW → discard
Automated Quality Gates
Source Credibility Factors:
- Author reputation in AI/tech domains
- Publication venue quality (arXiv, official docs, conferences)
- Content depth and technical specificity
- Presence of code examples or concrete implementations
Relevance Criteria:
- Alignment with current project portfolio
- Overlap assessment with existing wiki content
- Introduction of genuinely new concepts vs. rehashed basics
- Potential for compound learning and cross-referencing
Compound Learning Effect
The triage system improves over time through knowledge accumulation:
- Basic Transformer articles scored LOW when wiki already contains 15 related pages
- Novel techniques (Flash Attention, MoE architectures) scored HIGH
- Scoring accuracy increases as wiki index grows more comprehensive
- Domain understanding deepens, enabling better quality assessment
Real-World Performance
Live Implementation Results (June 2026):
- Karpathy's LLM wiki document: HIGH score (correctly identified as authoritative, novel)
- Automatic routing prevented manual review overhead
- Generated 5 wiki pages + 5 flashcards from single high-quality source
- Zero false positives in initial testing phase
Integration Patterns
Telegram Bot Integration
Direct scoring from mobile messaging:
- Send URL: bot fetches, scores, routes appropriately
- Send screenshot: vision model extracts text, scores content
- Real-time feedback on triage decisions
- Manual override capability for edge cases
Multi-Source Adaptation
Different scoring criteria per source type:
- Academic papers: Novelty + technical depth
- Project documentation: Direct applicability + completeness
- Social media: Signal-to-noise ratio + source authority
- Audio transcripts: Actionable insights + speaker expertise
Technical Implementation
LLM Prompt Structure:
# Compare against existing wiki index
# Apply domain relevance criteria from SCHEMA.md
# Score using established quality gates
# Return: HIGH/MEDIUM/LOW + reasoning
File System Routing:
raw/sources/→ HIGH content for immediate processingraw/pending/→ MEDIUM content awaiting manual reviewraw/discarded/→ LOW content archived with reasoning
This systematic approach enables high-volume content processing while maintaining knowledge base quality, essential for ai-engineering-wiki automation at scale.
See also
- multi-source-ingestion
- ai-engineering-wiki
- content-quality-assessment
- llm-wiki-pattern