~/wiki

Intelligent Content Triage

Confiance : high
content-triageautomationrelevance-scoringknowledge-managementllm-filteringquality-assessmentcontent-pipelinesource-evaluationhigh-medium-low-scoringcompound-learningscreenshot-filteringurl-evaluationdomain-relevancetelegram-integrationdiscord-preventionquality-gatesmeta-learningreal-time-scoringfrench-developer-workflow

Automated system for evaluating and routing incoming content based on relevance, quality, and potential value to knowledge base. Essential component of scalable knowledge management systems to prevent information overload while ensuring valuable content is captured and processed.

Triage Framework

Three-Tier Scoring System

HIGH (Auto-ingest):

  • Introduces concepts/techniques not yet in wiki
  • Contradicts or significantly nuances existing content
  • Directly relates to active projects
  • From high-quality sources (papers, official docs, recognized authors)
  • Contains actionable code patterns or architecture decisions

MEDIUM (Manual review queue):

  • Potentially relevant but uncertain value
  • Adjacent to current domains but not core
  • Interesting but shallow content that might need deeper sourcing
  • Routed to raw/pending/ folder for later evaluation

LOW (Archive/discard):

  • Not related to domains of interest
  • Too superficial (3-line social media posts without substance)
  • Duplicate of already-ingested content
  • Promotional content disguised as technical content
  • Outdated information already superseded in wiki

Implementation Architecture

Screenshot Filtering Pipeline

For iPhone social media screenshots - most common source requiring filtering:

  1. Quick Domain Check: "Is this tech/AI related?" - fast binary filter
  2. Content Analysis: OCR + vision model text extraction
  3. Relevance Scoring: Compare against existing wiki index
  4. Quality Assessment: Source credibility, depth, actionability
  5. Routing Decision: HIGH → ingest, MEDIUM → pending, LOW → discard

Automated Quality Gates

Source Credibility Factors:

  • Author reputation in AI/tech domains
  • Publication venue quality (arXiv, official docs, conferences)
  • Content depth and technical specificity
  • Presence of code examples or concrete implementations

Relevance Criteria:

  • Alignment with current project portfolio
  • Overlap assessment with existing wiki content
  • Introduction of genuinely new concepts vs. rehashed basics
  • Potential for compound learning and cross-referencing

Compound Learning Effect

The triage system improves over time through knowledge accumulation:

  • Basic Transformer articles scored LOW when wiki already contains 15 related pages
  • Novel techniques (Flash Attention, MoE architectures) scored HIGH
  • Scoring accuracy increases as wiki index grows more comprehensive
  • Domain understanding deepens, enabling better quality assessment

Real-World Performance

Live Implementation Results (June 2026):

  • Karpathy's LLM wiki document: HIGH score (correctly identified as authoritative, novel)
  • Automatic routing prevented manual review overhead
  • Generated 5 wiki pages + 5 flashcards from single high-quality source
  • Zero false positives in initial testing phase

Integration Patterns

Telegram Bot Integration

Direct scoring from mobile messaging:

  • Send URL: bot fetches, scores, routes appropriately
  • Send screenshot: vision model extracts text, scores content
  • Real-time feedback on triage decisions
  • Manual override capability for edge cases

Multi-Source Adaptation

Different scoring criteria per source type:

  • Academic papers: Novelty + technical depth
  • Project documentation: Direct applicability + completeness
  • Social media: Signal-to-noise ratio + source authority
  • Audio transcripts: Actionable insights + speaker expertise

Technical Implementation

LLM Prompt Structure:

# Compare against existing wiki index
# Apply domain relevance criteria from SCHEMA.md
# Score using established quality gates
# Return: HIGH/MEDIUM/LOW + reasoning

File System Routing:

  • raw/sources/ → HIGH content for immediate processing
  • raw/pending/ → MEDIUM content awaiting manual review
  • raw/discarded/ → LOW content archived with reasoning

This systematic approach enables high-volume content processing while maintaining knowledge base quality, essential for ai-engineering-wiki automation at scale.

See also