~/wiki

LLM Reliability

Mis à jour le 2025-12-30Confiance : high
llm-reliabilityuncertainty-acknowledgmentfactual-accuracyhallucination-detectionmodel-confidenceevaluation-frameworksextrinsic-hallucinationin-context-hallucination

The ability of Large Language Models to provide accurate, consistent, and trustworthy outputs while appropriately expressing uncertainty when knowledge is incomplete or confidence is low. A critical aspect of responsible AI deployment that encompasses multiple dimensions of model behavior.

Core Components

Factual Accuracy

Models must generate content grounded in verifiable knowledge when they possess it, avoiding extrinsic-hallucination where output is fabricated rather than based on training data or world knowledge.

Context Faithfulness

Models must remain consistent with provided source content, avoiding in-context-hallucination where outputs contradict explicit contextual information.

Uncertainty Acknowledgment

As lilian-weng emphasizes, "when the model does not know about a fact, it should say so" rather than fabricating confident-sounding but ungrounded responses.

Hallucination Taxonomy

lilian-weng's precise taxonomy distinguishes:

This framework helps separate genuine hallucination (fabricated content) from general model mistakes.

Assessment Challenges

  • Context-Based: In-context hallucination can be verified against provided information
  • Knowledge-Based: Extrinsic hallucination requires verification against world knowledge, which is "too expensive to retrieve and identify conflicts per generation"

Requirements for Reliable Systems

To achieve reliability, LLMs must:

  1. Generate factual content when knowledge exists
  2. Practice uncertainty-acknowledgment when knowledge is insufficient
  3. Maintain consistency with provided context
  4. Distinguish between different types of potential failures

See also