Model Calibration
Quality measurement for language models based on how well their predicted probabilities align with actual correctness. A well-calibrated model assigns highest probabilities to correct answers.
Definition
Model calibration measures whether a model's confidence scores accurately reflect its likelihood of being correct. In well-calibrated models, when the model assigns 80% probability to an answer, it should be correct approximately 80% of the time.
Evaluation Approach
Calibration is assessed through log-likelihood evaluations:
- Calculate probability distributions over answer choices
- Compare predicted probabilities with actual correctness
- Analyze whether highest probability answers are indeed correct most often
Significance
Quality Indicator: Better calibrated models provide more reliable uncertainty estimates, crucial for applications where confidence matters.
Practical Applications: Important for systems that need to know when they're uncertain and should abstain from answering or seek human input.
Measurement Context
Typically evaluated using multiple choice or cloze formulation tasks where ground truth is available. Log-likelihood evaluation methods enable direct probability comparison.
Related Concepts
Part of broader model reliability assessment alongside accuracy metrics. Complements traditional performance measures by adding uncertainty quantification dimension.