Statistical Significance
Confiance : high
statistical-significancehypothesis-testingp-valuesconfidence-intervalsstatistical-powereffect-sizemultiple-comparisonsbayesian-analysis
The likelihood that observed differences in model performance are due to genuine differences in capabilities rather than random variation. A fundamental concept in rigorous llm-evaluation that ensures reliable conclusions about model comparisons and performance claims.
Core Concepts
Hypothesis Testing Framework
- Null Hypothesis (H₀): Assumption that models perform equally well
- Alternative Hypothesis (H₁): Assumption that models differ in performance
- Test Statistic: Quantitative measure of difference between models
- P-Value: Probability of observing results at least as extreme under null hypothesis
Type I and Type II Errors
- Type I Error (α): Falsely concluding models differ when they don't (false positive)
- Type II Error (β): Failing to detect genuine differences (false negative)
- Statistical Power (1-β): Probability of correctly detecting true differences
- Effect Size: Magnitude of difference between models
Statistical Tests for Model Evaluation
Comparison Tests
Paired t-Test
- Compares mean performance across matched test instances
- Assumes normally distributed performance differences
- Appropriate for continuous metrics (accuracy, F1 score)
- Accounts for instance-level correlation between models
Wilcoxon Signed-Rank Test
- Non-parametric alternative to paired t-test
- Robust to non-normal distributions
- Based on rank differences rather than raw values
- Suitable for ordinal or skewed performance metrics
McNemar's Test
- Specialized for binary classification accuracy comparison
- Focuses on disagreement cases between models
- Accounts for marginal performance similarities
- Commonly used in NLP evaluation
Multiple Comparison Corrections
Bonferroni Correction
- Adjusts significance threshold for multiple tests
- Conservative approach: α_adjusted = α / number_of_tests
- Controls family-wise error rate
- May be overly conservative with many comparisons
False Discovery Rate (FDR)
- Controls expected proportion of false discoveries
- Less conservative than Bonferroni correction
- Benjamini-Hochberg procedure for FDR control
- Better statistical power for exploratory analysis
Sample Size and Power Analysis
Power Calculation
- Minimum Effect Size: Smallest difference worth detecting
- Desired Power: Typically 0.8 (80% chance of detecting true effects)
- Significance Level: Usually 0.05 (5% false positive rate)
- Sample Size: Number of test instances required
Effect Size Measures
Cohen's d
- Standardized difference between group means
- Small (0.2), medium (0.5), large (0.8) effect sizes
- Useful for comparing across different metrics
- Independent of sample size
Practical Significance
- Business or scientific importance of observed differences
- May differ from statistical significance
- Considers cost-benefit tradeoffs
- Domain-specific thresholds
Confidence Intervals
Interpretation
- Range of plausible values for true performance difference
- Quantifies uncertainty around point estimates
- More informative than p-values alone
- Enables effect size assessment
Bootstrap Confidence Intervals
- Non-parametric approach using resampling
- Robust to distribution assumptions
- Suitable for complex evaluation metrics
- Accounts for evaluation uncertainty
Bayesian Credible Intervals
- Probability that true parameter lies in interval
- Incorporates prior knowledge
- More intuitive interpretation than frequentist intervals
- Natural handling of uncertainty propagation
Common Pitfalls
Multiple Testing Issues
- P-Hacking: Selective reporting of significant results
- Data Snooping: Repeated testing until significance is found
- Optional Stopping: Ending experiments when significance is reached
- Subgroup Analysis: Post-hoc testing of population subsets
Misinterpretation of P-Values
- P-value ≠ probability that null hypothesis is true
- Statistical significance ≠ practical importance
- Non-significance ≠ evidence for null hypothesis
- P-values depend on sample size, not just effect size
Inadequate Sample Sizes
- Underpowered studies failing to detect true differences
- Overreliance on pilot studies with small samples
- Publication bias favoring significant results
- Overestimation of effect sizes in underpowered studies
Best Practices
Experimental Design
Pre-Registration
- Specify hypotheses and analysis plan before data collection
- Prevents p-hacking and selective reporting
- Increases confidence in reported results
- Standard practice in medical research
Power Analysis
- Calculate required sample size before evaluation
- Consider practical constraints and resources
- Report power calculations in methodology
- Discuss implications of achieved power
Result Reporting
Complete Statistical Reporting
- Report effect sizes alongside p-values
- Include confidence intervals for all estimates
- Describe statistical methods and assumptions
- Report negative results and non-significant findings
Uncertainty Quantification
- Bootstrap or cross-validation for robust estimates
- Multiple random seeds for evaluation stability
- Sensitivity analysis for key assumptions
- Discussion of limitations and threats to validity
Advanced Topics
Bayesian Model Comparison
- Model evidence and Bayes factors
- Posterior probability of model superiority
- Hierarchical modeling for multiple tasks
- Integration of prior knowledge
Non-Parametric Methods
- Permutation tests for hypothesis testing
- Rank-based methods for ordinal data
- Distribution-free confidence intervals
- Robust statistics for outlier resilience
Meta-Analysis
- Combining results across multiple studies
- Fixed-effects vs. random-effects models
- Assessment of heterogeneity
- Publication bias detection and correction
Practical Implementation
Statistical Software
- R: comprehensive statistical analysis capabilities
- Python: scipy.stats, statsmodels, scikit-learn
- Specialized packages: pingouin, bootstrapped
- Bayesian tools: PyMC3, Stan
Evaluation Workflows
from scipy import stats
import numpy as np
def compare_models(scores_a, scores_b):
# Paired t-test for mean comparison
t_stat, p_val = stats.ttest_rel(scores_a, scores_b)
# Effect size calculation
diff = np.array(scores_a) - np.array(scores_b)
cohen_d = np.mean(diff) / np.std(diff)
# Bootstrap confidence interval
n_bootstrap = 10000
bootstrap_diffs = []
for _ in range(n_bootstrap):
indices = np.random.choice(len(scores_a), len(scores_a))
boot_diff = np.mean(np.array(scores_a)[indices] - np.array(scores_b)[indices])
bootstrap_diffs.append(boot_diff)
ci_lower = np.percentile(bootstrap_diffs, 2.5)
ci_upper = np.percentile(bootstrap_diffs, 97.5)
return {
'p_value': p_val,
'effect_size': cohen_d,
'confidence_interval': (ci_lower, ci_upper)
}
Related Concepts
- llm-evaluation
- evaluation-frameworks
- benchmarking
- evaluation-metrics
- hypothesis-testing
- experimental-design