~/wiki

Statistical Significance

Confiance : high
statistical-significancehypothesis-testingp-valuesconfidence-intervalsstatistical-powereffect-sizemultiple-comparisonsbayesian-analysis

The likelihood that observed differences in model performance are due to genuine differences in capabilities rather than random variation. A fundamental concept in rigorous llm-evaluation that ensures reliable conclusions about model comparisons and performance claims.

Core Concepts

Hypothesis Testing Framework

  • Null Hypothesis (H₀): Assumption that models perform equally well
  • Alternative Hypothesis (H₁): Assumption that models differ in performance
  • Test Statistic: Quantitative measure of difference between models
  • P-Value: Probability of observing results at least as extreme under null hypothesis

Type I and Type II Errors

  • Type I Error (α): Falsely concluding models differ when they don't (false positive)
  • Type II Error (β): Failing to detect genuine differences (false negative)
  • Statistical Power (1-β): Probability of correctly detecting true differences
  • Effect Size: Magnitude of difference between models

Statistical Tests for Model Evaluation

Comparison Tests

Paired t-Test

  • Compares mean performance across matched test instances
  • Assumes normally distributed performance differences
  • Appropriate for continuous metrics (accuracy, F1 score)
  • Accounts for instance-level correlation between models

Wilcoxon Signed-Rank Test

  • Non-parametric alternative to paired t-test
  • Robust to non-normal distributions
  • Based on rank differences rather than raw values
  • Suitable for ordinal or skewed performance metrics

McNemar's Test

  • Specialized for binary classification accuracy comparison
  • Focuses on disagreement cases between models
  • Accounts for marginal performance similarities
  • Commonly used in NLP evaluation

Multiple Comparison Corrections

Bonferroni Correction

  • Adjusts significance threshold for multiple tests
  • Conservative approach: α_adjusted = α / number_of_tests
  • Controls family-wise error rate
  • May be overly conservative with many comparisons

False Discovery Rate (FDR)

  • Controls expected proportion of false discoveries
  • Less conservative than Bonferroni correction
  • Benjamini-Hochberg procedure for FDR control
  • Better statistical power for exploratory analysis

Sample Size and Power Analysis

Power Calculation

  • Minimum Effect Size: Smallest difference worth detecting
  • Desired Power: Typically 0.8 (80% chance of detecting true effects)
  • Significance Level: Usually 0.05 (5% false positive rate)
  • Sample Size: Number of test instances required

Effect Size Measures

Cohen's d

  • Standardized difference between group means
  • Small (0.2), medium (0.5), large (0.8) effect sizes
  • Useful for comparing across different metrics
  • Independent of sample size

Practical Significance

  • Business or scientific importance of observed differences
  • May differ from statistical significance
  • Considers cost-benefit tradeoffs
  • Domain-specific thresholds

Confidence Intervals

Interpretation

  • Range of plausible values for true performance difference
  • Quantifies uncertainty around point estimates
  • More informative than p-values alone
  • Enables effect size assessment

Bootstrap Confidence Intervals

  • Non-parametric approach using resampling
  • Robust to distribution assumptions
  • Suitable for complex evaluation metrics
  • Accounts for evaluation uncertainty

Bayesian Credible Intervals

  • Probability that true parameter lies in interval
  • Incorporates prior knowledge
  • More intuitive interpretation than frequentist intervals
  • Natural handling of uncertainty propagation

Common Pitfalls

Multiple Testing Issues

  • P-Hacking: Selective reporting of significant results
  • Data Snooping: Repeated testing until significance is found
  • Optional Stopping: Ending experiments when significance is reached
  • Subgroup Analysis: Post-hoc testing of population subsets

Misinterpretation of P-Values

  • P-value ≠ probability that null hypothesis is true
  • Statistical significance ≠ practical importance
  • Non-significance ≠ evidence for null hypothesis
  • P-values depend on sample size, not just effect size

Inadequate Sample Sizes

  • Underpowered studies failing to detect true differences
  • Overreliance on pilot studies with small samples
  • Publication bias favoring significant results
  • Overestimation of effect sizes in underpowered studies

Best Practices

Experimental Design

Pre-Registration

  • Specify hypotheses and analysis plan before data collection
  • Prevents p-hacking and selective reporting
  • Increases confidence in reported results
  • Standard practice in medical research

Power Analysis

  • Calculate required sample size before evaluation
  • Consider practical constraints and resources
  • Report power calculations in methodology
  • Discuss implications of achieved power

Result Reporting

Complete Statistical Reporting

  • Report effect sizes alongside p-values
  • Include confidence intervals for all estimates
  • Describe statistical methods and assumptions
  • Report negative results and non-significant findings

Uncertainty Quantification

  • Bootstrap or cross-validation for robust estimates
  • Multiple random seeds for evaluation stability
  • Sensitivity analysis for key assumptions
  • Discussion of limitations and threats to validity

Advanced Topics

Bayesian Model Comparison

  • Model evidence and Bayes factors
  • Posterior probability of model superiority
  • Hierarchical modeling for multiple tasks
  • Integration of prior knowledge

Non-Parametric Methods

  • Permutation tests for hypothesis testing
  • Rank-based methods for ordinal data
  • Distribution-free confidence intervals
  • Robust statistics for outlier resilience

Meta-Analysis

  • Combining results across multiple studies
  • Fixed-effects vs. random-effects models
  • Assessment of heterogeneity
  • Publication bias detection and correction

Practical Implementation

Statistical Software

  • R: comprehensive statistical analysis capabilities
  • Python: scipy.stats, statsmodels, scikit-learn
  • Specialized packages: pingouin, bootstrapped
  • Bayesian tools: PyMC3, Stan

Evaluation Workflows

from scipy import stats
import numpy as np

def compare_models(scores_a, scores_b):
    # Paired t-test for mean comparison
    t_stat, p_val = stats.ttest_rel(scores_a, scores_b)
    
    # Effect size calculation
    diff = np.array(scores_a) - np.array(scores_b)
    cohen_d = np.mean(diff) / np.std(diff)
    
    # Bootstrap confidence interval
    n_bootstrap = 10000
    bootstrap_diffs = []
    for _ in range(n_bootstrap):
        indices = np.random.choice(len(scores_a), len(scores_a))
        boot_diff = np.mean(np.array(scores_a)[indices] - np.array(scores_b)[indices])
        bootstrap_diffs.append(boot_diff)
    
    ci_lower = np.percentile(bootstrap_diffs, 2.5)
    ci_upper = np.percentile(bootstrap_diffs, 97.5)
    
    return {
        'p_value': p_val,
        'effect_size': cohen_d,
        'confidence_interval': (ci_lower, ci_upper)
    }

Statistical significance provides the mathematical foundation for making reliable claims about model performance, requiring careful attention to experimental design, appropriate testing, and accurate interpretation of results.