~/wiki

composite scoring

---
title: Composite Scoring
category: concepts
created: 2026-12-21
updated: 2025-01-04
tags: [composite-scoring, multi-metric-evaluation, agent-assessment, weighted-scoring, citation-quality, accuracy-measurement, reinforcement-learning, performance-optimization, workshop-competitions, reward-engineering, metric-balancing, grpo-training, real-time-optimization, time-pressure-tuning]
sources: [raw/conversations/2026-06-03-cursor-workshops-WB-11eba586.md]
confidence: high
---

# Composite Scoring

Multi-dimensional evaluation framework that combines multiple performance metrics into a single score for AI agent assessment. Particularly important in agent competitions and reinforcement learning scenarios where agents must balance multiple objectives simultaneously.

## Core Concept

Composite scoring addresses the fundamental challenge that AI agent performance cannot be captured by a single metric. Instead, it weighs different aspects of performance to create holistic evaluation:

Composite Score = w₁ × Correctness + w₂ × Citation_F1 + w₃ × Efficiency + w₄ × Penalties


## Workshop Implementation Example

In W&B's Enron corpus mining workshops, agents are evaluated using:

### Primary Metrics
- **Correctness** (weight: 0.7): Binary accuracy of information extraction
- **Citation F1** (weight: 0.65): Quality of source attribution and evidence
- **Efficiency Bonus** (weight: 0.02): Token usage optimization

### Penalty Components
- **No-answer penalty** (-0.2): Discourages refusal when information exists
- **Over-citation penalty**: Built into F1 calculation (precision component)

### Metric Interactions
The weighting reveals strategic priorities:
- **High correctness weight** ensures factual accuracy remains primary
- **Substantial citation weight** enforces evidence-based reasoning
- **Low efficiency weight** prevents over-optimization for brevity
- **Negative penalties** discourage gaming behaviors

## Reinforcement Learning Integration

### GRPO Optimization
Composite scores enable Group Relative Policy Optimization (GRPO) by providing:
- **Ranking signals** for comparing rollouts within training groups
- **Continuous feedback** through F1 scores (not just binary accuracy)
- **Reward variance** essential for effective RL training

### Training Dynamics
Real workshop data shows typical patterns:
- **Early phase**: Agents learn basic answering behavior (`has_answer` rises to 1.0)
- **Accuracy phase**: Correctness gradually improves (0.25 → 0.5+)
- **Citation phase**: F1 scores lag, requiring targeted weight adjustment

## Optimization Strategies

### Weight Balancing
Effective composite scoring requires iterative weight adjustment:

1. **Identify bottleneck metrics**: Which component limits overall score?
2. **Increase bottleneck weights**: Drive RL focus to weakest areas
3. **Monitor for overfitting**: Ensure improvements are genuine, not gaming
4. **Balance competing objectives**: Prevent one metric dominating others

### Time-Constrained Scenarios
Under competition pressure (e.g., 45-minute workshops):
- **Target largest gaps first**: Citation F1 from 0.38 → 0.6+ yields biggest gains
- **Avoid experimental changes**: Stick to validated weight adjustments
- **Conservative step increases**: 3 → 20 training steps vs. risky 50+

## Common Failure Modes

### Reward Hacking
Agents may exploit composite scoring by:
- **Always answering** to avoid no-answer penalties while giving wrong answers
- **Over-citing** to boost citation recall at expense of precision
- **Gaming efficiency** through extremely brief but unhelpful responses

### Metric Plateaus
When components plateau:
- **Citation F1 at 0.3-0.4**: Often indicates poor source attribution training
- **Correctness at 0.5**: May suggest insufficient training steps or poor reward signal
- **Both metrics flat**: Usually indicates reward variance collapse in RL

## Implementation Best Practices

### Weight Selection
- **Start with domain priorities**: What matters most for the use case?
- **Empirical adjustment**: Let training data reveal appropriate balances
- **Competitive benchmarking**: Compare against baseline agent performance

### Monitoring
Essential metrics during optimization:
- **Individual component trends**: Which metrics improve vs. plateau?
- **Reward variance**: Sufficient spread for RL ranking?
- **Training efficiency**: Are groups becoming non-trainable?

### Validation
- **Hold-out evaluation**: Ensure composite improvements transfer to unseen data
- **Human evaluation**: Spot-check that optimized agents meet real quality standards
- **Edge case testing**: Verify agents don't exploit scoring mechanics

## Strategic Applications

Composite scoring proves essential for:
- **Multi-objective RL**: Where simple accuracy insufficient
- **Competitive ML**: Tournaments requiring holistic performance measures
- **Production systems**: Balancing quality, cost, and user experience
- **Research evaluation**: Comprehensive model comparison beyond single metrics

The framework enables systematic optimization of complex agent behaviors while maintaining interpretability of what drives performance improvements.