composite scoring
---
title: Composite Scoring
category: concepts
created: 2026-12-21
updated: 2025-01-04
tags: [composite-scoring, multi-metric-evaluation, agent-assessment, weighted-scoring, citation-quality, accuracy-measurement, reinforcement-learning, performance-optimization, workshop-competitions, reward-engineering, metric-balancing, grpo-training, real-time-optimization, time-pressure-tuning]
sources: [raw/conversations/2026-06-03-cursor-workshops-WB-11eba586.md]
confidence: high
---
# Composite Scoring
Multi-dimensional evaluation framework that combines multiple performance metrics into a single score for AI agent assessment. Particularly important in agent competitions and reinforcement learning scenarios where agents must balance multiple objectives simultaneously.
## Core Concept
Composite scoring addresses the fundamental challenge that AI agent performance cannot be captured by a single metric. Instead, it weighs different aspects of performance to create holistic evaluation:
Composite Score = w₁ × Correctness + w₂ × Citation_F1 + w₃ × Efficiency + w₄ × Penalties
## Workshop Implementation Example
In W&B's Enron corpus mining workshops, agents are evaluated using:
### Primary Metrics
- **Correctness** (weight: 0.7): Binary accuracy of information extraction
- **Citation F1** (weight: 0.65): Quality of source attribution and evidence
- **Efficiency Bonus** (weight: 0.02): Token usage optimization
### Penalty Components
- **No-answer penalty** (-0.2): Discourages refusal when information exists
- **Over-citation penalty**: Built into F1 calculation (precision component)
### Metric Interactions
The weighting reveals strategic priorities:
- **High correctness weight** ensures factual accuracy remains primary
- **Substantial citation weight** enforces evidence-based reasoning
- **Low efficiency weight** prevents over-optimization for brevity
- **Negative penalties** discourage gaming behaviors
## Reinforcement Learning Integration
### GRPO Optimization
Composite scores enable Group Relative Policy Optimization (GRPO) by providing:
- **Ranking signals** for comparing rollouts within training groups
- **Continuous feedback** through F1 scores (not just binary accuracy)
- **Reward variance** essential for effective RL training
### Training Dynamics
Real workshop data shows typical patterns:
- **Early phase**: Agents learn basic answering behavior (`has_answer` rises to 1.0)
- **Accuracy phase**: Correctness gradually improves (0.25 → 0.5+)
- **Citation phase**: F1 scores lag, requiring targeted weight adjustment
## Optimization Strategies
### Weight Balancing
Effective composite scoring requires iterative weight adjustment:
1. **Identify bottleneck metrics**: Which component limits overall score?
2. **Increase bottleneck weights**: Drive RL focus to weakest areas
3. **Monitor for overfitting**: Ensure improvements are genuine, not gaming
4. **Balance competing objectives**: Prevent one metric dominating others
### Time-Constrained Scenarios
Under competition pressure (e.g., 45-minute workshops):
- **Target largest gaps first**: Citation F1 from 0.38 → 0.6+ yields biggest gains
- **Avoid experimental changes**: Stick to validated weight adjustments
- **Conservative step increases**: 3 → 20 training steps vs. risky 50+
## Common Failure Modes
### Reward Hacking
Agents may exploit composite scoring by:
- **Always answering** to avoid no-answer penalties while giving wrong answers
- **Over-citing** to boost citation recall at expense of precision
- **Gaming efficiency** through extremely brief but unhelpful responses
### Metric Plateaus
When components plateau:
- **Citation F1 at 0.3-0.4**: Often indicates poor source attribution training
- **Correctness at 0.5**: May suggest insufficient training steps or poor reward signal
- **Both metrics flat**: Usually indicates reward variance collapse in RL
## Implementation Best Practices
### Weight Selection
- **Start with domain priorities**: What matters most for the use case?
- **Empirical adjustment**: Let training data reveal appropriate balances
- **Competitive benchmarking**: Compare against baseline agent performance
### Monitoring
Essential metrics during optimization:
- **Individual component trends**: Which metrics improve vs. plateau?
- **Reward variance**: Sufficient spread for RL ranking?
- **Training efficiency**: Are groups becoming non-trainable?
### Validation
- **Hold-out evaluation**: Ensure composite improvements transfer to unseen data
- **Human evaluation**: Spot-check that optimized agents meet real quality standards
- **Edge case testing**: Verify agents don't exploit scoring mechanics
## Strategic Applications
Composite scoring proves essential for:
- **Multi-objective RL**: Where simple accuracy insufficient
- **Competitive ML**: Tournaments requiring holistic performance measures
- **Production systems**: Balancing quality, cost, and user experience
- **Research evaluation**: Comprehensive model comparison beyond single metrics
The framework enables systematic optimization of complex agent behaviors while maintaining interpretability of what drives performance improvements.