Stacked Servers
model-context-protocol deployment pattern where multiple MCP servers are connected simultaneously, creating multiplicative schema-bloat effects that can consume massive context windows before any productive work begins.
Problem Scale
Common Enterprise Stack:
- GitHub MCP server (~55k tokens)
- Jira MCP server
- Database MCP server
- Microsoft Graph MCP server
Combined overhead: Easily exceeds 150k tokens before model performs any useful work.
Root Cause
Each MCP server must declare its complete tool schema upfront for action-discovery. While individual servers may be well-designed, the protocol's architecture creates unavoidable additive cost when multiple services are integrated.
Mitigation Strategies
lazy-loading: Only load schemas on-demand rather than upfront Selective Connection: Connect only necessary servers for specific tasks File-Based Outputs: Reduce round-tripping through context window Server Consolidation: Purpose-built servers covering multiple related services
Context Dependency
Stacked server overhead particularly problematic in:
- Token-constrained environments
- Rapid iteration workflows
- Single-user contexts prioritizing efficiency
More acceptable in enterprise-ai contexts where governance benefits justify cost.