~/wiki

Production Debugging

Confiance : high
production-debuggingerror-monitoringsentry-integrationenvironment-configurationsystematic-troubleshootingvercel-deployment

Systematic approach to diagnosing and resolving issues in live production systems, particularly focusing on configuration mismatches, monitoring blind spots, and error propagation patterns in web applications.

Core Methodology

Environment Configuration Audit

Critical first step in production debugging involves verifying that environment variables, service configurations, and third-party integrations match between development and production environments.

Common Configuration Drift Patterns:

  • Service project slugs or IDs differing between environments
  • API endpoints pointing to wrong regions (e.g., EU vs US Sentry instances)
  • Database connection strings or authentication tokens being outdated
  • Feature flags or service tiers varying between dev/prod

Error Monitoring Validation

Before investigating specific bugs, ensure monitoring systems are properly configured and capturing exceptions.

Validation Steps:

  1. Verify monitoring service configuration (project slugs, API keys, regions)
  2. Test API connectivity from production environment
  3. Check that caught exceptions are explicitly reported to monitoring
  4. Validate error boundaries are properly configured for React applications

Error Propagation Analysis

Understanding how errors flow through system layers helps identify where failures are being masked or improperly handled.

Investigation Pattern:

  1. Client-side errors: Network timeouts, parsing failures, rendering exceptions
  2. Server-side errors: API route exceptions, database connection failures, external service timeouts
  3. Infrastructure errors: Function timeouts, memory limits, deployment failures

Sentry Integration Debugging

Project Configuration Issues

Sentry projects often use automatically-generated slugs that differ from expected names:

  • Default Next.js projects typically use javascript-nextjs slug
  • Custom project names may not match slug format
  • Organization regions affect both ingestion and API endpoints

API Connectivity Testing

When Sentry dashboard shows no issues, test API connectivity directly:

# Test organization access
curl -H "Authorization: Bearer $SENTRY_TOKEN" \
  https://sentry.io/api/0/organizations/$ORG/

# Test project issues endpoint  
curl -H "Authorization: Bearer $SENTRY_TOKEN" \
  https://sentry.io/api/0/projects/$ORG/$PROJECT/issues/

Exception Capture Gaps

Common scenarios where exceptions aren't captured:

  • Caught exceptions in try/catch blocks without explicit Sentry.captureException()
  • Client-side network errors that don't propagate to error boundaries
  • Server-side route handlers returning error responses instead of throwing

Vercel Function Debugging

Timeout Patterns

Vercel function timeouts manifest differently based on plan limits:

  • Hobby plan: 10-second default, 300-second maximum
  • Pro/Enterprise: Configurable limits up to 900 seconds
  • Timeout errors often return HTML pages instead of JSON, breaking client parsing

Function Monitoring

  • Check Vercel dashboard for function invocation logs
  • Monitor memory usage and cold start times
  • Verify environment variables are properly set in production

Error Classification Framework

Network vs Application Errors

  • Network errors: Connection timeouts, DNS failures, SSL issues
  • Application errors: Logic bugs, data validation failures, type errors
  • Infrastructure errors: Memory limits, function timeouts, service unavailability

Client vs Server Boundary

Understanding where errors originate helps prioritize fixes:

  • Client errors: Often manifest as generic "connection error" messages
  • Server errors: Should be captured by monitoring if properly configured
  • Serialization errors: Occur during Next.js server-to-client data transfer

Best Practices

Defensive Error Handling

  • Always include explicit Sentry.captureException() calls in catch blocks
  • Implement proper null/undefined checks for data that crosses serialization boundaries
  • Use timeout wrappers for external API calls

Error Message Clarity

  • Distinguish between different failure types in user-facing messages
  • Include relevant context (operation, timestamp, user ID) in error logs
  • Provide actionable guidance when possible

Configuration Management

  • Use infrastructure-as-code for environment variable management
  • Implement configuration validation in application startup
  • Document expected vs actual configuration values for debugging

See also