Common production failures, cascading degradation, recovery runbooks, and defense-in-depth patterns.
Top Incident Archetypes
1. Cascading Upstream Timeout
- **Symptom**: External provider slowdown triggers worker thread starvation across internal services.
- **Root Cause**: Missing per-request deadline propagation and lack of global connection concurrency limits.
- **Fix**: Inject timeout contexts at gateway ingress and implement bulkhead resource isolation.
2. Semantic Drift Under Load
- **Symptom**: Model outputs degrade in accuracy or hallucinate formatting under heavy concurrency.
- **Root Cause**: Prompt truncation caused by unmanaged context length growth and dynamic input overflow.
- **Fix**: Enforce strict context chunking, token budgets, and automated regression guards.
3. Rate Limit Lockout
- **Symptom**: HTTP 429 flood during peak traffic surges.
- **Root Cause**: Lack of local token bucket rate limiting prior to outbound API calls.
- **Fix**: Deploy client-side distributed token buckets with Redis or Cloudflare Durable Objects.