Purpose
\nInvestigate elevated HTTP 4xx/5xx responses, application exceptions, failed dependencies, and rollout-related errors.
\nSymptoms
\n- \n
- Increased 5xx rate. \n
- Unexpected 4xx spike. \n
- Failed health checks. \n
- CrashLoopBackOff or frequent restarts. \n
- Increased timeout, reset, or connection-refused errors. \n
- Error concentration in one pod, node, version, or endpoint. \n
Initial Triage
\nkubectl get pods -n <namespace> -o wide\nkubectl get deploy <deployment> -n <namespace>\nkubectl rollout history deploy/<deployment> -n <namespace>\nkubectl get events -n <namespace> --sort-by=.lastTimestamp\nInspect logs:
\nkubectl logs -n <namespace> <pod> --since=30m\nkubectl logs -n <namespace> <pod> --previous\nCheck status-code distribution by:
\n- \n
- Endpoint. \n
- HTTP method. \n
- Pod. \n
- Application version. \n
- Dependency. \n
- Tenant or request type. \n
- Region or availability zone. \n
Classify the Errors
\n4xx Errors
\nInvestigate:
\n- \n
- Authentication or authorization failures. \n
- Invalid request payloads. \n
- Missing headers or cookies. \n
- API contract changes. \n
- Routing or ingress rules. \n
- Rate limiting. \n
- Client version incompatibility. \n
5xx Errors
\nInvestigate:
\n- \n
- Application exceptions. \n
- Database or cache failures. \n
- Downstream timeouts. \n
- Connection-pool exhaustion. \n
- Out-of-memory events. \n
- Failed readiness or liveness probes. \n
- Bad configuration or secret changes. \n
- Deployment regressions. \n
Kubernetes Checks
\nkubectl describe pod <pod> -n <namespace>\nkubectl describe svc <service> -n <namespace>\nkubectl describe ingress <ingress> -n <namespace>\nkubectl get endpoints <service> -n <namespace>\nkubectl get endpointslices -n <namespace>\nVerify:
\n- \n
- Ready endpoints exist. \n
- Service selectors match pod labels. \n
- Readiness probes are correct. \n
- Ingress routes to the expected service. \n
- Pods are not repeatedly removed from service. \n
- Network policies permit required traffic. \n
Java Diagnostics
\nCapture relevant exceptions and a thread dump:
\nkubectl exec -n <namespace> <pod> -- jcmd <java-pid> Thread.print\nUse JFR when errors appear related to:
\n- \n
- Lock contention. \n
- Slow I/O. \n
- GC pauses. \n
- Thread starvation. \n
- Excessive allocation. \n
Mitigation
\n- \n
- Roll back a confirmed bad deployment. \n
- Remove unhealthy pods only when replacement capacity is available. \n
- Correct configuration, secret, or routing errors. \n
- Restore failed dependencies or activate a fallback. \n
- Adjust probe settings only after validating application startup and health behavior. \n
- Rate-limit abusive or runaway clients. \n
- Disable a faulty feature flag if an approved control exists. \n
Validation
\nConfirm:
\n- \n
- 4xx/5xx rates return to baseline. \n
- Error signatures disappear or reduce materially. \n
- All intended endpoints have healthy backends. \n
- No new restart or probe-failure pattern appears. \n
- Recovery is consistent across pods and versions. \n
Evidence to Capture
\n- \n
- Error-rate graph and exact incident window. \n
- Representative request IDs. \n
- Application stack traces. \n
- Pod events and termination reasons. \n
- Deployment version and recent changes. \n
- Dependency health and latency. \n
- Before/after validation metrics. \n