Purpose
\nDiagnose elevated response times, slow endpoints, timeouts, and latency regressions in Java services.
\nSymptoms
\n- \n
- Increased p95, p99, or maximum latency. \n
- Request timeouts. \n
- Thread-pool saturation. \n
- Increased downstream dependency time. \n
- High GC pauses or CPU throttling. \n
- Queue growth in ingress, application, database, or messaging layers. \n
Initial Triage
\nkubectl top pods -n <namespace>\nkubectl get pods -n <namespace> -o wide\nkubectl get hpa -n <namespace>\nkubectl get events -n <namespace> --sort-by=.lastTimestamp\nReview latency by:
\n- \n
- Endpoint. \n
- HTTP method. \n
- Status code. \n
- Pod. \n
- Availability zone or node. \n
- Request size. \n
- Downstream dependency. \n
Determine Where Time Is Spent
\nBreak request latency into:
\n1. Ingress or network time.
\n2. Application queue time.
\n3. Controller/service execution.
\n4. Database or cache calls.
\n5. External API calls.
\n6. Serialization and response transfer.
\nCheck:
\n- \n
- Connection-pool wait time. \n
- Thread-pool queue length. \n
- Database query latency. \n
- Redis latency. \n
- Kafka or messaging lag. \n
- DNS and connection-establishment time. \n
JVM Diagnostics
\nCapture a thread dump during the incident:
\nkubectl exec -n <namespace> <pod> -- jcmd <java-pid> Thread.print\nUse JFR for a bounded recording:
\nkubectl exec -n <namespace> <pod> -- jcmd <java-pid> JFR.start name=latency-investigation duration=120s filename=/tmp/latency.jfr settings=profile\nReview:
\n- \n
- Blocked and waiting threads. \n
- Lock contention. \n
- Socket read/write durations. \n
- File I/O. \n
- GC pauses. \n
- Allocation hotspots. \n
- Thread-pool saturation. \n
Common Causes
\n- \n
- Slow database queries or missing indexes. \n
- Connection-pool exhaustion. \n
- Downstream service degradation. \n
- Lock contention. \n
- CPU throttling. \n
- Long GC pauses. \n
- Synchronous calls to slow dependencies. \n
- Retry amplification. \n
- Insufficient replicas. \n
- Uneven pod load. \n
- Large payloads or expensive serialization. \n
Mitigation
\n- \n
- Reduce or stop retry storms. \n
- Increase replicas if the bottleneck is stateless capacity. \n
- Tune connection and thread pools based on measurements. \n
- Enable timeouts and circuit breakers. \n
- Route around an unhealthy dependency where possible. \n
- Roll back a confirmed latency-causing release. \n
- Optimize slow queries and add validated indexes. \n
- Reduce payload size or paginate large responses. \n
Validation
\nConfirm recovery using:
\n- \n
- p50, p95, and p99 latency. \n
- Timeout rate. \n
- Dependency latency. \n
- Pool wait time. \n
- CPU and GC metrics. \n
- Request throughput. \n
- Error rate. \n
Do not close the incident based on average latency alone.
\n