sreroot

Runbooks / High latency

\n

High latency

\n

Kubernetes · Java · incident notes

\n
\n

Purpose

\n

Diagnose elevated response times, slow endpoints, timeouts, and latency regressions in Java services.

\n

Symptoms

\n\n

Initial Triage

\n
kubectl top pods -n <namespace>\nkubectl get pods -n <namespace> -o wide\nkubectl get hpa -n <namespace>\nkubectl get events -n <namespace> --sort-by=.lastTimestamp
\n

Review latency by:

\n\n

Determine Where Time Is Spent

\n

Break request latency into:

\n

1. Ingress or network time.

\n

2. Application queue time.

\n

3. Controller/service execution.

\n

4. Database or cache calls.

\n

5. External API calls.

\n

6. Serialization and response transfer.

\n

Check:

\n\n

JVM Diagnostics

\n

Capture a thread dump during the incident:

\n
kubectl exec -n <namespace> <pod> -- jcmd <java-pid> Thread.print
\n

Use JFR for a bounded recording:

\n
kubectl exec -n <namespace> <pod> --   jcmd <java-pid> JFR.start name=latency-investigation duration=120s filename=/tmp/latency.jfr settings=profile
\n

Review:

\n\n

Common Causes

\n\n

Mitigation

\n\n

Validation

\n

Confirm recovery using:

\n\n

Do not close the incident based on average latency alone.

\n