Purpose
\nInvestigate and mitigate sustained or sudden CPU utilization in Java services running on Kubernetes.
\nSymptoms
\n- \n
- CPU utilization above the service or node threshold. \n
- HPA scaling rapidly or reaching
\1. \n - Increased request latency or timeout rates. \n
- High thread activity, excessive retries, or busy loops. \n
- CPU throttling despite apparently low application CPU usage. \n
Initial Triage
\nkubectl top pods -n <namespace> --sort-by=cpu\nkubectl top nodes\nkubectl get hpa -n <namespace>\nkubectl describe pod <pod> -n <namespace>\nkubectl get events -n <namespace> --sort-by=.lastTimestamp\nCheck the configured requests and limits:
\nkubectl get deploy <deployment> -n <namespace> -o yaml\nLook for:
\n- \n
- Missing or undersized CPU requests. \n
- Very low CPU limits causing throttling. \n
- HPA targets that do not match actual workload behavior. \n
- Uneven traffic distribution across pods. \n
JVM and Thread Investigation
\nIdentify the Java process:
\nkubectl exec -n <namespace> <pod> -- sh -c 'ps -ef | grep java'\nCapture thread CPU usage:
\nkubectl exec -n <namespace> <pod> -- top -H -p <java-pid>\nConvert a hot Linux thread ID to hexadecimal:
\nprintf '%x\n' <thread-id>\nCapture a thread dump:
\nkubectl exec -n <namespace> <pod> -- jcmd <java-pid> Thread.print\nCorrelate the hexadecimal thread ID with the \1 value in the thread dump.
Java Flight Recorder
\nStart a short diagnostic recording where permitted:
\nkubectl exec -n <namespace> <pod> -- jcmd <java-pid> JFR.start name=cpu-investigation duration=120s filename=/tmp/cpu.jfr settings=profile\nCopy the recording:
\nkubectl cp <namespace>/<pod>:/tmp/cpu.jfr ./cpu.jfr\nReview:
\n- \n
- Hot methods and execution samples. \n
- Thread states and blocked time. \n
- Lock contention. \n
- Allocation pressure. \n
- Garbage-collection activity. \n
- Socket and file I/O. \n
Common Causes
\n- \n
- Inefficient loops or expensive algorithms. \n
- Excessive JSON serialization/deserialization. \n
- High-cardinality logging or debug logging. \n
- Retry storms. \n
- Synchronous blocking work on request threads. \n
- Excessive garbage creation. \n
- CPU throttling from container limits. \n
- Traffic imbalance or a single hot partition. \n
Mitigation
\n- \n
- Scale out temporarily if capacity is available. \n
- Reduce excessive logging and disable debug logging where safe. \n
- Correct CPU requests and limits. \n
- Tune HPA targets and stabilization windows. \n
- Stop runaway traffic or retry loops. \n
- Roll back a recent change if evidence points to a deployment. \n
- Apply a code fix for the identified hot path. \n
Validation
\nkubectl top pods -n <namespace>\nkubectl get hpa -n <namespace>\nConfirm:
\n- \n
- CPU returns to the expected range. \n
- Latency and error rates recover. \n
- HPA stabilizes. \n
- No new throttling or restart pattern appears. \n
Evidence to Capture
\n- \n
- Time window and affected namespace/deployment. \n
- Pod and node CPU metrics. \n
- HPA status. \n
- Deployment version and recent changes. \n
- Thread dump or JFR recording. \n
- Relevant application and access logs. \n
- Before/after graphs. \n