Purpose
\nInvestigate high memory usage, OOMKills, container restarts, and possible Java heap or native-memory leaks.
\nSymptoms
\n- \n
- Container memory near its limit. \n
\1termination reason. \n- Frequent pod restarts. \n
- Increasing old-generation occupancy. \n
- Long or frequent garbage-collection pauses. \n
\1in application logs. \n
Initial Triage
\nkubectl top pods -n <namespace> --sort-by=memory\nkubectl describe pod <pod> -n <namespace>\nkubectl get events -n <namespace> --sort-by=.lastTimestamp\nkubectl get pod <pod> -n <namespace> -o jsonpath='{.status.containerStatuses[*].lastState.terminated.reason}'\nCheck configured resources:
\nkubectl get deploy <deployment> -n <namespace> -o yaml\nCompare:
\n- \n
- Container memory limit. \n
- JVM heap settings. \n
- Non-heap and native-memory requirements. \n
- Number of threads. \n
- Direct-buffer usage. \n
- Metaspace usage. \n
JVM Memory Checks
\nkubectl exec -n <namespace> <pod> -- jcmd <java-pid> VM.flags\nkubectl exec -n <namespace> <pod> -- jcmd <java-pid> GC.heap_info\nkubectl exec -n <namespace> <pod> -- jcmd <java-pid> VM.native_memory summary\nNative Memory Tracking must be enabled at JVM startup for detailed output:
\n-XX:NativeMemoryTracking=summary\nCheck GC logs for:
\n- \n
- Increasing post-GC heap occupancy. \n
- Full GC frequency. \n
- Promotion failures. \n
- Humongous allocations. \n
- Long pause times. \n
Heap Dump
\nOnly capture a heap dump after assessing disk capacity and operational impact:
\nkubectl exec -n <namespace> <pod> -- jcmd <java-pid> GC.heap_dump /tmp/heap.hprof\nkubectl cp <namespace>/<pod>:/tmp/heap.hprof ./heap.hprof\nAnalyze with Eclipse MAT or another approved heap-analysis tool.
\nLook for:
\n- \n
- Dominator tree leaders. \n
- Large collections and maps. \n
- Retained objects. \n
- Class-loader leaks. \n
- Unbounded caches. \n
- Request/session objects retained too long. \n
Common Causes
\n- \n
- Unbounded in-memory caches. \n
- Static collections retaining objects. \n
- ThreadLocal leaks. \n
- Excessive direct buffers. \n
- Too many threads. \n
- Large response/request payloads. \n
- Class-loader leaks after redeployments. \n
- Heap sizing too close to the container limit. \n
- Native memory or off-heap allocations. \n
Mitigation
\n- \n
- Scale out if memory pressure is traffic-related. \n
- Reduce cache size or add expiry. \n
- Fix object-retention or listener leaks. \n
- Reduce payload sizes and buffering. \n
- Review thread-pool sizes. \n
- Increase container memory only after identifying the memory category. \n
- Keep headroom between JVM heap and container limit. \n
- Enable heap-dump-on-OOM only with sufficient disk capacity. \n
Validation
\nkubectl top pod <pod> -n <namespace>\nkubectl get pod <pod> -n <namespace> -w\nConfirm:
\n- \n
- Memory remains below the limit. \n
- No repeated OOMKills. \n
- Post-GC heap occupancy is stable. \n
- Restart rate returns to normal. \n
- Latency and error rates remain healthy. \n