FOR SREs & PLATFORM ENGINEERS
Updated 2026

Notes on strace
for production debugging

Field notes for diagnosing Linux applications, containers, and performance issues with system call tracing.

A handbook. Not a training programme.
19
Modules
15
Hands-on Labs
26
War Stories
Production-ready playbooks
Java, Python, Go, Node.js examples
Kubernetes & container debugging
Cheatsheet and quiz
WHY SREs NEED STRACE

When logs and metrics
are not enough

Application Hangs & Deadlocks

strace reveals exactly which futex or recvfrom call your process is blocked on — something no APM or log aggregator can show.

Mysterious CPU Spikes

Identify tight loops of failing stat() calls or excessive polling that monitoring dashboards only show as "high CPU".

Container & Startup Failures

Quickly diagnose ENOENT, EACCES, missing libraries, and port conflicts in Kubernetes pods where logs are empty.

19 SECTIONS

Guide sections

Module 1

Introduction to strace

What is strace, ptrace internals, syscall lifecycle, blocking vs non-blocking.

Module 2

Why Every SRE Should Know strace

Production incidents where logs/metrics fail. Comparison tables and real scenarios.

Module 3

strace Fundamentals

Attaching to processes, tracing children & threads, timestamps, durations, output files.

Module 4

Understanding Common Syste