From performance to reliability
The work started with performance: workloads, latency, capacity, and how an application changes when the load is no longer polite. Over time that sat next to production diagnosis.
Most days are ordinary. Read the dump. Check the probe. Compare this run with the last one. Write down what actually moved the needle. When the same diagnosis returns, a small tool or a short guide is enough.