$ whoami

Manjunatha Reddy T_

Site Reliability & Performance Engineer

Thirteen years in performance engineering, non-functional testing, and production diagnosis — workload models, latency and capacity, JVM dumps, and the quiet work of making the next incident shorter.

$ cat /notes/approach.txt

Read the system first. Follow the evidence. Change little. Leave a shorter path.
13+years in engineering
4core focus areas
6groups in the field kit
1working practice
01. JOURNEY

From performance to reliability

The work started with performance: workloads, latency, capacity, and how an application changes when the load is no longer polite. Over time that sat next to production diagnosis.

Most days are ordinary. Read the dump. Check the probe. Compare this run with the last one. Write down what actually moved the needle. When the same diagnosis returns, a small tool or a short guide is enough.

performance production diagnostics reliability notes
  • Studies

    BE in Computer Science. Learned to reason about programs before putting a load on one.

    2008 — 2012
  • Started with tests

    Performance engineer

    Performance testing and workload behaviour.

    ~2012
  • Closer to production

    Live issues, dumps and traces. Reliability next to performance.

    ~2015
  • Tests in the pipeline

    2013 — 2020

    JMeter in CI. Capacity forecasts. Bottlenecks across versions.

    ~2016
  • Distributed systems

    More hops. Same questions — latency, restarts, where it waits.

    ~2018
  • SRE and performance together

    Senior SRE & Performance

    NFR plus diagnosis across application, JVM, data and Kubernetes.

    2020 — now
  • Helpers for repeated steps

    Small tools and notes for diagnoses that kept returning.

    ~2021
  • Still studying

    Runtime diagnostics. Local models with Ollama for log summaries — early, and kept on the machine.

    ~2024 — now
02. FIELD KIT

Engineering toolkit BUILD

A collection of practical utilities for recurring analysis and troubleshooting workflows.

  • log and runtime analysis
  • performance investigation
  • reporting and workflow automation

Cloud operations BUILD

Public-safe work around workload visibility, controlled operations, and faster incident diagnosis.

  • workload health
  • operational diagnostics
  • role-based workflows

Automation & analysis EXPLORE

Small helpers for repeated investigation steps, including local models for AI-assisted analysis.

  • collect → correlate → analyze
  • local LLMs with Ollama
  • summaries that stay on the machine
03. LEARNING
01
Java / JVMJFR, GC logs, heap dumps, thread dumps.
02
PerformanceWorkload modeling, JMeter, latency, capacity.
03
KubernetesPod health, probes, limits, deployment issues.
04
ObservabilityPrometheus, Grafana, ELK — enough to follow a path.
05
LanguagesJava, Python, Bash, SQL.
06
Local modelsOllama for on-machine LLMs. Early use: summarise logs and dumps without sending them out.
04. PRINCIPLES
Read the system first.
Follow the evidence.
Change little.
Leave a shorter path.
  • Dumps and traces before a production change
  • Performance tests in CI when they can catch a regression
  • Automate only a diagnosis that has already repeated
  • Share notes when they help someone else start faster
05. LET'S CONNECT

Open to a technical conversation.

Professional questions, collaboration, and ideas around reliable systems. LinkedIn is the best place to reach me.