Lesson 016 · Linux Administration Learning Path

RHEL Performance Monitoring and Slow-Server Diagnosis

· Published · 7 min read

Labelled RHEL path from user symptom through CPU memory pressure storage network application evidence bottleneck correction and validation

“The server is slow” describes an experience, not a resource or cause. Start with the affected transaction, time window, population and normal baseline. Then decide whether demand changed or capacity, dependency or software behavior changed. High utilization can be healthy throughput; saturation means work is waiting. A low CPU percentage does not rule out a single-thread, lock, storage or remote dependency bottleneck.

Measure demand, utilization, saturation and errors

Use several time scales. A one-second snapshot can miss a recurring burst, while a daily average hides a ten-minute queue. Record the application result and compare host telemetry for the same interval. Pressure Stall Information shows time tasks wait for CPU, memory or I/O; it complements utilization rather than replacing workload metrics.

Do not clear caches, drop page cache, renice, kill or reboot merely to see whether a graph improves. Those actions change evidence and can harm other workloads. Prefer read-only captures, application profiling and representative controlled tests.

Build an evidence map

LayerQuestion to answerEvidence
WorkloadWhich transaction, rate and population changed?Request latency, throughput, queue and error time
CPUAre tasks running or waiting for CPU?run queue, utilization, PSI and per-thread evidence
MemoryIs reclaim, paging or OOM pressure present?available memory, PSI, faults and swap I/O
Storage/networkWhere is latency or retransmission introduced?device latency/queue, sockets and path evidence
Application/dependencyWhich lock, query or remote call dominates?profiles, traces, logs and dependency metrics

Operating sequence

  1. Define the affected operation, baseline, start time and impact.
  2. Capture load, PSI, CPU, memory, swap, storage and network without changing state.
  3. Segment by process, thread, cgroup, device, endpoint and dependency.
  4. Correlate the first deviation rather than selecting the largest current number.
  5. Reproduce in a representative safe test or use trace/profile evidence.
  6. Apply one bounded correction with a rollback trigger.
  7. Compare equal workload windows and confirm errors and downstream impact.

Commands and expected evidence

uptime
cat /proc/pressure/cpu
cat /proc/pressure/memory
cat /proc/pressure/io
vmstat 1 10
mpstat -P ALL 1 5
pidstat -dur 1 5
iostat -xz 1 5
free -h
swapon --show
ss -s
sar -n DEV 1 5
systemd-cgtop --iterations=5
  • Install sysstat through approved repositories when its tools are required.
  • Load average includes tasks waiting in uninterruptible states and must be interpreted with CPU count and wait evidence.
  • wa is not a complete device-latency metric; inspect per-device await, queue and application I/O.
  • Production profiling and packet capture need authorization, bounded duration and protected output.

Evidence and acceptance criteria

EvidenceHealthy resultFailure meaning
User symptomExact transaction and time reproduce or correlateHost metric may be unrelated
CPU/memory pressureWait evidence identifies constrained resourceUtilization alone is ambiguous
Device/pathSpecific device/endpoint latency or errors alignWrong layer selected
ApplicationTrace/profile explains time consumptionHost tuning may mask code/dependency defect
After changeEqual workload improves with guardrails stableApparent gain came from lower demand or lost work

Worked operating scenario

A host shows 40% total CPU while one customer workflow is slow. Per-thread evidence shows one application thread at 100% of one core and a growing request queue. Adding memory or dropping caches cannot remove that serialization point. The team profiles the bounded code path, removes an accidental global lock and compares equal traffic cohorts.

Acceptance includes queue depth, tail latency, error rate and CPU per completed request. A higher total CPU after the fix is healthy because more work completes concurrently.

Practical how-to cases

Case 1: Separate load from CPU saturation

Observe run queue, per-CPU use and pressure during a controlled CPU workload. Compare with a workload baseline and change one causal constraint at a time.

uptime
vmstat 1 10
mpstat -P ALL 1 5
pidstat -u 1 5
cat /proc/pressure/cpu
CheckpointWhat to establish
Expected resultRunnable work, CPU utilization and pressure agree on whether tasks wait for CPU.
If it failsLoad average includes uninterruptible work; a high number is not automatically CPU saturation.
Safe recoveryRemove the lab load or tuning override, restore the recorded sysctl/unit configuration, and prove latency and pressure return to baseline.

Case 2: Diagnose memory pressure

Relate available memory, reclaim, swap and PSI to workload latency. Compare with a workload baseline and change one causal constraint at a time.

free -h
vmstat 1 10
swapon --show
cat /proc/pressure/memory
pidstat -r 1 5
journalctl -k | grep -i -E 'oom|out of memory'
CheckpointWhat to establish
Expected resultThe evidence distinguishes cache use from sustained reclaim, swapping or an OOM event.
If it failsLow free memory alone is normal when cache is reclaimable.
Safe recoveryRemove the lab load or tuning override, restore the recorded sysctl/unit configuration, and prove latency and pressure return to baseline.

Case 3: Diagnose storage latency

Measure device latency and process I/O before blaming filesystem capacity. Compare with a workload baseline and change one causal constraint at a time.

iostat -xz 1 10
pidstat -d 1 10
cat /proc/pressure/io
df -hT
df -ih
sudo lsof +L1
CheckpointWhat to establish
Expected resultA device and workload correlate with elevated await, queue or I/O pressure.
If it failsHigh utilization metrics differ by device; compare latency and workload behavior.
Safe recoveryRemove the lab load or tuning override, restore the recorded sysctl/unit configuration, and prove latency and pressure return to baseline.

Case 4: Diagnose network application delay

Measure sockets, retransmissions and application timing as separate layers. Compare with a workload baseline and change one causal constraint at a time.

ss -s
nstat -az | grep -i retrans
sar -n DEV,TCP,ETCP 1 5
curl -sS -o /dev/null -w 'dns=%{time_namelookup} connect=%{time_connect} first=%{time_starttransfer} total=%{time_total}\n' https://example.com/
tracepath example.com
CheckpointWhat to establish
Expected resultThe delay is assigned to name resolution, connection, server response or transfer with matching host counters.
If it failsOne curl timing is a sample; use repeated tests and application monitoring before concluding.
Safe recoveryRemove the lab load or tuning override, restore the recorded sysctl/unit configuration, and prove latency and pressure return to baseline.

Independent practice tasks

  1. Create controlled CPU, memory and I/O load separately and identify each.
  2. Use tuned-adm to inspect, not blindly change, the active profile.
  3. Set a systemd resource limit on a lab service and prove its effect.
  4. Write a slow-server incident report with before/after evidence.

For this lesson on RHEL Performance Troubleshooting, complete each task without copying the worked command sequence. Record the initial state, exact change, verification, negative test and recovery command. A task is unfinished if it works now but does not survive a reboot where persistence is required.

Troubleshooting by symptom

SymptomInspect firstDefensible next action
High load, idle CPUI/O wait, blocked tasks and wait channelFind storage/lock/dependency wait
Swap used, no active pagingPaging rate and PSI over timeDo not disable swap solely from used bytes
Disk 100% busyawait, queue, throughput and workload patternIdentify device/workload and latency objective
Network “slow”retransmits, drops, RTT, socket queues and app timingLocate host/path/protocol layer
Fix improves average onlytail latency, errors and dropped workUse complete service-level acceptance

Unsafe operations and recovery boundaries

  • Unsafe: dropping caches on production changes global memory behavior and can create an I/O storm.
  • Unsafe: running unbounded benchmarks competes with the workload and can fill disk or network.
  • Unsafe: killing the largest process destroys evidence and may terminate the healthy workload rather than the cause.

Rewritten knowledge checks

What is saturation?
Work is waiting because a resource or serialized path cannot serve demand immediately.
Why is CPU average misleading?
A single core/thread can saturate while the machine-wide average remains low.
What does PSI measure?
Time tasks are stalled for CPU, memory or I/O resources.
Does swap used prove memory pressure?
No; inspect active paging, PSI, available memory and workload latency.
Why correlate time windows?
Metrics outside the user symptom window may describe unrelated normal work.
What is tail latency?
High-percentile response time showing the slow end hidden by averages.
Why compare equal workload?
Lower demand can make a change look faster without improving efficiency.
What is a complete performance fix?
A causal correction that improves target outcomes without unacceptable errors, cost or downstream harm.

Guided lab and acceptance test

  1. Create a baseline for a small CPU and file-I/O workload.
  2. Introduce one bounded CPU worker and observe per-core and PSI evidence.
  3. Introduce a capped file-I/O job on disposable storage and inspect device latency.
  4. Create memory pressure within a cgroup limit and observe PSI without affecting the host broadly.
  5. Write a symptom-to-layer diagnosis for each case.
  6. Stop lab workloads and prove resources, files and units are removed.

Primary references

Advertisement