RHEL Performance Monitoring and Slow-Server Diagnosis
“The server is slow” describes an experience, not a resource or cause. Start with the affected transaction, time window, population and normal baseline. Then decide whether demand changed or capacity, dependency or software behavior changed. High utilization can be healthy throughput; saturation means work is waiting. A low CPU percentage does not rule out a single-thread, lock, storage or remote dependency bottleneck.
Measure demand, utilization, saturation and errors
Use several time scales. A one-second snapshot can miss a recurring burst, while a daily average hides a ten-minute queue. Record the application result and compare host telemetry for the same interval. Pressure Stall Information shows time tasks wait for CPU, memory or I/O; it complements utilization rather than replacing workload metrics.
Do not clear caches, drop page cache, renice, kill or reboot merely to see whether a graph improves. Those actions change evidence and can harm other workloads. Prefer read-only captures, application profiling and representative controlled tests.
Build an evidence map
| Layer | Question to answer | Evidence |
|---|---|---|
| Workload | Which transaction, rate and population changed? | Request latency, throughput, queue and error time |
| CPU | Are tasks running or waiting for CPU? | run queue, utilization, PSI and per-thread evidence |
| Memory | Is reclaim, paging or OOM pressure present? | available memory, PSI, faults and swap I/O |
| Storage/network | Where is latency or retransmission introduced? | device latency/queue, sockets and path evidence |
| Application/dependency | Which lock, query or remote call dominates? | profiles, traces, logs and dependency metrics |
Operating sequence
- Define the affected operation, baseline, start time and impact.
- Capture load, PSI, CPU, memory, swap, storage and network without changing state.
- Segment by process, thread, cgroup, device, endpoint and dependency.
- Correlate the first deviation rather than selecting the largest current number.
- Reproduce in a representative safe test or use trace/profile evidence.
- Apply one bounded correction with a rollback trigger.
- Compare equal workload windows and confirm errors and downstream impact.
Commands and expected evidence
uptime
cat /proc/pressure/cpu
cat /proc/pressure/memory
cat /proc/pressure/io
vmstat 1 10
mpstat -P ALL 1 5
pidstat -dur 1 5
iostat -xz 1 5
free -h
swapon --show
ss -s
sar -n DEV 1 5
systemd-cgtop --iterations=5- Install sysstat through approved repositories when its tools are required.
- Load average includes tasks waiting in uninterruptible states and must be interpreted with CPU count and wait evidence.
wais not a complete device-latency metric; inspect per-device await, queue and application I/O.- Production profiling and packet capture need authorization, bounded duration and protected output.
Evidence and acceptance criteria
| Evidence | Healthy result | Failure meaning |
|---|---|---|
| User symptom | Exact transaction and time reproduce or correlate | Host metric may be unrelated |
| CPU/memory pressure | Wait evidence identifies constrained resource | Utilization alone is ambiguous |
| Device/path | Specific device/endpoint latency or errors align | Wrong layer selected |
| Application | Trace/profile explains time consumption | Host tuning may mask code/dependency defect |
| After change | Equal workload improves with guardrails stable | Apparent gain came from lower demand or lost work |
Worked operating scenario
A host shows 40% total CPU while one customer workflow is slow. Per-thread evidence shows one application thread at 100% of one core and a growing request queue. Adding memory or dropping caches cannot remove that serialization point. The team profiles the bounded code path, removes an accidental global lock and compares equal traffic cohorts.
Acceptance includes queue depth, tail latency, error rate and CPU per completed request. A higher total CPU after the fix is healthy because more work completes concurrently.
Practical how-to cases
Case 1: Separate load from CPU saturation
Observe run queue, per-CPU use and pressure during a controlled CPU workload. Compare with a workload baseline and change one causal constraint at a time.
uptime
vmstat 1 10
mpstat -P ALL 1 5
pidstat -u 1 5
cat /proc/pressure/cpu| Checkpoint | What to establish |
|---|---|
| Expected result | Runnable work, CPU utilization and pressure agree on whether tasks wait for CPU. |
| If it fails | Load average includes uninterruptible work; a high number is not automatically CPU saturation. |
| Safe recovery | Remove the lab load or tuning override, restore the recorded sysctl/unit configuration, and prove latency and pressure return to baseline. |
Case 2: Diagnose memory pressure
Relate available memory, reclaim, swap and PSI to workload latency. Compare with a workload baseline and change one causal constraint at a time.
free -h
vmstat 1 10
swapon --show
cat /proc/pressure/memory
pidstat -r 1 5
journalctl -k | grep -i -E 'oom|out of memory'| Checkpoint | What to establish |
|---|---|
| Expected result | The evidence distinguishes cache use from sustained reclaim, swapping or an OOM event. |
| If it fails | Low free memory alone is normal when cache is reclaimable. |
| Safe recovery | Remove the lab load or tuning override, restore the recorded sysctl/unit configuration, and prove latency and pressure return to baseline. |
Case 3: Diagnose storage latency
Measure device latency and process I/O before blaming filesystem capacity. Compare with a workload baseline and change one causal constraint at a time.
iostat -xz 1 10
pidstat -d 1 10
cat /proc/pressure/io
df -hT
df -ih
sudo lsof +L1| Checkpoint | What to establish |
|---|---|
| Expected result | A device and workload correlate with elevated await, queue or I/O pressure. |
| If it fails | High utilization metrics differ by device; compare latency and workload behavior. |
| Safe recovery | Remove the lab load or tuning override, restore the recorded sysctl/unit configuration, and prove latency and pressure return to baseline. |
Case 4: Diagnose network application delay
Measure sockets, retransmissions and application timing as separate layers. Compare with a workload baseline and change one causal constraint at a time.
ss -s
nstat -az | grep -i retrans
sar -n DEV,TCP,ETCP 1 5
curl -sS -o /dev/null -w 'dns=%{time_namelookup} connect=%{time_connect} first=%{time_starttransfer} total=%{time_total}\n' https://example.com/
tracepath example.com| Checkpoint | What to establish |
|---|---|
| Expected result | The delay is assigned to name resolution, connection, server response or transfer with matching host counters. |
| If it fails | One curl timing is a sample; use repeated tests and application monitoring before concluding. |
| Safe recovery | Remove the lab load or tuning override, restore the recorded sysctl/unit configuration, and prove latency and pressure return to baseline. |
Independent practice tasks
- Create controlled CPU, memory and I/O load separately and identify each.
- Use tuned-adm to inspect, not blindly change, the active profile.
- Set a systemd resource limit on a lab service and prove its effect.
- Write a slow-server incident report with before/after evidence.
For this lesson on RHEL Performance Troubleshooting, complete each task without copying the worked command sequence. Record the initial state, exact change, verification, negative test and recovery command. A task is unfinished if it works now but does not survive a reboot where persistence is required.
Troubleshooting by symptom
| Symptom | Inspect first | Defensible next action |
|---|---|---|
| High load, idle CPU | I/O wait, blocked tasks and wait channel | Find storage/lock/dependency wait |
| Swap used, no active paging | Paging rate and PSI over time | Do not disable swap solely from used bytes |
| Disk 100% busy | await, queue, throughput and workload pattern | Identify device/workload and latency objective |
| Network “slow” | retransmits, drops, RTT, socket queues and app timing | Locate host/path/protocol layer |
| Fix improves average only | tail latency, errors and dropped work | Use complete service-level acceptance |
Unsafe operations and recovery boundaries
- Unsafe: dropping caches on production changes global memory behavior and can create an I/O storm.
- Unsafe: running unbounded benchmarks competes with the workload and can fill disk or network.
- Unsafe: killing the largest process destroys evidence and may terminate the healthy workload rather than the cause.
Rewritten knowledge checks
Guided lab and acceptance test
- Create a baseline for a small CPU and file-I/O workload.
- Introduce one bounded CPU worker and observe per-core and PSI evidence.
- Introduce a capped file-I/O job on disposable storage and inspect device latency.
- Create memory pressure within a cgroup limit and observe PSI without affecting the host broadly.
- Write a symptom-to-layer diagnosis for each case.
- Stop lab workloads and prove resources, files and units are removed.