Lesson 012 · Jenkins Learning Path

Monitor Jenkins, Audit Changes and Plan Build Capacity

· Published · 6 min read

Labelled Jenkins CI/CD path separating untrusted pull request validation from trusted immutable artifact approval deployment monitoring and rollback

Jenkins performance is a queueing and control-plane problem before it is a CPU graph. Operators must separate demand waiting in the queue, agent provisioning, executor occupancy, controller responsiveness, storage latency, plugin failures and downstream service time.

Start from a known controller checkpoint

Use the accepted controller and agent checkpoint from the prior lesson. Record the Jenkins version, Java runtime, active configuration, plugin inventory and current Git revision before changing this boundary.

java -version
sudo systemctl is-active jenkins
sudo journalctl -u jenkins -n 50 --no-pager
curl -fsS http://127.0.0.1:8080/login >/dev/null

Understand the operating boundary

DecisionImplementationEvidence
ScopeName the controller, folder, job, node and environmentThe selected boundary is visible
InputUse reviewed source and scoped credentialsRevision and credential ID are attributable
ExecutionSet label, timeout and concurrency policyQueue and node evidence match intent
RecoveryPreserve the last known working stateRollback is rehearsed before promotion

Prepare the lab

Record a normal-hour baseline and a known busy period. Identify controller, static agents, cloud agents, storage and reverse proxy separately. Define service objectives for UI/API availability, queue delay and critical Pipeline completion.

uptime
free -h
df -hT /var/lib/jenkins
df -ih /var/lib/jenkins
sudo systemctl status jenkins --no-pager
sudo journalctl -u jenkins --since '-1 hour' --no-pager
curl -fsS http://127.0.0.1:8080/login >/dev/null

Implement it step by step

  1. Measure queue length and age by required label, not only total count.
  2. Track online/offline agents, executor occupancy, provisioning delay and failed launches.
  3. Observe controller JVM heap, garbage collection, threads, HTTP latency and restart causes.
  4. Alert on JENKINS_HOME space and inode headroom before writes fail.
  5. Retain job configuration, credential, security, plugin and node changes with actor and time through an approved audit mechanism.
  6. Correlate Pipeline stage duration with external SCM, registry, test and deployment services.

Adding executors does not create CPU, memory, disk bandwidth or license capacity. It may increase contention and make every build slower. A label-specific queue with idle executors often means the idle nodes do not satisfy the label, are reserved by trust policy or cannot provision the required environment.

# Read-only API samples with a scoped user token
curl --fail --user 'observer:API_TOKEN' \
 'https://jenkins.example.test/queue/api/json?tree=items[id,inQueueSince,why,task[name,url]]'
curl --fail --user 'observer:API_TOKEN' \
 'https://jenkins.example.test/computer/api/json?tree=computer[displayName,offline,temporarilyOffline,numExecutors,busyExecutors]'
# Host evidence
pid=$(systemctl show -p MainPID --value jenkins)
ps -p "$pid" -o pid,etimes,%cpu,%mem,rss,vsz,cmd
sudo ls -l /proc/"$pid"/fd | wc -l

Capacity decisions start with a workload model: arrival rate, typical and high-percentile duration, label, concurrency safety, resource request and time waiting for external systems. Separate trusted deployment executors from elastic test capacity. Keep controller executors at zero and give ephemeral agents startup-time and failure metrics so provisioning delay is visible rather than charged to the build stage.

Audit data needs a retention and access policy. Console logs show job execution but may not explain who changed authorization, a credential, a plugin or a node. Use supported audit logging or external identity/proxy events, synchronize time and protect logs from the administrators whose actions they record where separation of duties requires it.

Verify the positive path

Inject one bounded queue condition by taking the lab agent temporarily offline, one disk threshold alert with a small test filesystem or synthetic metric, and one failed agent launch. Confirm each alert identifies the layer and clears after recovery.

date -u
curl --fail --user 'observer:API_TOKEN' 'https://jenkins.example.test/queue/api/json'
curl --fail --user 'observer:API_TOKEN' 'https://jenkins.example.test/computer/api/json'
sudo journalctl -u jenkins --since '-15 minutes' --no-pager
findmnt -T /var/lib/jenkins
df -h /var/lib/jenkins; df -i /var/lib/jenkins
CheckExpected evidenceReject when
ConfigurationMatches reviewed sourceUI drift or unresolved placeholder remains
ExecutionRuns on the intended isolated nodeController or wrong trust zone executes code
EvidenceRevision, result and outputs are retainedGreen status has no attributable output
RecoveryKnown state can be restored and verifiedRecovery depends on an improvised manual edit

A long queue appears while total executors are idle

Create a job requiring linux-deploy while only linux-build nodes are online. The correct diagnosis is capability or trust mismatch, not more global executors. Retain the queue reason, bring the approved deploy node online and prove the item schedules there.

# Pipeline lab stage
agent { label 'linux-deploy' }
options { timeout(time: 5, unit: 'MINUTES') }
# Observe queue why field; do not relax the label to an untrusted node

Troubleshoot by failed layer

SymptomInspectCorrection
Queued or unavailableLabel, executor, node and networkRepair the failed scheduling or transport layer
Configuration rejectedController log, syntax and plugin ownershipCorrect source; do not bypass validation
Job fails unexpectedlyFirst causal console error and agent logsFix one layer and rerun the smallest scope
Second run differsMutable dependency, workspace or UI driftPin inputs and remove hidden retained state

Unsafe shortcuts

  • Unsafe: granting administrator access to avoid designing a narrow permission removes accountability.
  • Unsafe: binding protected secrets around untrusted repository code permits exfiltration despite masking.
  • Unsafe: changing controller state without a verified backup and rollback turns a small error into an outage.

Operate, recover and retain evidence

Own the configuration, plugin and credential dependencies explicitly. Record the controller version, Git revision, immutable tool or artifact identity, initiator, approver and acceptance result. Rehearse the failure path on a disposable controller before adopting it as production procedure.

sudo systemctl stop jenkins
sudo rsync -aHAX --numeric-ids /backup/jenkins-known-good/ /var/lib/jenkins/
sudo restorecon -RF /var/lib/jenkins 2>/dev/null || true
sudo systemctl start jenkins
sudo journalctl -u jenkins -n 150 --no-pager

Worked use cases

SituationDesign choiceAcceptance
Lab rolloutApply to one disposable controller or folderPositive, negative and recovery results are retained
Team rolloutPromote the same reviewed revisionPermissions and behavior remain consistent
Production changeUse backup, change window and acceptanceFailure stays bounded and rollback is tested

Knowledge checks

What is queue age?
Time runnable work has waited before receiving a suitable executor.
Why monitor inodes?
Jenkins can fail to create files even when byte capacity remains.
Can more executors reduce throughput?
No; beyond resource capacity they increase contention.
Why keep controller configuration in source?
It provides review, attribution and repeatable recovery.
Why record the exact Jenkins version?
Core and plugin behavior depends on the running baseline.
Does a successful process prove service acceptance?
No; verify the user-facing or downstream result.
Why test one negative case?
It proves the control rejects an invalid or unauthorized path.
Why use a disposable rehearsal?
Controller changes can prevent the same interface from repairing itself.

Independent lab

  1. Capture the starting version, configuration and plugin evidence.
  2. Implement the change on a disposable controller or folder.
  3. Run one successful case and one controlled failure.
  4. Restore the accepted state and prove service behavior, not only process state.
  5. Repeat from a clean source checkout without relying on remembered UI actions.

Official references

Advertisement