Design Preflight, Post Actions, Timeouts and Recovery
Pre steps prove that a build or deployment is safe to start. Post steps publish evidence, clean resources, notify operators and restore service after success or failure. Keep substantial logic in reviewed repository scripts while Jenkins defines ordering, credentials, conditions and evidence.
Map the lifecycle
| Phase | Examples | Failure behavior |
|---|---|---|
| Preflight | Required variables, tools, target, disk, lock, change window | Stop before mutation |
| Main action | Build, test, publish, deploy | Fail fast and preserve diagnostics |
| Post success | Health proof, release record, success notification | Mark failure if acceptance is incomplete |
| Post failure | Logs, rollback, incident notification | Do not hide the original failure |
| Always/cleanup | Reports, temporary file removal, lock release | Must be idempotent |
Create a repository preflight script
#!/usr/bin/env bash
set -euo pipefail
: "${DEPLOY_ENV:?DEPLOY_ENV is required}"
: "${ARTIFACT:?ARTIFACT is required}"
case "$DEPLOY_ENV" in dev|staging|production) ;; *) echo 'invalid environment' >&2; exit 2;; esac
test -r "$ARTIFACT"
command -v sha256sum >/dev/null
command -v curl >/dev/null
sha256sum -c "${ARTIFACT}.sha256"
df -Pk . | awk 'NR==2 && $4 < 1048576 {exit 1}'
install -d -m 0755 ci
chmod 0755 ci/preflight.sh
git add ci/preflight.sh Jenkinsfile
git commit -m 'Add deployment preflight checks'
Use stage-level pre and post behavior
pipeline {
agent { label 'deploy' }
options {
timestamps()
timeout(time: 30, unit: 'MINUTES')
disableConcurrentBuilds()
}
environment {
DEPLOY_ENV = 'staging'
ARTIFACT = "dist/app-${BUILD_NUMBER}.tgz"
}
stages {
stage('Preflight') {
steps { sh 'set -eu; ./ci/preflight.sh' }
}
stage('Deploy') {
steps {
retry(2) { sh 'set -eu; ./ci/deploy.sh' }
}
post {
success { sh 'set -eu; ./ci/verify-deployment.sh' }
failure { sh 'set +e; ./ci/collect-deploy-diagnostics.sh' }
}
}
}
post {
always {
junit testResults: 'reports/*.xml', allowEmptyResults: true
archiveArtifacts artifacts: 'diagnostics/**', allowEmptyArchive: true
}
failure { echo "FAILED: ${env.JOB_NAME} #${env.BUILD_NUMBER}" }
cleanup { deleteDir() }
}
}
Retry only operations known to be idempotent or safely resumable. Retrying a partially applied database migration or payment action can multiply damage.
Understand Declarative post conditions
always: reports and evidence needed for every result.success: acceptance checks and success-only publication.failure: diagnostics, rollback request and alert.unstable: test or quality result that did not fully pass.changed: notify when the result differs from the prior run.cleanup: final cleanup after other post conditions.
Use shell traps inside one script
#!/usr/bin/env bash
set -euo pipefail
tmp_dir="$(mktemp -d)"
cleanup() {
rc=$?
rm -rf -- "$tmp_dir"
exit "$rc"
}
trap cleanup EXIT INT TERM
./ci/render-config.sh >"$tmp_dir/config"
./ci/apply.sh "$tmp_dir/config"
The trap preserves the action's exit status and cleans its own validated temporary directory. Jenkins post still handles cross-stage evidence and notifications.
Add a production approval
stage('Approve production') {
when { branch 'main' }
options { timeout(time: 30, unit: 'MINUTES') }
input { message 'Promote the tested artifact to production?' }
steps { echo 'Approval recorded by Jenkins' }
}
The approval must reference the already-tested artifact digest. Do not rebuild after approval.
Unsafe patterns
Unsafe: ./deploy.sh || true converts a failed deployment into a successful step. Use set +e only for best-effort diagnostics in a post-failure block, and retain the original build result.
Unsafe: placing secrets directly in Groovy-interpolated shell strings can expose them before masking. Bind credentials to environment variables and use a single-quoted shell script so the shell expands them at execution.
Troubleshoot post actions
| Problem | Cause | Correction |
|---|---|---|
| Post step hides primary error | Cleanup also fails | Make cleanup idempotent and capture diagnostic failure separately. |
| Notification duplicates | Stage and Pipeline handlers overlap | Give each handler one responsibility. |
| Timeout leaves process | Child ignores termination or daemonizes | Keep processes attached and add server-side reconciliation. |
| Retry repeats mutation | Operation is not idempotent | Remove retry; add status detection or transaction/rollback. |
Preserve the original failure while recovery runs
Preflight must reject missing inputs, wrong targets, unavailable capacity and invalid artifacts before mutation. Post actions should publish evidence even after failure, but a successful cleanup or rollback must not silently turn a failed release green. Put complex behavior in versioned repository scripts with strict error handling and unit tests; keep Jenkins responsible for order, credentials, timeouts and result policy. Distinguish retryable transport errors from deterministic validation failures so retry does not repeat a harmful action.
stage('Deploy') {
options { timeout(time: 15, unit: 'MINUTES') }
steps {
lock(resource: "deploy-${params.ENVIRONMENT}") {
sh './ci/preflight.sh'
sh './ci/deploy.sh'
sh './ci/acceptance.sh'
}
}
post {
unsuccessful { sh './ci/collect-diagnostics.sh' }
cleanup { deleteDir() }
}
}Evidence and failure exercise
| Layer | Evidence | Failure response |
|---|---|---|
| Preflight | Target and digest checked before mutation | Stop without deployment |
| Action | Release ID and bounded command result | Preserve first causal failure |
| Acceptance | Client-side health and version response | Run reviewed rollback |
| Cleanup | Lock, temp files and reports handled | Do not erase diagnostics before publish |
Make acceptance fail after a successful release switch. Capture diagnostics, execute rollback and deliberately leave the Jenkins build failed because the requested release did not succeed. A later recovery build may be green, but it must reference the failed build and restored release identity.
Independent operating checks
- Cancel the approval and confirm no executor or environment lock leaks.
- Cause a timeout and verify child processes terminate.
- Make diagnostics collection fail and preserve the original error.
- Run preflight twice and prove it creates no state.