Scenario 06 — Backup verification / restore drills¶
An untested backup is a hope, not a guarantee. A backup you have never restored might be encrypted with a lost password, pointed at a dead bucket, or quietly capturing an empty volume. You find out during the outage.
A verification drill catches that on your schedule instead: periodically restore the latest snapshot into a throwaway PVC, assert it completed, then clean up.
There are two layers of verification, and they answer different questions:
| Layer | What it proves | How |
|---|---|---|
Built-in verification on the SnapshotPolicy |
The repository blobs are intact and, optionally, a scratch-restore of the latest snapshot succeeds. | A field on the recipe. The operator runs it on its own cron. Start here. |
| A full restore drill (this scenario) | An end-to-end Restore into a real PVC completes, mountable and app-checkable. |
A CronJob that creates Restore CRs. The deepest, app-level proof. |
Start with the built-in capability. Reach for the drill when you want a true end-to-end restore, and an app-level check on the restored data.
Built-in verification (SnapshotPolicy.spec.verification)¶
Kopiur has first-class, opt-in verification. Add a verification block to the recipe and the operator runs it on a schedule. No CronJob, no extra RBAC:
spec:
verification:
quick: # blob-level `kopia snapshot verify`, often
schedule: { cron: "0 4 * * *", jitter: 30m }
parallel: 8 # --parallel: verification parallelism (kopia default: 8)
maxErrors: 0 # --max-errors: stop after this many errors (0 = stop at first)
deep: # scratch-restore the latest snapshot into a throwaway volume, rarely
schedule: { cron: "0 5 * * 0", jitter: 1h }
capacity: 100Gi # size a fresh ephemeral PVC for the restore (omit = emptyDir)
storageClassName: fast-ssd # StorageClass for that PVC (omit = cluster default)
parallel: 4 # restore --parallel: deep verify IS a restore under the hood
successExpr: "stats.files > 0 && stats.errors == 0" # CEL pass/fail predicate
verifyFilesPercent: 10 # how much of each file `quick` actually reads
quickanddeepare two tiers.quickis a cheap, frequent blob-level integrity check.deepis a rare full scratch-restore into a throwaway volume, which is then discarded. Both tiers nest their cron underschedule:, asquick.scheduleanddeep.schedule.deepalso carries the scratch-volume knobs below. Schedule each independently.quicktuning.parallel,fileParallelism,fileQueueLength, andmaxErrorsmap directly ontokopia snapshot verify's own flags. All are optional; leaving one out keeps kopia's default.deep.parallelis the matching knob for the scratch-restore, mapping torestore --parallel.- Verification waits until there is something to verify. A brand-new policy does not spawn a verify Job before it has a first successful backup. On an adopted repository, discovered snapshots already in it count too. See Backups → verification scheduling.
deepscratch sizing. Setcapacityto provision a fresh generic-ephemeral PVC, sized to hold the restored snapshot; it is auto-deleted with the Job.storageClassNameplaces it, and omitting it uses the cluster default. Omitcapacityand scratch is a node-ephemeralemptyDir: zero-config, but bounded by node disk, so prefer a sized PVC for large snapshots. AstorageClassNamewith nocapacitydoes nothing, because anemptyDirhas no StorageClass, and the operator flags it as aScratchStorageClassIgnoredcondition on theSnapshotPolicy. Set the size and class once for all policies viamoverDefaults.scratchon the repository;verification.deephere overrides it field by field.successExpris a CEL predicate over the result. It seesstats{files,bytes,errors},snapshot, and for deep verification only,restored{files,checksumMatches}. It kills the silent "0 files" success. It is validated at admission, so a typo is rejected onkubectl apply.verifyFilesPercentsets how much of each filequickreads in full. The rest is checked at the index and blob level.
The most recent successful verify lands in status.lastVerified, shows in the LAST-VERIFIED printer column, and exports the kopiur_snapshot_verified_timestamp metric. Alert on its staleness exactly like kopiur_snapshot_last_success_timestamp_seconds. The full field reference is in Backups → verification.
The full-restore drill¶
When you want the deepest, app-level proof, meaning a real Restore into a real PVC you can mount and check, run a drill.
Kopiur has no RestoreSchedule kind, because restores are one-shot operations. So the cadence comes from a tiny CronJob that creates Restore CRs. The bundle has two halves you can use independently.
Half A — run one drill by hand¶
The first object in the file is a plain Restore you can kubectl apply right now: latest snapshot into a throwaway PVC, with onMissingSnapshot: Fail so a drill that finds nothing is a failed drill. That failure is the alarm. Watch it, eyeball the result, delete the PVC.
Half B — the automated nightly drill¶
The rest of the file is a ServiceAccount, a Role, a RoleBinding, and a CronJob. Each night the CronJob creates a timestamped drill Restore, waits for it to reach Completed, then deletes both the Restore and its throwaway PVC. If the restore fails or times out, the Job fails, and that is what your monitoring alerts on.
Least-privilege RBAC
The drill runner can only create, get, and delete Restore CRs, and delete PVCs, in its own namespace. It is not the operator and holds none of the operator's repository or mover permissions.
The CronJob image is the upstream registry.k8s.io/kubectl. Any image with kubectl ≥ 1.23 works, because it uses kubectl wait --for=jsonpath.
# Scenario 06 — Backup verification / restore drills (an untested backup isn't one)
#
# A backup you have never restored is a hope, not a guarantee. There are two ways
# to prove restorability on a schedule:
#
# 0) FIRST-CLASS verification (RECOMMENDED). Put a `verification:`
# block on the SnapshotPolicy and the operator runs it for you: a frequent
# blob-level `kopia snapshot verify` (quick) and a rarer scratch-restore test
# (deep) into an ephemeral PVC. An optional CEL `successExpr` asserts the
# result is good (killing the silent "0 files" success). Surfaces
# `status.lastVerified` + the kopiur_snapshot_verified_timestamp metric.
# A/B) DIY DRILL. If you want full control of the drill, a tiny CronJob that
# creates Restore CRs works too (shown below) — restore into a THROWAWAY PVC,
# assert Completed, clean up.
#
# This file has three parts:
# 0) The SnapshotPolicy with a native `verification:` block (the easy path).
# A) A run-one-now drill Restore you can `kubectl apply` by hand.
# B) The automation: a ServiceAccount + RBAC + CronJob that creates a
# timestamped drill each night, waits for Completed, and deletes both the
# Restore and its throwaway PVC.
#
# Pair this with alerting on the operator's metrics (PromQL examples are in the
# docs page): kopiur_snapshot_last_success_timestamp_seconds going stale, and the
# drill CronJob's own failures.
#
# Field shapes verified against crates/api (verification.{quick,deep,successExpr,
# verifyFilesPercent}; source.fromPolicy, target.pvc). The RBAC/CronJob are plain
# core/k8s objects.
---
# ── 0) First-class verification on the recipe (the recommended path) ──────────
apiVersion: kopiur.home-operations.com/v1alpha1
kind: SnapshotPolicy
metadata:
name: postgres-data
namespace: billing
spec:
repository:
name: postgres-primary
sources:
- pvc:
name: postgres-data
retention:
keepDaily: 14
# The operator verifies restorability on this cadence — no CronJob needed.
verification:
# Quick: blob-level `kopia snapshot verify`, often. Its cron lives under
# `quick.schedule` (matching `deep.schedule`).
quick:
schedule:
cron: "0 4 * * *"
jitter: 30m
# Tuning knobs map directly onto `kopia snapshot verify`'s own flags; all
# optional, absent leaves kopia's default.
parallel: 8 # --parallel (kopia default: 8)
fileParallelism: 4 # --file-parallelism
fileQueueLength: 20000 # --file-queue-length (kopia default: 20000)
maxErrors: 0 # --max-errors (kopia default: 0, stop at the first error)
# Deep: scratch-restore the latest snapshot into a throwaway volume, weekly.
deep:
schedule:
cron: "0 5 * * 0"
jitter: 1h
# capacity provisions a fresh ephemeral PVC (auto-deleted with the Job),
# sized to hold the restored snapshot. Omit it and scratch is a node-ephemeral
# emptyDir. storageClassName only applies when capacity is set (omit = default).
capacity: 100Gi
storageClassName: fast-ssd
# restore --parallel: deep verify IS a restore under the hood.
parallel: 4
# CEL pass/fail over the result — a verify with 0 files or any error FAILS.
successExpr: "stats.files > 0 && stats.errors == 0"
verifyFilesPercent: 10 # how much of each file `quick` reads fully
---
# ── A) Run a single drill by hand ────────────────────────────────────────────
apiVersion: kopiur.home-operations.com/v1alpha1
kind: Restore
metadata:
name: postgres-drill
namespace: billing
spec:
source:
fromPolicy:
name: postgres-data
offset: 0 # latest snapshot
target:
pvc:
name: postgres-drill # throwaway — delete it after you've checked it
storageClassName: fast-ssd
capacity: 100Gi
accessModes:
- ReadWriteOnce
policy:
# A drill that finds NO snapshot is a FAILED drill — that's the whole alarm.
onMissingSnapshot: Fail
waitTimeout: 10m
---
# ── B) Automated nightly drill ───────────────────────────────────────────────
apiVersion: v1
kind: ServiceAccount
metadata:
name: kopiur-drill-runner
namespace: billing
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: kopiur-drill-runner
namespace: billing
rules:
# Create/observe/clean up Restore CRs in this namespace.
- apiGroups: ["kopiur.home-operations.com"]
resources: ["restores"]
verbs: ["create", "get", "list", "watch", "delete"]
# Delete the throwaway PVC the drill created.
- apiGroups: [""]
resources: ["persistentvolumeclaims"]
verbs: ["get", "list", "delete"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: kopiur-drill-runner
namespace: billing
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: Role
name: kopiur-drill-runner
subjects:
- kind: ServiceAccount
name: kopiur-drill-runner
namespace: billing
---
apiVersion: batch/v1
kind: CronJob
metadata:
name: kopiur-restore-drill
namespace: billing
spec:
schedule: "30 4 * * *" # nightly, after the 02:xx backups have landed
concurrencyPolicy: Forbid
jobTemplate:
spec:
backoffLimit: 0 # a failed drill should stay failed and alert, not retry-mask
activeDeadlineSeconds: 1800
template:
spec:
serviceAccountName: kopiur-drill-runner
restartPolicy: Never
containers:
- name: drill
# Any image with kubectl >= 1.23 (for `wait --for=jsonpath`).
image: registry.k8s.io/kubectl:v1.33.0
command: ["/bin/sh", "-c"]
args:
- |
set -euo pipefail
ts="$(date +%Y%m%d-%H%M%S)"
name="postgres-drill-${ts}"
ns="billing"
cleanup() {
kubectl delete restore "$name" -n "$ns" --ignore-not-found
kubectl delete pvc "$name" -n "$ns" --ignore-not-found
}
trap cleanup EXIT
cat <<EOF | kubectl apply -f -
apiVersion: kopiur.home-operations.com/v1alpha1
kind: Restore
metadata:
name: ${name}
namespace: ${ns}
spec:
source:
fromPolicy:
name: postgres-data
offset: 0
target:
pvc:
name: ${name}
storageClassName: fast-ssd
capacity: 100Gi
accessModes: [ReadWriteOnce]
policy:
onMissingSnapshot: Fail
waitTimeout: 10m
EOF
# Succeeds only if the restore reaches Completed; a Failed phase
# or a timeout makes the Job (and the drill) fail loudly.
kubectl wait --for=jsonpath='{.status.phase}'=Completed \
"restore/${name}" -n "$ns" --timeout=15m
echo "drill ${name} restored successfully"
Alert on the operator's metrics¶
The drill proves a full restore works. Pair it with cheap, always-on alerts on the operator's Prometheus metrics, all named kopiur_* and scraped from /metrics. See Observability. That way you also catch a backup that simply stopped running:
# A backup hasn't succeeded in over 26h (a missed nightly + margin).
time() - kopiur_snapshot_last_success_timestamp_seconds > 26 * 3600
# A schedule is racking up consecutive failures.
kopiur_snapshot_consecutive_failures > 2
# Built-in `verification` hasn't passed in over a week (deep verify is weekly + margin).
time() - kopiur_snapshot_verified_timestamp > 8 * 24 * 3600
Alert on the drill itself by watching the CronJob's Job failures. With kube-state-metrics that is kube_job_status_failed{job_name=~"kopiur-restore-drill.*"} > 0. You can also alert on the drill restore's duration via kopiur_restore_duration_seconds.
What "verified" should mean to you
Completed proves kopia could decrypt the repository and write the bytes back.
For the strongest guarantee, go one level further. Have the drill, or a follow-on Job, mount the restored PVC and run an app-level check: pg_verifybackup, a checksum of known files, a test query. A restore that completes but produces unreadable data is rare, and a drill that opens the data rules it out entirely.
See also¶
- Backups → verification: the built-in
quick/deep/successExprfield reference. - Observability: the full
kopiur_*metric surface and how to scrape it. - Restores:
fromPolicy,onMissingSnapshot, and restore phases. - Scenario 02, recover from data loss: the real restore your drills are rehearsing.