Scenario 06 — Backup verification / restore drills¶
An untested backup is a hope, not a guarantee. A backup you have never restored might be encrypted with a lost password, pointed at a dead bucket, or quietly capturing an empty volume — and you find out during the outage. A verification drill catches that on your schedule: periodically restore the latest snapshot into a throwaway PVC, assert it completed, then clean up.
There are two layers of verification, and they answer different questions:
| Layer | What it proves | How |
|---|---|---|
Built-in verification on the SnapshotPolicy |
The repository blobs are intact and (optionally) a scratch-restore of the latest snapshot succeeds. | A field on the recipe — the operator runs it on its own cron. Start here. |
| A full restore drill (this scenario) | An end-to-end Restore into a real PVC completes, mountable and app-checkable. |
A CronJob that creates Restore CRs. The deepest, app-level proof. |
Start with the built-in capability; reach for the drill when you want a true end-to-end restore (and an app-level check on the restored data).
Built-in verification (SnapshotPolicy.spec.verification)¶
Kopiur has first-class, opt-in verification. Add a verification
block to the recipe and the operator runs it on a schedule — no CronJob, no
extra RBAC:
spec:
verification:
quick: # blob-level `kopia snapshot verify`, often
schedule: { cron: "0 4 * * *", jitter: 30m }
parallel: 8 # --parallel: verification parallelism (kopia default: 8)
maxErrors: 0 # --max-errors: stop after this many errors (0 = stop at first)
deep: # scratch-restore the latest snapshot into a throwaway volume, rarely
schedule: { cron: "0 5 * * 0", jitter: 1h }
capacity: 100Gi # size a fresh ephemeral PVC for the restore (omit = emptyDir)
storageClassName: fast-ssd # StorageClass for that PVC (omit = cluster default)
parallel: 4 # restore --parallel: deep verify IS a restore under the hood
successExpr: "stats.files > 0 && stats.errors == 0" # CEL pass/fail predicate
verifyFilesPercent: 10 # how much of each file `quick` actually reads
quickis a cheap, frequent blob-level integrity check;deepis a rare full scratch-restore into a throwaway volume (then discarded). Both tiers nest their cron underschedule:(quick.schedule,deep.schedule);deepadditionally carries the scratch-volume knobs below. Schedule each independently.quicktuning —parallel/fileParallelism/fileQueueLength/maxErrorsmap directly ontokopia snapshot verify's own flags; all optional, absent leaves kopia's default.deep.parallelis the analogous knob for the scratch-restore (restore --parallel).- Gated until there's something to verify — a brand-new policy does not spawn a verify Job before it has a first successful backup (or, for an adopted repository, discovered snapshots already in it). See Backups → verification scheduling.
deepscratch sizing — setcapacityto provision a fresh generic-ephemeral PVC (auto-deleted with the Job), sized to hold the restored snapshot;storageClassNameplaces it (omit = cluster default). Omitcapacityand scratch is a node-ephemeralemptyDir— zero-config, but bounded by node disk, so prefer a sized PVC for large snapshots. AstorageClassNamewith nocapacityis a no-op (anemptyDirhas no StorageClass) — the operator flags it as aScratchStorageClassIgnoredcondition on theSnapshotPolicy. Set the size/class once for all policies viamoverDefaults.scratchon the repository;verification.deephere overrides it field-wise.successExpris a CEL predicate over the result (stats{files,bytes,errors},snapshot, and — deep only —restored{files,checksumMatches}) — it kills the silent "0 files" success. It is validated at admission, so a typo is rejected onkubectl apply.verifyFilesPercentsets how much of each filequickreads in full (the rest is checked at the index/blob level).
The most recent successful verify lands in status.lastVerified, shows in the
LAST-VERIFIED printer column, and exports the kopiur_snapshot_verified_timestamp
metric — alert on its staleness exactly like kopiur_snapshot_last_success_timestamp_seconds.
The full field reference is in Backups → verification.
The full-restore drill¶
When you want the deepest, app-level proof — a real Restore into a real PVC you
can mount and check — run a drill. Kopiur has no RestoreSchedule kind (restores
are one-shot operations), so the cadence is a tiny CronJob that creates Restore
CRs. The bundle has two halves you can use independently.
Half A — run one drill by hand¶
The first object in the file is a plain Restore you can kubectl apply right
now: latest snapshot → throwaway PVC, with onMissingSnapshot: Fail so a drill
that finds nothing is a failed drill (that's the alarm). Watch it, eyeball the
result, delete the PVC.
Half B — the automated nightly drill¶
The rest of the file is a ServiceAccount + Role + RoleBinding + CronJob.
Each night the CronJob creates a timestamped drill Restore, waits for it to
reach Completed, then deletes both the Restore and its throwaway PVC. If the
restore fails or times out, the Job fails — which is what your monitoring alerts
on.
Least-privilege RBAC
The drill runner can only create/get/delete Restore CRs and delete PVCs
in its own namespace — it is not the operator and holds none of the
operator's repository or mover permissions. The CronJob image is the upstream
registry.k8s.io/kubectl (any image with kubectl ≥ 1.23 works — it uses
kubectl wait --for=jsonpath).
# Scenario 06 — Backup verification / restore drills (an untested backup isn't one)
#
# A backup you have never restored is a hope, not a guarantee. There are two ways
# to prove restorability on a schedule:
#
# 0) FIRST-CLASS verification (RECOMMENDED). Put a `verification:`
# block on the SnapshotPolicy and the operator runs it for you: a frequent
# blob-level `kopia snapshot verify` (quick) and a rarer scratch-restore test
# (deep) into an ephemeral PVC. An optional CEL `successExpr` asserts the
# result is good (killing the silent "0 files" success). Surfaces
# `status.lastVerified` + the kopiur_snapshot_verified_timestamp metric.
# A/B) DIY DRILL. If you want full control of the drill, a tiny CronJob that
# creates Restore CRs works too (shown below) — restore into a THROWAWAY PVC,
# assert Completed, clean up.
#
# This file has three parts:
# 0) The SnapshotPolicy with a native `verification:` block (the easy path).
# A) A run-one-now drill Restore you can `kubectl apply` by hand.
# B) The automation: a ServiceAccount + RBAC + CronJob that creates a
# timestamped drill each night, waits for Completed, and deletes both the
# Restore and its throwaway PVC.
#
# Pair this with alerting on the operator's metrics (PromQL examples are in the
# docs page): kopiur_snapshot_last_success_timestamp_seconds going stale, and the
# drill CronJob's own failures.
#
# Field shapes verified against crates/api (verification.{quick,deep,successExpr,
# verifyFilesPercent}; source.fromPolicy, target.pvc). The RBAC/CronJob are plain
# core/k8s objects.
---
# ── 0) First-class verification on the recipe (the recommended path) ──────────
apiVersion: kopiur.home-operations.com/v1alpha1
kind: SnapshotPolicy
metadata:
name: postgres-data
namespace: billing
spec:
repository:
name: postgres-primary
sources:
- pvc:
name: postgres-data
retention:
keepDaily: 14
# The operator verifies restorability on this cadence — no CronJob needed.
verification:
# Quick: blob-level `kopia snapshot verify`, often. Its cron lives under
# `quick.schedule` (matching `deep.schedule`).
quick:
schedule:
cron: "0 4 * * *"
jitter: 30m
# Tuning knobs map directly onto `kopia snapshot verify`'s own flags; all
# optional, absent leaves kopia's default.
parallel: 8 # --parallel (kopia default: 8)
fileParallelism: 4 # --file-parallelism
fileQueueLength: 20000 # --file-queue-length (kopia default: 20000)
maxErrors: 0 # --max-errors (kopia default: 0, stop at the first error)
# Deep: scratch-restore the latest snapshot into a throwaway volume, weekly.
deep:
schedule:
cron: "0 5 * * 0"
jitter: 1h
# capacity provisions a fresh ephemeral PVC (auto-deleted with the Job),
# sized to hold the restored snapshot. Omit it and scratch is a node-ephemeral
# emptyDir. storageClassName only applies when capacity is set (omit = default).
capacity: 100Gi
storageClassName: fast-ssd
# restore --parallel: deep verify IS a restore under the hood.
parallel: 4
# CEL pass/fail over the result — a verify with 0 files or any error FAILS.
successExpr: "stats.files > 0 && stats.errors == 0"
verifyFilesPercent: 10 # how much of each file `quick` reads fully
---
# ── A) Run a single drill by hand ────────────────────────────────────────────
apiVersion: kopiur.home-operations.com/v1alpha1
kind: Restore
metadata:
name: postgres-drill
namespace: billing
spec:
source:
fromPolicy:
name: postgres-data
offset: 0 # latest snapshot
target:
pvc:
name: postgres-drill # throwaway — delete it after you've checked it
storageClassName: fast-ssd
capacity: 100Gi
accessModes:
- ReadWriteOnce
policy:
# A drill that finds NO snapshot is a FAILED drill — that's the whole alarm.
onMissingSnapshot: Fail
waitTimeout: 10m
---
# ── B) Automated nightly drill ───────────────────────────────────────────────
apiVersion: v1
kind: ServiceAccount
metadata:
name: kopiur-drill-runner
namespace: billing
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: kopiur-drill-runner
namespace: billing
rules:
# Create/observe/clean up Restore CRs in this namespace.
- apiGroups: ["kopiur.home-operations.com"]
resources: ["restores"]
verbs: ["create", "get", "list", "watch", "delete"]
# Delete the throwaway PVC the drill created.
- apiGroups: [""]
resources: ["persistentvolumeclaims"]
verbs: ["get", "list", "delete"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: kopiur-drill-runner
namespace: billing
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: Role
name: kopiur-drill-runner
subjects:
- kind: ServiceAccount
name: kopiur-drill-runner
namespace: billing
---
apiVersion: batch/v1
kind: CronJob
metadata:
name: kopiur-restore-drill
namespace: billing
spec:
schedule: "30 4 * * *" # nightly, after the 02:xx backups have landed
concurrencyPolicy: Forbid
jobTemplate:
spec:
backoffLimit: 0 # a failed drill should stay failed and alert, not retry-mask
activeDeadlineSeconds: 1800
template:
spec:
serviceAccountName: kopiur-drill-runner
restartPolicy: Never
containers:
- name: drill
# Any image with kubectl >= 1.23 (for `wait --for=jsonpath`).
image: registry.k8s.io/kubectl:v1.33.0
command: ["/bin/sh", "-c"]
args:
- |
set -euo pipefail
ts="$(date +%Y%m%d-%H%M%S)"
name="postgres-drill-${ts}"
ns="billing"
cleanup() {
kubectl delete restore "$name" -n "$ns" --ignore-not-found
kubectl delete pvc "$name" -n "$ns" --ignore-not-found
}
trap cleanup EXIT
cat <<EOF | kubectl apply -f -
apiVersion: kopiur.home-operations.com/v1alpha1
kind: Restore
metadata:
name: ${name}
namespace: ${ns}
spec:
source:
fromPolicy:
name: postgres-data
offset: 0
target:
pvc:
name: ${name}
storageClassName: fast-ssd
capacity: 100Gi
accessModes: [ReadWriteOnce]
policy:
onMissingSnapshot: Fail
waitTimeout: 10m
EOF
# Succeeds only if the restore reaches Completed; a Failed phase
# or a timeout makes the Job (and the drill) fail loudly.
kubectl wait --for=jsonpath='{.status.phase}'=Completed \
"restore/${name}" -n "$ns" --timeout=15m
echo "drill ${name} restored successfully"
Alert on the operator's metrics¶
The drill proves a full restore works. Pair it with cheap, always-on alerts on
the operator's Prometheus metrics (all kopiur_*, scraped from /metrics — see
Observability) so you also catch a backup that simply
stopped running:
# A backup hasn't succeeded in over 26h (a missed nightly + margin).
time() - kopiur_snapshot_last_success_timestamp_seconds > 26 * 3600
# A schedule is racking up consecutive failures.
kopiur_snapshot_consecutive_failures > 2
# Built-in `verification` hasn't passed in over a week (deep verify is weekly + margin).
time() - kopiur_snapshot_verified_timestamp > 8 * 24 * 3600
And alert on the drill itself by watching the CronJob's Job failures (e.g.
kube_job_status_failed{job_name=~"kopiur-restore-drill.*"} > 0 if you run
kube-state-metrics), or on the drill restore's duration via
kopiur_restore_duration_seconds.
What "verified" should mean to you
Completed proves kopia could decrypt the repo and write the bytes back. For the
strongest guarantee, go one level further: have the drill (or a follow-on Job)
mount the restored PVC and run an app-level check — pg_verifybackup, a checksum
of known files, a test query. A restore that completes but produces unreadable
data is rare, but a drill that opens the data rules it out entirely.
See also¶
- Backups → verification — the built-in
quick/deep/successExprfield reference. - Observability — the full
kopiur_*metric surface and how to scrape it. - Restores —
fromPolicy,onMissingSnapshot, and restore phases. - Scenario 02 — recover from data loss — the real restore your drills are rehearsing.