Installing Kopiur¶
Kopiur is a Kopia-native Kubernetes backup operator (Rust / kube-rs). This guide covers installing the operator with the Helm chart and verifying it.
The chart is published as an OCI artifact at
oci://ghcr.io/home-operations/charts/kopiur — that is the preferred way to
install. Every published version is cosign-signed and ships with all three
component images (controller, webhook, mover) pinned by digest to the exact
release build. The in-repo chart (deploy/helm/kopiur) is the development
copy: its image digests are empty and its tags float, so use it only when
working from a checkout.
Status: alpha — API group
kopiur.home-operations.com, versionv1alpha1. The CRD surface may still change between releases.
Prerequisites¶
- Kubernetes >= 1.24. The deploy-or-restore volume-populator path (
Restore+PVC.spec.dataSourceRef) relies on theAnyVolumeDataSourcefeature, available from 1.24. - Helm 3 or 4.
- A kopia repository backend you can reach: S3/MinIO, Azure Blob, GCS, B2, filesystem (PVC), SFTP, WebDAV, or rclone.
- A CSI snapshot stack (a
snapshot-controllerplus aVolumeSnapshotClassfor your driver) for the defaultcopyMethod: Snapshot. Many distributions bundle it (EKS, GKE, AKS, Talos, k3s add-ons); where yours does not, install the home-operationssnapshot-controllerchart (oci://ghcr.io/home-operations/charts/snapshot-controller). Not needed if everySnapshotPolicysetscopyMethod: Direct. See Copy methods → What it requires. - (Optional) cert-manager — only if you prefer it to manage the admission webhook's certificate. It is not required: by default the operator manages the webhook cert itself (see Webhook TLS).
- (Optional) volume-data-source-validator — recommended alongside CSI populators so a malformed
dataSourceRefis surfaced as an event rather than a silently-stuck PVC. - (Optional) Prometheus Operator — if you want the chart's
ServiceMonitor.
Quickstart¶
# 1. Install the published chart, creating the operator namespace. No extra
# flags needed — the webhook cert is self-managed by default (no
# cert-manager required). Add --version x.y.z to pin a release.
helm install kopiur oci://ghcr.io/home-operations/charts/kopiur \
--namespace kopiur-system --create-namespace
# 2. Wait for rollout. (The webhook pod stays in ContainerCreating until the
# controller mints its serving Secret — a few seconds after the controller
# becomes ready.)
kubectl -n kopiur-system rollout status deploy/kopiur-controller
kubectl -n kopiur-system rollout status deploy/kopiur-webhook
# 3. Confirm the 8 CRDs are registered.
kubectl get crd | grep kopiur.home-operations.com
GitOps install (Flux / Argo)¶
Point your GitOps tool at the same OCI chart — never at the in-repo
deploy/helm/kopiur path, whose image digests are empty and whose tags float
(a git-sourced install renders image tags from the chart's appVersion, so a
checkout can pull a tag that was never published). The published chart pins all
three images by digest.
apiVersion: v1
kind: Namespace
metadata:
name: kopiur-system
---
apiVersion: source.toolkit.fluxcd.io/v1
kind: OCIRepository
metadata:
name: kopiur
namespace: kopiur-system
spec:
interval: 1h
url: oci://ghcr.io/home-operations/charts/kopiur
ref:
tag: 0.8.0 # pin the chart version
---
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: kopiur
namespace: kopiur-system
spec:
interval: 1h
chartRef:
kind: OCIRepository
name: kopiur
# values: {} # override chart values here
The OCIRepository and HelmRelease are namespaced, so kopiur-system must
exist before they apply; that is why the Namespace is included above (Flux's
install.createNamespace only creates the release's target namespace, not
the one the HelmRelease object itself lives in). In a real Flux repo the
namespace usually comes from the parent Kustomization; keep it here so the
snippet applies stand-alone.
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: kopiur
namespace: argocd
spec:
project: default
source:
repoURL: ghcr.io/home-operations/charts # no oci:// prefix here
chart: kopiur
targetRevision: 0.8.0
# helm:
# valuesObject: {} # override chart values here
destination:
server: https://kubernetes.default.svc
namespace: kopiur-system
syncPolicy:
syncOptions:
- CreateNamespace=true
helm upgrade and GitOps reconciles never update the CRDs
The 8 CRDs ship in the chart's special crds/ directory, which Helm installs
once and never touches on upgrade (see CRD lifecycle below).
A GitOps tool with a CreateReplace / server-side-apply CRD policy handles
schema changes automatically; otherwise apply deploy/crds/ yourself.
Webhook TLS¶
The admission webhook always serves TLS, and the API server must trust it. webhook.tls.mode chooses how the serving certificate is provisioned:
webhook.tls.mode |
What happens | Needs cert-manager? |
|---|---|---|
self (default) |
The operator mints its own CA + serving cert, writes the Secret, injects the caBundle, and auto-rotates the cert (zero-downtime hot-reload). |
No |
cert-manager |
cert-manager issues the cert; its ca-injector populates the caBundle. |
Yes |
manual |
You pre-create the Secret and supply webhook.caBundle. |
No |
self (default)¶
Nothing to configure — this is the helm install above. The operator grants itself the minimal extra RBAC to write its one serving Secret and patch the caBundle of its two webhook configurations (resourceName-scoped).
cert-manager¶
helm install kopiur oci://ghcr.io/home-operations/charts/kopiur \
--namespace kopiur-system \
--set webhook.tls.mode=cert-manager
# Optionally point at your own Issuer/ClusterIssuer instead of the chart's
# self-signed one: --set webhook.certManager.issuerRef.name=my-issuer
manual¶
# create a kubernetes.io/tls Secret named per webhook.tls.secretName, then:
helm install kopiur oci://ghcr.io/home-operations/charts/kopiur \
--namespace kopiur-system \
--set webhook.tls.mode=manual \
--set webhook.tls.secretName=kopiur-webhook-tls \
--set webhook.caBundle="$(base64 -w0 ca.crt)"
Or disable the webhook entirely (validation then relies on the controller's defensive checks only — not recommended):
helm install kopiur oci://ghcr.io/home-operations/charts/kopiur -n kopiur-system --set webhook.enabled=false
Install scope¶
| Mode | --set installScope= |
RBAC | Manages | ClusterRepository |
|---|---|---|---|---|
| Cluster (default) | cluster |
ClusterRole | cluster-wide | reconciled |
| Namespaced | namespaced |
Role | release namespace only | not reconciled |
Cluster is the default so a shared platform repository (ClusterRepository) referenced by many tenant namespaces works out of the box — a namespace-scoped Role silently disables the cluster-scoped ClusterRepository kind. See deploy/examples/02-cluster-repository.yaml. Choose namespaced as the explicit least-privilege opt-down for a single-team install.
In namespaced scope the controller's watches are narrowed to the release
namespace to match the Role-only RBAC (the chart passes
--namespace={{ .Release.Namespace }}), and cluster-scoped kinds are skipped
entirely: ClusterRepository is not reconciled, and the
privileged-movers namespace opt-in check fails open
(the operator is already confined to admin-chosen namespaces there).
Features that need cluster RBAC are refused in namespaced scope
Two features read or write cluster-scoped objects (PersistentVolumes, StorageClasses, VolumeSnapshotClasses) that a Role can never grant, so a namespaced install refuses them up front with an actionable message instead of retrying forever:
target.populatorrestores — usetarget.pvc/target.pvcRef, or install withinstallScope: cluster.copyMethod: Snapshot/Clonebackups (CSI staging) — setcopyMethod: Directon theSnapshotPolicy. NotecopyMethoddefaults toSnapshot, so an untouched policy hits this refusal; the message spells out the fix.
RBAC for the optional web UI
If you use the web UI (spec.server on a Repository/ClusterRepository), the controller manages Deployments, Services, ConfigMaps, and Secrets for the kopia server pod. That RBAC ships with the chart in both scopes — nothing extra to grant. The feature is off until you add a spec.server block.
Runtime tuning¶
The chart's defaults are sized for a typical homelab-to-mid-size cluster; four top-level values tune the controller's runtime footprint and API-server load (full rationale in the developer docs):
| Value | Default | What it does |
|---|---|---|
workerThreads |
2 |
Tokio worker threads. The controller is I/O-bound; raise only for a reconcile-heavy deployment. |
streamingLists |
true |
Stream cluster-wide re-lists via the WatchList API (lower apiserver + controller memory). Auto-downgrades to paged lists on a pre-1.32 apiserver — or when the startup version probe fails. |
reconcileConcurrency |
8 |
Per-controller cap on concurrent reconciles. Bounds API-server load and file descriptors during re-list storms and API-server outages. 0 = unbounded (not recommended). |
maxConcurrentDeleteJobs |
0 (uncapped) |
Opt-in backstop on concurrent snapshot-delete batch Jobs; batching per repository is the primary protection. |
Why reconcileConcurrency is bounded by default
Unbounded reconcile concurrency let an API-server outage exhaust the controller's file descriptors within seconds (every re-listed object reconciling — and failing — at once). The default of 8 per controller clears a few-hundred-object re-list in seconds while keeping the controller a good citizen toward a struggling control plane.
CRD lifecycle¶
The 8 CRDs ship in the chart's special crds/ directory. Helm treats that directory specially: helm install installs the CRDs, but helm upgrade never touches them. There is no toggle for this.
Caution: because
helm upgradeskips thecrds/directory, a helm-CLI upgrade that carries a schema change does not apply it — you must apply the new CRDs yourself:# Server-side apply is required: the SnapshotPolicy CRD embeds a full JobSpec # (runJob hook) and is too large for the client-side last-applied annotation. kubectl apply --server-side -f deploy/crds/A GitOps flow (Flux/Argo with a
CreateReplacesync policy) applies CRD changes automatically. Anyone managing CRDs entirely out of band just appliesdeploy/crds/themselves —helm installskips the ones that already exist.
The CRDs and RBAC shipped by the chart are generated by cargo xtask gen-crds / cargo xtask gen-rbac and checked in under deploy/crds/ and deploy/rbac/. Those xtasks are the source of truth.
Upgrading from 0.5.x to 0.6.0 is a special case: the CRDs used to be release-owned templates and moved into this crds/ directory, so the crossing removes and re-installs them (cascade-deleting your CRs) unless you pin them first. See Upgrading → 0.5.x → 0.6.0 before you upgrade.
First backup¶
After install, create a repository and start backing up a PVC. The smallest end-to-end example is deploy/examples/01-single-pvc-scheduled.yaml:
kubectl apply -f deploy/examples/01-single-pvc-scheduled.yaml
kubectl get repositories,snapshotpolicies,snapshotschedules -n billing
Eight runnable walkthroughs live in deploy/examples/:
| File | Pattern |
|---|---|
01-single-pvc-scheduled.yaml |
Single PVC, scheduled daily |
02-cluster-repository.yaml |
Shared platform ClusterRepository (cluster scope) |
03-restore-by-backup.yaml |
Restore by picking a Snapshot |
04-multi-pvc-selector.yaml |
Multi-PVC label selector + group snapshot |
05-deploy-or-restore-gitops.yaml |
Deploy-or-restore (PVC dataSourceRef) |
06-manual-backup.yaml |
Manual one-shot Snapshot |
07-restore-discovered.yaml |
Restore a discovered / foreign snapshot |
08-maintenance.yaml |
kopia maintenance schedule + ownership lease |
Observability¶
- The controller serves
/metrics,/healthz, and/readyzon its single port (metrics.port, default:8081); the webhook serves/metricson its TLS port (webhook.port, default:8443). All metrics are under thekopiur_namespace. - The controller and webhook default to the dual-stack wildcard bind (
[::]:8081/[::]:8443), which serves both IPv4 and IPv6 kubelets/API servers, so probes work out of the box on either. The chart renders[::]:<port>intoKOPIUR_HTTP_ADDR(controller) andKOPIUR_WEBHOOK_ADDR(webhook). On a host where IPv6 is disabled in the pod network namespace a[::]bind fails outright — force an IPv4-only bind by addingKOPIUR_HTTP_ADDR=0.0.0.0:8081through the controller'sextraEnv. See Helm chart values → Controller Deployment for the full detail. - The controller's metrics
Serviceis always created (the:8081listener co-hosts/metricswith the probes, so there's nothing to disable). monitoring.serviceMonitor.enabled=truecreates a Prometheus-OperatorServiceMonitor(requires the Prometheus-Operator CRDs);monitoring.prometheusRule.enabled=trueships the kopiur alert rules.monitoring.dashboards.enabled=trueships the Grafana dashboard as a sidecar-discoverableConfigMap(source:deploy/dashboards/kopiur.json). Setmonitoring.dashboards.grafanaOperator.enabled=trueto instead render a grafana-operatorGrafanaDashboardCR from the same JSON (usemonitoring.dashboards.grafanaOperator.matchLabelsto select the Grafana instance).observability.otlp.enabled=true(withobservability.otlp.endpoint) additionally exports OTLP traces, logs, and a metrics push from the controller, webhook, and mover Jobs. Off by default.
Turn it all on with the ready-made overlay:
helm upgrade kopiur oci://ghcr.io/home-operations/charts/kopiur -n kopiur-system \
-f deploy/observability-values.yaml
See docs/dev/observability.md for the full metric list, OTLP details, and a sample collector config.
Upgrade / uninstall¶
helm upgrade kopiur oci://ghcr.io/home-operations/charts/kopiur -n kopiur-system # does NOT touch crds/ — apply schema changes yourself (see CRD lifecycle)
helm uninstall kopiur -n kopiur-system # leaves the crds/ CRDs (and your CRs) in place
Upgrading from 0.5.x needs a one-time pre-step so the CRD move doesn't delete your resources — see Upgrading.
See also¶
- Every chart value, walked through: Helm chart values
- Chart values & modes:
deploy/helm/kopiur/README.md