Scenario 09 — Share one repository across clusters¶
Several clusters back up to (or restore from) the SAME kopia repository at the same time — a platform team running one S3 bucket for every cluster's backups, an active-active pair, or a warm DR-standby cluster that also reads from the primary's history. Unlike migrating an app, where the app runs in exactly one place before and after, this is a genuinely concurrent shape: more than one cluster's operator writes to (or reads from) the same physical repository, on an ongoing basis.
Why this needs its own knob¶
Kopia's own identity model (username@hostname:path) defaults hostname to
the namespace. Point two clusters at the same bucket without anything
else and every same-named namespace (billing on cluster "east", billing
on cluster "west") writes under the identical kopia identity — three
things break at once:
| Collision surface | Without identityDefaults.cluster |
With it set |
|---|---|---|
| Identity dedup / retention | Both clusters' Snapshot CRs for billing share one kopia source; anything that matches a snapshot by path alone (retention, pin/unpin re-matching, a Snapshot CR's own deletion) can't tell which cluster's snapshot is which. |
Each cluster's default hostname becomes <namespace>.<cluster> — distinct identities, so retention, pin/unpin re-matching, and CR deletion never cross-match another cluster's snapshot of the same path. |
| Catalog cross-materialization | Every cluster's catalog scan sees the OTHER cluster's snapshots as unrecognized, same-namespace history and would materialize them as its own discovered Snapshot CRs — noise at best, a misleading duplicate at worst. |
catalog.foreignSnapshots classifies a .<cluster> suffix that isn't this cluster's own: Ignore (default) drops those rows — still counted in status.catalog.foreignSnapshotCount, never silently invisible — Fallback collects them into one namespace instead. |
| Maintenance lease | The default-managed Maintenance's lease derives from the repository name alone, byte-identical on every cluster — there is no way for the yield/claim logic to tell "mine" from "the other cluster's" claim. |
The lease is cluster-qualified (kopiur/<cluster>/...); each cluster gets a distinct lease and kopia owner, so exactly one cluster's Maintenance actually claims and runs. See Maintenance → Ownership and shared repositories. |
identityDefaults.cluster (an RFC 1123 label, at most 32 characters, no
dots) is the single field that fixes all three at once — see
Repositories → identityDefaults.cluster
for the field itself.
Greenfield setup — two supported shapes¶
Starting a shared repository from scratch, pick one of two shapes:
Shape A — active-active, both clusters write¶
Every cluster gets its own ClusterRepository (or Repository) object
pointing at the same bucket, each with a distinct identityDefaults.cluster,
and maintenance enabled on exactly one of them:
# Scenario 09 — Share one repository across clusters
#
# Two clusters ("east" and "west") back up to the SAME S3 bucket, each through
# its OWN ClusterRepository object with a distinct identityDefaults.cluster so
# same-named namespaces on either cluster never collide on kopia identity (the
# default hostname becomes <namespace>.<cluster> instead of bare <namespace>).
#
# IMPORTANT — this file shows BOTH clusters' manifests together for
# comparison, but they are NEVER applied to the same cluster:
# * The FIRST ClusterRepository (identityDefaults.cluster: east) is applied
# to cluster "east".
# * The SECOND (identityDefaults.cluster: west) is applied to cluster
# "west".
# They deliberately reuse the SAME ClusterRepository name ("shared-primary") —
# that's what a real rollout looks like: each cluster's operator only ever
# sees its OWN copy of the object, in its OWN etcd, so there is no naming
# collision to avoid.
#
# Maintenance ownership (see docs/maintenance.md "pick ONE owner" — kopia has
# no cross-cluster lock of its own beyond kopiur's lease): "east" is the ONE
# cluster that runs maintenance here (maintenance.enabled defaults to true).
# "west" explicitly disables its own managed Maintenance (enabled: false)
# rather than leaving it enabled to yield forever, which is safe but noisy.
#
# REQUIRES the operator (and mover image) upgraded to a version that supports
# identityDefaults.cluster on BOTH clusters BEFORE this is applied to either —
# see docs/scenarios/shared-repository-multi-cluster.md step 0. Also requires
# installScope=cluster on both (ClusterRepository is cluster-scoped).
#
# Field shapes verified against crates/api: identityDefaults.cluster is an
# RFC 1123 label (<=32 chars, no dots) on IdentityDefaults, shared by
# Repository and ClusterRepository.
---
# ============================================================================
# CLUSTER "east" — the maintenance owner
# ============================================================================
apiVersion: v1
kind: Secret
metadata:
name: kopia-shared-creds
namespace: kopiur-system
type: Opaque
stringData:
AWS_ACCESS_KEY_ID: "REPLACE_ME"
AWS_SECRET_ACCESS_KEY: "REPLACE_ME"
# MUST be the byte-identical value on every cluster sharing this repository —
# kopia bakes the resolved password into the repository format at creation.
KOPIA_PASSWORD: "REPLACE_ME_same_value_on_every_cluster"
---
apiVersion: kopiur.home-operations.com/v1alpha1
kind: ClusterRepository
metadata:
name: shared-primary
spec:
backend:
s3:
bucket: org-kopia-repo # the SAME bucket every cluster points at
prefix: ""
endpoint: s3.us-east-1.amazonaws.com
region: us-east-1
auth:
secretRef:
name: kopia-shared-creds
namespace: kopiur-system
encryption:
passwordSecretRef:
name: kopia-shared-creds
namespace: kopiur-system
key: KOPIA_PASSWORD
allowedNamespaces:
all: true
identityDefaults:
cluster: east # THIS cluster's identity suffix — distinct per cluster
catalog:
# Recommended on a shared repo: keep discovering peers' snapshots as they
# write, rather than only ever seeing them at the next spec change.
periodicRefresh: true
refreshInterval: 1h
maintenance:
namespace: kopiur-system
# enabled defaults to true — "east" is the ONE cluster that runs it.
---
# ============================================================================
# CLUSTER "west" — NOT the maintenance owner
# ============================================================================
apiVersion: v1
kind: Secret
metadata:
name: kopia-shared-creds
namespace: kopiur-system
type: Opaque
stringData:
AWS_ACCESS_KEY_ID: "REPLACE_ME"
AWS_SECRET_ACCESS_KEY: "REPLACE_ME"
KOPIA_PASSWORD: "REPLACE_ME_same_value_on_every_cluster"
---
apiVersion: kopiur.home-operations.com/v1alpha1
kind: ClusterRepository
metadata:
name: shared-primary # SAME name as "east" — this cluster's OWN copy of it
spec:
backend:
s3:
bucket: org-kopia-repo
prefix: ""
endpoint: s3.us-east-1.amazonaws.com
region: us-east-1
auth:
secretRef:
name: kopia-shared-creds
namespace: kopiur-system
encryption:
passwordSecretRef:
name: kopia-shared-creds
namespace: kopiur-system
key: KOPIA_PASSWORD
allowedNamespaces:
all: true
identityDefaults:
cluster: west # distinct from "east" above — never reuse a cluster suffix
catalog:
periodicRefresh: true
refreshInterval: 1h
maintenance:
# "west" is NOT the maintenance owner: disable it here rather than leaving
# the default takeoverPolicy: Never to yield on every cron slot forever
# (safe, but a noisy MaintenanceYielding warning + a wasted mover Job each
# time). See docs/maintenance.md "pick ONE owner".
enabled: false
Shape B — one primary writes, a secondary is read-only¶
If only ONE cluster should ever back up (a true DR-standby that only restores,
or a read replica), skip the second cluster's own maintenance/writer role
entirely: set mode: ReadOnly
on the secondary's repository object instead of also making it a writer.
spec:
mode: ReadOnly # this cluster only ever restores; never backs up, never bootstraps write
identityDefaults:
cluster: west # still set, so any catalog/placement work stays correct
A ReadOnly repository never launches a backup mover, never stamps or
restamps a maintenance owner (M6's ReadOnly owner-stamp gating — a read-only
consumer connecting read-write to self-heal an owner was exactly the bug this
closes), and skips maintenance projection. It can still discover and restore
the primary's snapshots.
Recommended on any shared repository
Set catalog.periodicRefresh: true (off by default) on every cluster sharing
the repository, so each one keeps discovering the others' snapshots as they
write, instead of only re-scanning on its own next spec change.
Turning it on for a repository already in production¶
This is the higher-stakes path: one cluster already has a live repository with real snapshot history, and you're adding one or more peers. The order below is load-bearing — each step exists to close a specific window the step before it would otherwise leave open.
Step 0 — upgrade every cluster's operator (and CRDs) FIRST¶
identityDefaults.cluster is a new field. Upgrade the operator and mover
image on every cluster that will share the repository before you set it
anywhere. A helm upgrade bumps the running images but does not touch
the CRDs — Helm's special crds/ chart directory is only ever applied by
helm install (see Installation → CRD lifecycle)
— so on a helm-CLI install you must also re-apply the CRDs by hand:
Verify the field actually resolves on every cluster before proceeding:
$ kubectl explain clusterrepository.spec.identityDefaults.cluster
# or: kubectl explain repository.spec.identityDefaults.cluster
If this errors or omits cluster, that cluster's CRD is still stale — fix it
before touching any repository's spec. A cluster with the old CRD would
silently strip identityDefaults.cluster at admission (structural-schema
pruning of an unrecognized field), and you'd have no error to tell you the
flip didn't take.
Step 1 — pick names and the ONE maintenance cluster, before any identity change¶
Decide, in order:
- Every cluster's
clustername (an RFC 1123 label, ≤ 32 chars, no dots). - Which ONE cluster owns maintenance going forward.
Then, on every non-owner cluster's repository, before touching
identityDefaults at all:
And remove ownership.takeoverPolicy: Force from every cluster except the
one true owner, if any of them had it set. Doing this first means that when
the identity flip (next step) changes the maintenance lease format, there is
only ever ONE cluster racing to claim it — never a multi-way scramble during
the transition itself. See Maintenance → pick ONE owner
for the full rationale.
Step 2 — annotate, then set cluster¶
On the repository that already has history (and on every peer being added), in the SAME apply:
metadata:
annotations:
kopiur.home-operations.com/allow-identity-change: "intentional"
spec:
identityDefaults:
cluster: east
Setting identityDefaults.cluster for the first time on a repository whose
consumers already have history IS an identity change — the webhook rejects it
fleet-wide without the annotation (see Repositories → identityDefaults.cluster
and Backups → identity).
foreignSnapshots is REQUIRED, explicitly, if you already have a fallbackNamespace
If catalog.fallbackNamespace was already configured (e.g. from adopting a
repository before it had a cluster identity), the validator now requires
catalog.foreignSnapshots to be set explicitly in the same edit — either
Ignore or Fallback. This is deliberate: adopting a cluster identity must
never silently repurpose or silently keep an existing fallback collector
without you saying which behavior you actually want. Choosing Ignore never
touches kopia data — any discovered Snapshot CR the fallback namespace was
previously holding for what now classifies as another cluster's snapshot is
simply expired (the CR is deleted; discovered rows are forced
deletionPolicy: Retain, so the underlying kopia snapshot is untouched
either way) on this same re-scan, since it no longer matches a placeable
classification.
This spec edit is itself a change, so it triggers an immediate re-scan (no
need to separately bump catalog.retain/wait for periodicRefresh) — the
same deterministic-rescan-on-spec-change behavior any other catalog config
edit gets.
Step 3 — what changes on the very next backup¶
- Re-pin. Every consumer
SnapshotPolicy'sstatus.resolved.identity.hostnameupdates to<namespace>.<cluster>on its next reconcile (identity re-resolves from the live repository — it was never frozen at admission). - New lineage. The next
Snapshotthis policy takes writes under that new hostname — a different kopia source than everything before it. - Content-dedup, not a re-upload. Kopia deduplicates by content hash across the WHOLE repository regardless of identity, so the first post-flip snapshot still dedups against everything already uploaded — this is a metadata re-pin, not a fresh full backup.
- Legacy
SnapshotCRs keep pruning normally. Kopiur's own GFS retention selects over everySnapshotCR labeled to a policy regardless of which identity each one recorded, so pre-flip CRs are not abandoned by their own policy's retention — they continue aging out exactly as before. See the residual window below for the multi-cluster-specific hazard this does not cover.
Step 4 — the residual window (and how to close it)¶
Before the flip, if more than one cluster was ALREADY (unsafely) writing
under the same bare-namespace identity, each cluster still holds its own
local Snapshot CRs pointing into that one shared, pre-migration lineage.
Every cluster's GFS retention prunes its own CRs independently — so until
every cluster's pre-migration CRs have aged out of their own GFS windows,
more than one cluster can still independently decide to delete a kopia
snapshot from that shared pre-migration history. This is exactly the
identity dedup / retention hazard the flip
fixes going forward — it just doesn't retroactively fix history that
predates it.
Mitigation: on every cluster except the one you want to keep pruning that
old shared lineage, flip the legacy (pre-flip) Snapshot CRs to
deletionPolicy: Orphan — the finalizer then removes only the Kubernetes
object when GFS ages them out, never touching the kopia snapshot:
$ kubectl get snapshot -n billing -l kopiur.home-operations.com/config=postgres-data -o json \
| jq -r '.items[] | select((.status.snapshot.identity.hostname // "") | contains(".") | not) | .metadata.name' \
| xargs -r -n1 kubectl patch snapshot -n billing --type=merge -p '{"spec":{"deletionPolicy":"Orphan"}}'
(The filter selects rows whose recorded identity hostname has no . — the
pre-cluster-identity, bare-namespace form — so only PRE-flip CRs are touched;
run it once, right after the flip, before any new-identity CRs exist.)
Step 5 — maintenance's lease upgrades itself¶
The owner cluster's managed Maintenance upgrades its lease to the
cluster-qualified format on its very next claim, automatically: the operator
records the pre-cluster lease as a recognized ownerAliases entry, so the
run recognizes kopia's already-recorded owner as itself, claims cleanly, and
re-stamps it to the new format. No manual takeoverPolicy: Force is needed
for this step specifically — see
Maintenance → self-healing a stale owner.
Step 6 — reaching the old history afterward¶
The pre-flip lineage doesn't disappear — restore from it at any time via the
raw identity, whether or not its Snapshot CR (or catalog row) still exists:
# Example 13 — Restore by raw kopia identity
#
# When there is no `Snapshot` CR to point at — a snapshot written by a foreign kopia
# client, or one that aged out of the catalog — restore by the raw kopia identity
# (`username@hostname:path`). This mode REQUIRES an explicit `spec.repository`
# (there's no Backup/SnapshotPolicy to infer it from).
#
# # find identities/snapshots in a repo with the kopia CLI (or `kubectl get snapshots`
# # for catalogued ones):
# kopia snapshot list --all
#
# Default identity is `<snapshotPolicyName>@<namespace>:/pvc/<pvcName>` — see
# docs/concepts/how-kopia-works.md.
#
# Field shapes verified against crates/api: source.identity is
# { username, hostname, sourcePath?, snapshotID?, asOf?, offset? }; `snapshotID`
# has a capital ID on the wire.
---
apiVersion: kopiur.home-operations.com/v1alpha1
kind: Restore
metadata:
name: postgres-restore-foreign
namespace: billing
spec:
# Identity restores can't infer the repo — name it explicitly.
repository:
kind: Repository
name: nas-primary
namespace: backups
source:
identity:
username: postgres-data # kopia username to match
hostname: billing # kopia hostname to match
sourcePath: /data # optional; omit to match any path for this identity
snapshotID: k1f1ec0a8 # pin an exact snapshot (or use asOf / offset below)
# asOf: 2026-05-01T00:00:00Z # newest snapshot at/before this instant
# offset: 0 # 0 = latest, 1 = previous, ...
target:
pvc:
name: postgres-data-restored
storageClassName: fast-ssd
capacity: 100Gi
accessModes:
- ReadWriteOnce
policy:
onMissingSnapshot: Fail # explicit sources fail-closed
waitTimeout: 5m
Edge notes¶
- The fork guard's window. Like the per-policy guard, the repository-edit
guard can only fire on an UPDATE (it diffs against
oldObject); a delete + re-apply of the repository is a CREATE with no prior object to diff, so it bypasses the guard by design. This is a known, accepted limitation of admission-time guards in general — don't rely on delete + re-apply as a way around the annotation. - Explicit
spec.identitypolicies keep the ORIGINAL hazard. ASnapshotPolicythat pins its ownidentity.{username,hostname}explicitly never consults the repository'sidentityDefaultsat all — an explicit override always wins. If two clusters both hand-pin the SAME identity, the cluster flip does nothing for them; give each cluster's pinned identity its own suffix by hand instead. - A custom, dotted
hostnameExprforgoes catalog placement — restores still work.classify_hostnamesplits any hostname at its FIRST.and compares the suffix tocluster— it has no idea whether the prefix is a real namespace. A hand-writtenhostnameExprthat embeds a dot for its own reasons (not the<namespace>.<cluster>convention) will misclassify under this scheme, and the catalog's placement pass will likely give up on it (fallback or unplaced) — butRestore.spec.source.identitymatches the rawusername@hostname:pathdirectly and is unaffected. Prefer the default<namespace>.<cluster>scheme (nohostnameExprat all) unless you have a specific reason not to. - Verification goes quiet for one cycle, not forever.
verification.deeprestores the latest snapshot for the policy's currently resolved identity — if a scheduled deep verify lands in the gap between the flip and the first post-flip backup, it targets the new identity and finds nothing yet to restore. Expect (at most) one harmless miss until the first new-lineage snapshot completes;verification.quickand the overall verification gate are unaffected (they don't reset on an identity change). - Stale legacy kopia policy objects are cosmetic. Kopiur pins kopia's own
native retention to effectively-infinite per identity it manages
(
kopia policy set, scoped to that identity) — after a flip, the OLD identity's policy object just becomes unused clutter inkopia policy list. It's harmless to leave;kopia policy delete <old-username>@<old-hostname>(viakubectl execinto any mover/kopia-shell pod connected to the repository) tidies it up if you care.
Verification checklist¶
# 1. The consumer's resolved identity carries the new suffix:
$ kubectl get snapshotpolicy postgres-data -n billing \
-o jsonpath='{.status.resolved.identity.hostname}'
billing.east
# 2. Exactly one cluster's Maintenance shows itself as the owner; the OWNER
# column is the cluster-qualified lease (kopiur/<cluster>/...), and the
# others show LeaseOwned=False, reason=LeaseHeldByOther (or don't exist at
# all, if maintenance.enabled: false there):
$ kubectl get maintenance -A
# 3. Peers' snapshots are being counted (not silently dropped or duplicated):
$ kubectl get clusterrepository shared-primary \
-o jsonpath='{.status.catalog.foreignSnapshotCount}'
# 4. The CRD schema is current on every cluster (step 0, re-checked):
$ kubectl explain clusterrepository.spec.identityDefaults.cluster
See also¶
- Repositories →
identityDefaults.cluster— the field itself, and the identity-change warning. - Backups → identity — how identity resolves live, and both fork guards.
- Maintenance → Ownership and shared repositories — the lease format table,
ownerAliases, and self-healing rules in full. - Migrate an app across clusters or namespaces — the one-time-move cousin of this scenario.
- Troubleshooting →
LeaseHeldByOtheron a shared repository — telling expected from stale. deploy/examples/13-restore-by-identity.yaml— restoring old-lineage history by raw identity.