Disclaimer: This project is not part of the Strimzi project or the CNCF. It is an independent, community-built operator designed to work with Strimzi-managed Kafka clusters.
A Kubernetes operator for Kafka backup and disaster recovery of Strimzi-managed Apache Kafka clusters. Provides dedicated CRDs for automated Kafka backup scheduling, point-in-time recovery, and multi-cloud storage — designed for the Strimzi ecosystem.
Current release: 0.3.1 — default job image osodevops/kafka-backup:v0.22.0.
Strimzi makes running Apache Kafka on Kubernetes straightforward, but backup and disaster recovery remain unsolved problems in the Strimzi ecosystem:
- MirrorMaker2 is not a backup — it requires a full secondary cluster, expensive cross-cluster replication, and complex client failover procedures.
- PVC snapshots are fragile — deleting Strimzi CRDs triggers garbage collection of PVCs, and node failures can result in permanent data loss.
- No point-in-time recovery — there is no native mechanism to restore a Strimzi Kafka cluster to a specific moment in time.
- No Strimzi-compatible backup CRD — backup workflows are entirely manual or require external tools that don't integrate with the Strimzi operator model.
The Kafka Backup Operator solves these problems with a purpose-built Kubernetes operator that follows Strimzi conventions and is designed to work with Strimzi-managed clusters, providing first-class backup and restore capabilities.
- Strimzi-compatible CRDs —
KafkaBackupandKafkaRestorecustom resources under thekafkabackup.comAPI group, following Strimzi conventions for status conditions, labels, and finalizers - Auto-discovery of Strimzi resources — automatically resolves bootstrap servers, TLS certificates, and KafkaUser credentials from your existing Strimzi
KafkaCRs - Scheduled backups — cron-based scheduling with timezone support via Kubernetes CronJobs
- Point-in-time recovery (PITR) — restore your Kafka cluster to any millisecond-precision timestamp
- Multi-cloud storage — back up to Amazon S3, Azure Blob Storage, Google Cloud Storage, or any S3-compatible store (MinIO, Ceph RGW)
- Topic filtering — include/exclude topics using glob or regex patterns, for both backup and restore
- Topic mapping — rename topics during restore for migration or testing scenarios
- Consumer group offset restore — restore consumer group offsets with optional group remapping
- Retention policies — automatic pruning of old backups by count or age
- Compression — gzip, snappy, lz4, or zstd compression for storage efficiency
- Prometheus metrics — built-in observability with backup/restore counters, duration histograms, and storage gauges
- Pod template customisation — full control over backup/restore Job pods (affinity, tolerations, host aliases, service account, security context, environment variables)
- Azure Workload Identity — native support for passwordless Azure authentication
The Kafka Backup Operator creates Kubernetes Jobs (or CronJobs for scheduled backups) that run the kafka-backup CLI to perform backup and restore operations. This provides resource isolation, failure isolation, and pod-level customisation.
+------------------------------------------------------------+
| Kubernetes Cluster |
| |
| +------------------+ +------------------------+ |
| | Strimzi | | Kafka Backup | |
| | Cluster | <------ | Operator | |
| | Operator | reads | (watches KafkaBackup | |
| | | Kafka | & KafkaRestore CRs) | |
| +--------+---------+ CRs +-----------+------------+ |
| | | |
| v v creates |
| +------------------+ +------------------------+ |
| | Kafka CR | | Backup/Restore Jobs | |
| | + Brokers | <------ | (kafka-backup CLI) | |
| | + ZooKeeper | reads | | |
| +------------------+ data +-----------+------------+ |
| | |
| v writes/reads |
| +------------------------+ |
| | Object Storage | |
| | (S3 / Azure / GCS) | |
| +------------------------+ |
+------------------------------------------------------------+
- Kubernetes 1.27+
- Strimzi Cluster Operator installed with a running Kafka cluster (
kafka.strimzi.io/v1and legacyv1beta2resources are supported) - Helm 3.x
# Add the OSO DevOps Helm repository
helm repo add oso-devops https://osodevops.github.io/helm-charts/
helm repo update
# Install the operator
helm install strimzi-backup-operator oso-devops/strimzi-backup-operator \
--namespace kafka \
--create-namespaceapiVersion: kafkabackup.com/v1
kind: KafkaBackup
metadata:
name: my-cluster-backup
namespace: kafka
spec:
strimziClusterRef:
name: my-cluster
storage:
type: s3
s3:
bucket: my-kafka-backups
region: eu-west-1
accessKeySecret:
name: aws-credentials
key: access-key-id
secretKeySecret:
name: aws-credentials
key: secret-access-key
topics:
include:
- "orders-*"
- "payments-*"
exclude:
- "__*"
backup:
compression: zstd
parallelism: 4
includeOffsetHeaders: true # default; adds x-original-offset / x-original-timestamp to every archived record
logging:
level: warn
format: json
modules:
kafka_backup: warn
rdkafka: info
env:
- name: RUST_LOG
value: "kafka_backup=warn,rdkafka=info"
schedule:
cron: "0 2 * * *" # Daily at 2 AM
timezone: "Europe/London"
retention:
maxBackups: 30
maxAge: "90d"
pruneOnSchedule: trueapiVersion: kafkabackup.com/v1
kind: KafkaRestore
metadata:
name: my-cluster-restore
namespace: kafka
spec:
strimziClusterRef:
name: my-cluster
backupRef:
name: my-cluster-backup
topics:
include:
- "orders-*" # glob, or "~orders-\d+" for regex
exclude:
- "*-internal"
pointInTime:
timestamp: "2026-02-12T14:30:00.000Z"
logging:
level: info
format: json
topicMapping:
- sourceTopic: orders
targetTopic: orders-restored
consumerGroups:
restore: true
restore:
stripOffsetHeaders: false # true = restore header-for-header identical to the source (kafka-backup >= v0.19.0)kafka-backup restores keys, values, timestamps and headers verbatim (null and empty
values stay distinct — job images >= v0.18.0). The one deliberate addition is a pair
of offset-tracking headers, controlled by three fields:
| Field | kafka-backup key | Default | Effect |
|---|---|---|---|
spec.backup.includeOffsetHeaders |
backup.include_offset_headers |
true |
Adds x-original-offset / x-original-timestamp (little-endian i64) — and x-source-cluster when sourceClusterId is set — to every archived record. Needed for header-based consumer offset recovery. |
spec.restore.includeOriginalOffsetHeader |
restore.include_original_offset_header |
false |
Adds x-original-offset, x-original-timestamp and x-source-partition to every restored record (also implied by the header-based strategy). |
spec.restore.stripOffsetHeaders |
restore.strip_offset_headers |
false |
Removes all of the above from archived records before producing, so a restore of an archive taken with the default includeOffsetHeaders: true is header-for-header identical to the source. Requires a job image >= v0.19.0. Offset mapping is unaffected — the source offset is stored natively in the segment. |
Any other native key can be passed through spec.backup.config / spec.restore.config.
Known limitation: duplicate header keys on one record are collapsed to the last one
(kafka-backup#156).
| CRD | Short Name | API Group | Description |
|---|---|---|---|
KafkaBackup |
kb |
kafkabackup.com/v1 (v1alpha1 served, deprecated) |
Defines a backup configuration with scheduling, retention, and storage |
KafkaRestore |
kr |
kafkabackup.com/v1 (v1alpha1 served, deprecated) |
Defines a restore operation with PITR, topic mapping, and consumer group restore |
Both resources accept spec.image, the kafka-backup image the Job runs:
spec:
image: osodevops/kafka-backup:v0.19.2 # example: any 0.x release newer than the defaultLeave it unset to use the operator-wide default (Helm backupJobs.image, else
the release's compiled-in engine). See Compatibility for
which engines are supported and how the default is chosen; the image each Job
actually ran with is recorded in status.lastBackup.image /
status.restore.image.
The typed fields under spec.backup / spec.restore cover the common
kafka-backup options with camelCase names (for example segmentSize maps to
segment_max_bytes). Every other option in the
kafka-backup config reference
can be set through the free-form config map using kafka-backup's native
snake_case key names — following the same pattern as Strimzi's
spec.kafka.config:
spec:
backup:
compression: zstd
config:
fetch_max_bytes: 16777216 # native kafka-backup key names
segment_max_records: 2000000Keys set in config are passed through verbatim to the generated job config
and take precedence over the typed fields (so config.segment_max_bytes wins
over segmentSize). Keys the kafka-backup binary does not recognize are
logged as warnings at job startup (kafka-backup >= v0.16.0) instead of being
silently ignored. spec.restore.config works the same way for the restore:
section.
Adding spec.offsetStorage to a KafkaBackup turns each scheduled run into an
incremental one: the operator keeps backup_id equal to the resource name
so every run resumes from the offsets the previous run saved, and the engine
merges manifests across runs.
spec:
offsetStorage:
backend: sqlite # the only implemented backend
# dbPath: /data/offsets.db # optional; see below before setting it
# syncIntervalSecs: 30 # how often the offset DB is synced to storage (engine >= 0.22.0)How the offset database moves between runs:
- The remote copy is authoritative. The engine uploads the SQLite offset
database to
<storage prefix>/<backup_id>/offsets.dbalongside the manifest, and on start-up downloads it whenever the local database is empty. - Job pods are ephemeral. The operator mounts no volume for the database, so it lives in the container's writable layer and disappears with the pod. That is by design: every run starts from the remote copy, and losing a pod — or the node — is harmless.
dbPathon a persistent volume changes which copy wins. If you mount a PVC and pointdbPathat it, the local database is no longer empty on the next run, so it is used even when the remote copy is newer (for example after a manualkafka-backup pruneor a run from another pod). LeavedbPathunset unless you understand that trade-off.syncIntervalSecsis honoured by kafka-backup 0.22.0 and later; older engines usebackup.sync_interval_secs(30s default).s3Keyis deprecated and ignored by kafka-backup 0.22.0 and later: the remote key is always<prefix>/<backup_id>/offsets.db, whichkafka-backup prune,statusand this operator's retention all rely on. The operator logs a warning when it is set.
See the kafka-backup incremental backups guide for the engine-side details.
KafkaBackup and KafkaRestore support Strimzi's standard pause annotation.
While its value is "true", the operator reports a ReconciliationPaused
condition but does not add a finalizer, resolve dependencies, or create/update
ConfigMaps, Jobs, or CronJobs. Removing the annotation or setting it to
"false" resumes normal reconciliation.
kubectl annotate kafkabackup restore-source strimzi.io/pause-reconciliation="true"
kubectl annotate kafkabackup restore-source strimzi.io/pause-reconciliation-storage:
type: s3
s3:
bucket: my-kafka-backups
region: eu-west-1
prefix: production/
accessKeySecret:
name: aws-credentials
key: access-key-id
secretKeySecret:
name: aws-credentials
key: secret-access-keystorage:
type: azure
azure:
storageAccount: myaccount
container: kafka-backups
prefix: production/
accountKeySecret:
name: azure-credentials
key: account-keyOn AKS, prefer Workload Identity over an account key — no secret in the
cluster, and the Job pod authenticates with a federated token:
spec.storage.azure.useWorkloadIdentity: true on the resource, the Job's
ServiceAccount annotated with azure.workload.identity/client-id, and the
Job pod labelled azure.workload.identity/use: "true". A complete example,
including the Azure-side federated credential, is in
config/examples/kafka-backup-azure-workload-identity.yaml.
storage:
type: gcs
gcs:
bucket: my-kafka-backups
prefix: production/
credentialsSecret:
name: gcs-credentials
key: service-account.jsonstorage:
type: s3
s3:
bucket: kafka-backups
endpoint: https://minio.example.com
forcePathStyle: true
allowHttp: false
accessKeySecret:
name: minio-credentials
key: access-key-id
secretKeySecret:
name: minio-credentials
key: secret-access-keyThe operator automatically discovers TLS certificates and authentication credentials from your Strimzi cluster. You can also reference KafkaUser CRs directly:
spec:
strimziClusterRef:
name: my-cluster
authentication:
type: tls
kafkaUserRef:
name: backup-user # References a Strimzi KafkaUser CRspec:
authentication:
type: tls
certificateAndKey:
secretName: my-tls-secret
certificate: user.crt
key: user.keyspec:
authentication:
type: scram-sha-512
username: backup-user
passwordSecret:
name: backup-user-password
key: passwordBackup and restore jobs connect through the Kafka listener whose
authentication.type matches the resource's spec.authentication — SCRAM
credentials go to a scram-sha-512 listener, client certificates to a tls
listener, and resources without authentication use an unauthenticated
listener. Among matching listeners, in-cluster types (internal,
cluster-ip) are preferred over external ones, and TLS-encrypted listeners
over plaintext. If no listener matches, reconciliation fails with a condition
listing the cluster's listeners.
To bypass the automatic selection, name a listener explicitly:
spec:
strimziClusterRef:
name: my-cluster
listener: external # connect via this listener, as declared in the Kafka CR| Parameter | Description | Default |
|---|---|---|
image.repository |
Operator container image | ghcr.io/osodevops/strimzi-backup-operator |
image.tag |
Image tag | Chart appVersion |
image.pullPolicy |
Image pull policy | Always |
replicaCount |
Number of operator replicas (extra replicas are warm standbys; needs leaderElection.enabled) |
1 |
updateStrategy |
Deployment update strategy; the default (RollingUpdate, maxSurge: 0, maxUnavailable: 1) deletes the outgoing pod before creating its replacement |
see values.yaml |
watchNamespaces |
Namespaces to watch (empty = all) | [] |
logging.level |
Rust log filter | info,kafka_backup_operator=debug |
logging.format |
Log output format | json |
serviceAccount.create |
Create a service account | true |
backupJobs.serviceAccountName |
Service account used by backup/restore job pods (empty = operator service account) | "" |
backupJobs.image |
Default kafka-backup image for job pods (empty = the release's compiled-in engine; spec.image overrides per resource) — see Compatibility |
"" |
backupJobs.imagePullPolicy |
imagePullPolicy for job pods (empty = Kubernetes default) |
"" |
azureWorkloadIdentity.enabled |
Enable Azure Workload Identity | false |
azureWorkloadIdentity.clientId |
Azure Managed Identity client ID | "" |
metrics.enabled |
Enable Prometheus metrics | true |
metrics.serviceMonitor.enabled |
Create a Prometheus ServiceMonitor | false |
metrics.jobPodMonitor.enabled |
Create a PodMonitor for backup/restore job metrics | false |
leaderElection.enabled |
Only the replica holding the <release>-leader Lease reconciles |
true |
leaderElection.leaseDuration |
Time a standby waits without seeing a lease change before taking over | 15s |
leaderElection.renewDeadline |
Time the leader keeps retrying renewals before it exits | 10s |
leaderElection.retryPeriod |
Interval between acquire/renew attempts | 2s |
resources.requests.cpu |
CPU request | 100m |
resources.requests.memory |
Memory request | 128Mi |
resources.limits.cpu |
CPU limit | 500m |
resources.limits.memory |
Memory limit | 512Mi |
Backup and restore Jobs run in the namespace of the KafkaBackup/KafkaRestore
resource, and by default reference the service account named by
backupJobs.serviceAccountName (falling back to the operator's own service
account). Service accounts are namespace-scoped, so for resources created
outside the operator's namespace, set spec.template.pod.serviceAccountName
to a service account that exists in that namespace:
spec:
template:
pod:
serviceAccountName: kafka-backup-jobsThe service account only needs to exist (job pods don't call the Kubernetes API), but it should carry any workload-identity annotations (IRSA, Azure Workload Identity) your storage backend requires.
Each operator release ships with a default kafka-backup engine image. It is
compiled in (DEFAULT_BACKUP_IMAGE in src/engine.rs), named in the header of
this README, logged at start-up (default_job_image), exposed as the
strimzi_backup_operator_engine_image_info metric, and is what every Job runs
unless told otherwise. The operator generates the engine's YAML config and
runs the engine in a Job; backup and restore behaviour lives in the engine.
Choosing the engine. Per resource: spec.image on KafkaBackup /
KafkaRestore. For the whole installation: Helm backupJobs.image (env
BACKUP_JOB_IMAGE). Precedence: spec.image → backupJobs.image →
compiled-in default. The image a Job actually used is recorded in
status.lastBackup.image / status.restore.image.
Policy.
- Default engine — tested in CI on every commit (the configs the operator generates are run through it) and the only combination we guarantee.
- Newer engine — any kafka-backup
0.xrelease newer than the default is supported: pin it withspec.imageorbackupJobs.imageto pick up engine fixes without waiting for an operator release. Since kafka-backup 0.16.0 an unknown config key is warned about instead of failing the run, andspec.backup.config/spec.restore.configare passed through verbatim, so new engine options are usable immediately. A nightly CI leg runs the generated configs against the latest engine release. - Older engine — supported down to the minimum in the table, with degraded
behaviour: options the older engine does not know are ignored with a
warning (see the feature table). Below the minimum the operator still runs
the Job but sets
EngineVersionSupported=Falseon the resource (reasonEngineOlderThanMinimum) and counts it instrimzi_backup_operator_engine_version_unsupported_total. Images whose tag is not a release (latest, a digest, a custom tag) getEngineVersionSupported=Truewith reasonEngineVersionUnknown. - Behaviour changes — an engine bump can change data semantics (0.18.0
stopped flattening null header values to empty; 0.19.0 added
strip_offset_headers). Every default-image bump has a CHANGELOG entry saying what changed and whether existing archives need re-taking. Pinspec.imageto keep an older engine. - Engine 1.x will require an operator release; 0.x does not support it.
| Operator | Default engine | Minimum engine | Notes |
|---|---|---|---|
| 0.3.0 – 0.3.1 | v0.22.0 | v0.16.0 | kafkabackup.com/v1 API, templated CRDs; engine 0.22: prune/backup.retention, on_missing_topic, syncIntervalSecs honoured |
| 0.2.25 | v0.19.1 | v0.16.0 | backupJobs.image, EngineVersionSupported condition, status.*.image |
| 0.2.22 – 0.2.24 | v0.19.1 | v0.16.0 | |
| 0.2.21 | v0.19.0 | v0.16.0 | stripOffsetHeaders needs ≥ v0.19.0 |
| 0.2.20 | v0.16.0 | v0.16.0 | spec.backup.config passthrough needs ≥ v0.16.0 |
| ≤ 0.2.19 | v0.15.x | — | not supported |
| Feature | Needs engine |
|---|---|
spec.backup.config / spec.restore.config passthrough |
≥ v0.16.0 |
| null header values preserved (not flattened to empty) | ≥ v0.18.0 |
spec.restore.stripOffsetHeaders |
≥ v0.19.0 |
per-run incremental progress gauges (kafka_backup_snapshot_records_*) |
≥ v0.19.1 |
Bumping the default is a scripted, gated step — see RELEASING.md.
The operator is a single writer: exactly one replica may reconcile at a time,
otherwise two versions can race on the same child resources (issue #62 — a
scheduled backup's CronJob kept the previous version's job image after
helm upgrade). Two mechanisms enforce that:
updateStrategy(defaultRollingUpdatewithmaxSurge: 0,maxUnavailable: 1) — the outgoing pod is deleted before its replacement is created, so at most one operator pod is scheduled at any time; the outgoing pod may still be draining (finishing in-flight reconciles, up toterminationGracePeriodSeconds) while the new one starts, which is what the lease below covers.type: Recreateadditionally waits for the old pod to be fully gone, but an existing release managed with server-side apply (Helm 4) cannot switch to it in place — Kubernetes forbids the API-defaultedrollingUpdateblock together withRecreateand SSA cannot clear an unowned default. Use it on fresh installs, or remove the block first:kubectl patch deploy <release> --type=json -p '[{"op":"remove","path":"/spec/strategy/rollingUpdate"}]'.- Leader election (
leaderElection.enabled, defaulttrue) — replicas compete for the Lease<release fullname>-leaderin the operator namespace (kubectl get lease -n <ns>). Only the holder runs the controllers; the others stand by. On shutdown the leader drains its reconciles first and then releases the lease, so the successor takes over within about oneretryPeriod; after a crash the standby waitsleaseDuration(15s). A leader that cannot renew withinrenewDeadlineexits and is restarted by the kubelet as a candidate. Lease expiry is judged on the observing pod's own clock, so clock skew between nodes does not cause premature takeovers.
For a warm standby run replicaCount: 2 (optionally with maxSurge: 1 so
a rollout keeps two pods up): the incoming pod becomes Ready as a standby
(/readyz returns standby) and acquires the lease as soon as the outgoing
leader releases it. /readyz returns 503 leader election pending until a replica
has observed the lease at least once, so an install whose ServiceAccount lacks
the coordination.k8s.io/leases rule (rendered by the chart when leader
election is enabled) fails helm upgrade --wait instead of running silently.
Set leaderElection.enabled=false to opt out; the maxSurge: 0 rollout and
the CronJob watch below still cover plain upgrades. The gauge strimzi_backup_operator_leader{identity}
is 1 on the leader.
On start-up (and on every out-of-band change to an owned CronJob) the backup controller re-applies the desired CronJob, and it reconciles every resource again 5s and 60s after start, so a stale CronJob is corrected within seconds even if the mechanisms above are disabled.
Since 0.3.0 the chart renders the CRDs as templates (crds.install, default
true), so helm upgrade applies CRD changes — including the addition of the
v1 API version — and crds.keep (default true) marks them
helm.sh/resource-policy: keep so helm uninstall never deletes your
KafkaBackup/KafkaRestore resources. Installations of 0.2.x used Helm's
static crds/ directory, which installs CRDs without release ownership
metadata, so a plain helm upgrade to 0.3.0 stops with
CustomResourceDefinition "kafkabackups.kafkabackup.com" … cannot be imported into the current release: invalid ownership metadata. Let Helm adopt them —
once, on the upgrade from 0.2.x:
# Helm 3.17+ / Helm 4
helm upgrade <release> oso/strimzi-backup-operator -n <namespace> --take-ownership
# Older Helm: label the two CRDs for adoption first, then upgrade as usual
for crd in kafkabackups.kafkabackup.com kafkarestores.kafkabackup.com; do
kubectl label crd "$crd" app.kubernetes.io/managed-by=Helm --overwrite
kubectl annotate crd "$crd" meta.helm.sh/release-name=<release> meta.helm.sh/release-namespace=<namespace> --overwrite
donePrefer not to let Helm manage the CRDs at all? Install with
--set crds.install=false and apply them yourself:
kubectl apply --server-side -f https://github.com/osodevops/strimzi-backup-operator/releases/download/v0.3.0/crds.yamlExisting v1alpha1 objects keep working; they are read back as v1 (identical
schema) and every v1alpha1 request prints a deprecation warning.
The operator deployment and the backup/restore job pods are configured separately.
Configure the operator deployment log level with Helm:
logging:
level: "info,kafka_backup_operator=debug"
format: jsonConfigure kafka-backup job logging on each KafkaBackup or KafkaRestore:
apiVersion: kafkabackup.com/v1
kind: KafkaBackup
metadata:
name: debug-backup
namespace: kafka
spec:
strimziClusterRef:
name: my-kafka-cluster
storage:
type: s3
s3:
bucket: my-kafka-backups
region: eu-west-1
logging:
level: warn
format: json
output: stderr
modules:
kafka_backup: warn
rdkafka: infoFor environment-based logging, use top-level spec.env; these entries are added
to backup and restore job containers:
spec:
env:
- name: RUST_LOG
value: "kafka_backup=warn,rdkafka=info"There are two independent metrics endpoints:
- The operator exposes controller health metrics on port
9090at/metrics. - Each
kafka-backupbackup/restore pod exposes operation and progress metrics on port8080by default whenspec.metrics.enabledis notfalse.
The operator does not proxy or copy job metrics. A ServiceMonitor only scrapes
the operator Service; enable the chart's PodMonitor to discover job pods
directly across namespaces.
Operator metrics include:
| Metric | Type | Description |
|---|---|---|
strimzi_backup_operator_build_info |
Gauge | Running operator version |
strimzi_backup_operator_reconciliations_total |
Counter | Reconciliations by controller and result |
strimzi_backup_operator_reconciliation_duration_seconds |
Histogram | Reconciliation latency by controller and result |
Job metrics include kafka_backup_lag_records, the low-cardinality
kafka_backup_lag_records_sum, snapshot progress gauges
kafka_backup_snapshot_records_target and
kafka_backup_snapshot_records_remaining, kafka_backup_records_total,
kafka_backup_bytes_total, and the runtime's storage, throughput, compression,
error, and restore metric families.
# values.yaml
metrics:
enabled: true
serviceMonitor:
enabled: true
interval: 30s
scrapeTimeout: 10s
jobPodMonitor:
enabled: true
interval: 30s
scrapeTimeout: 10sThe PodMonitor requires the Prometheus Operator CRDs and defaults to
namespaceSelector.any: true because backup and restore resources may live
outside the Helm release namespace. Its endpoint path defaults to /metrics;
if a CR customizes spec.metrics.path, provide a matching custom PodMonitor.
For a one-shot backup or restore, keep the metrics endpoint alive long enough for at least one scrape. A practical minimum is twice the PodMonitor interval:
spec:
metrics:
enabled: true
keepAliveSeconds: 60
maxPartitionLabels: 100This setting is supported by the default job image (see
Compatibility).
maxPartitionLabels limits unique topic/partition series; set it to 0 only
when unlimited per-partition cardinality is intentional.
Durable last-success reporting should still come from the CR status or a
service-level batch metric store rather than an operator proxy.
1. SCHEDULE 2. BACKUP 3. DISASTER 4. RESTORE
+-----------+ +-----------+ +-----------+ +-----------+
| CronJob |----->| Kafka |----->| Data | | PITR |
| triggers |reads | Backup |stores| Safe |apply | Restore |
| backup |data | Job |to S3 | in S3 |----->| Job |
+-----------+ +-----------+ +-----------+ +-----------+
- Schedule — The operator creates a CronJob based on your
KafkaBackupschedule - Backup — The Job reads data from Kafka and writes it to object storage with compression
- Disaster — Your data is safe in durable, versioned object storage
- Restore — Apply a
KafkaRestoreCR with an optional point-in-time timestamp to recover
- Rust 1.88+ (stable)
- A Kubernetes cluster with Strimzi installed (for integration testing)
# Build the operator
cargo build --release
# Run tests
cargo test --all-features
# Run clippy
cargo clippy --all-features -- -D warnings
# Generate CRDs
cargo run --release --bin crdgen
# Build Docker image
docker build -t strimzi-backup-operator .# Run the operator locally against your kubeconfig
RUST_LOG=debug cargo run
# Install CRDs
kubectl apply -f deploy/crds/- API versions —
kafkabackup.com/v1is the stable API (operator 0.3.0+);v1alpha1is still served with an identical schema but deprecated. Whatv1guarantees, the deprecation window forv1alpha1, and how to migrate (changeapiVersion, nothing else) are in docs/api-stability.md. - Support — what is covered, severity levels and response targets, the supported-version window and the security-fix policy are in SUPPORT.md. Support is included with the kafka-backup Enterprise licence.
- Security — how to report a vulnerability and how advisories are handled: SECURITY.md.
- Maintenance and continuity — the release process, who can cut a release, and the source-availability commitment are in SUPPORT.md#maintenance-and-continuity.
Contributions are welcome. Please open an issue or submit a pull request.
- Fork the repository
- Create a feature branch (
git checkout -b feature/my-feature) - Commit your changes
- Push to the branch (
git push origin feature/my-feature) - Open a pull request
Releases, including how the default kafka-backup engine image is bumped, are
described in RELEASING.md.
Apache License 2.0 — see LICENSE for details.
- Strimzi — Kafka on Kubernetes
- kafka-backup — The backup engine used by this operator
- OSO DevOps — Enterprise Kafka and Kubernetes consulting