Skip to content

Repository files navigation

Kafka Backup Operator for Strimzi

License Rust Kubernetes

Disclaimer: This project is not part of the Strimzi project or the CNCF. It is an independent, community-built operator designed to work with Strimzi-managed Kafka clusters.

A Kubernetes operator for Kafka backup and disaster recovery of Strimzi-managed Apache Kafka clusters. Provides dedicated CRDs for automated Kafka backup scheduling, point-in-time recovery, and multi-cloud storage — designed for the Strimzi ecosystem.

Current release: 0.3.1 — default job image osodevops/kafka-backup:v0.22.0.

Why Kafka Backup?

Strimzi makes running Apache Kafka on Kubernetes straightforward, but backup and disaster recovery remain unsolved problems in the Strimzi ecosystem:

  • MirrorMaker2 is not a backup — it requires a full secondary cluster, expensive cross-cluster replication, and complex client failover procedures.
  • PVC snapshots are fragile — deleting Strimzi CRDs triggers garbage collection of PVCs, and node failures can result in permanent data loss.
  • No point-in-time recovery — there is no native mechanism to restore a Strimzi Kafka cluster to a specific moment in time.
  • No Strimzi-compatible backup CRD — backup workflows are entirely manual or require external tools that don't integrate with the Strimzi operator model.

The Kafka Backup Operator solves these problems with a purpose-built Kubernetes operator that follows Strimzi conventions and is designed to work with Strimzi-managed clusters, providing first-class backup and restore capabilities.

Features

  • Strimzi-compatible CRDs — KafkaBackup and KafkaRestore custom resources under the kafkabackup.com API group, following Strimzi conventions for status conditions, labels, and finalizers
  • Auto-discovery of Strimzi resources — automatically resolves bootstrap servers, TLS certificates, and KafkaUser credentials from your existing Strimzi Kafka CRs
  • Scheduled backups — cron-based scheduling with timezone support via Kubernetes CronJobs
  • Point-in-time recovery (PITR) — restore your Kafka cluster to any millisecond-precision timestamp
  • Multi-cloud storage — back up to Amazon S3, Azure Blob Storage, Google Cloud Storage, or any S3-compatible store (MinIO, Ceph RGW)
  • Topic filtering — include/exclude topics using glob or regex patterns, for both backup and restore
  • Topic mapping — rename topics during restore for migration or testing scenarios
  • Consumer group offset restore — restore consumer group offsets with optional group remapping
  • Retention policies — automatic pruning of old backups by count or age
  • Compression — gzip, snappy, lz4, or zstd compression for storage efficiency
  • Prometheus metrics — built-in observability with backup/restore counters, duration histograms, and storage gauges
  • Pod template customisation — full control over backup/restore Job pods (affinity, tolerations, host aliases, service account, security context, environment variables)
  • Azure Workload Identity — native support for passwordless Azure authentication

Architecture

The Kafka Backup Operator creates Kubernetes Jobs (or CronJobs for scheduled backups) that run the kafka-backup CLI to perform backup and restore operations. This provides resource isolation, failure isolation, and pod-level customisation.

+------------------------------------------------------------+
|                    Kubernetes Cluster                       |
|                                                            |
|  +------------------+         +------------------------+   |
|  |  Strimzi         |         |  Kafka Backup          |   |
|  |  Cluster         | <------ |  Operator              |   |
|  |  Operator        |  reads  |  (watches KafkaBackup  |   |
|  |                  |  Kafka  |   & KafkaRestore CRs)  |   |
|  +--------+---------+   CRs  +-----------+------------+   |
|           |                               |                |
|           v                               v creates        |
|  +------------------+         +------------------------+   |
|  |  Kafka CR        |         |  Backup/Restore Jobs   |   |
|  |  + Brokers       | <------ |  (kafka-backup CLI)    |   |
|  |  + ZooKeeper     |  reads  |                        |   |
|  +------------------+  data   +-----------+------------+   |
|                                           |                |
|                                           v writes/reads   |
|                               +------------------------+   |
|                               |  Object Storage        |   |
|                               |  (S3 / Azure / GCS)    |   |
|                               +------------------------+   |
+------------------------------------------------------------+

Quick Start

Prerequisites

  • Kubernetes 1.27+
  • Strimzi Cluster Operator installed with a running Kafka cluster (kafka.strimzi.io/v1 and legacy v1beta2 resources are supported)
  • Helm 3.x

Install with Helm

# Add the OSO DevOps Helm repository
helm repo add oso-devops https://osodevops.github.io/helm-charts/
helm repo update

# Install the operator
helm install strimzi-backup-operator oso-devops/strimzi-backup-operator \
  --namespace kafka \
  --create-namespace

Create a Backup

apiVersion: kafkabackup.com/v1
kind: KafkaBackup
metadata:
  name: my-cluster-backup
  namespace: kafka
spec:
  strimziClusterRef:
    name: my-cluster
  storage:
    type: s3
    s3:
      bucket: my-kafka-backups
      region: eu-west-1
      accessKeySecret:
        name: aws-credentials
        key: access-key-id
      secretKeySecret:
        name: aws-credentials
        key: secret-access-key
  topics:
    include:
      - "orders-*"
      - "payments-*"
    exclude:
      - "__*"
  backup:
    compression: zstd
    parallelism: 4
    includeOffsetHeaders: true   # default; adds x-original-offset / x-original-timestamp to every archived record
  logging:
    level: warn
    format: json
    modules:
      kafka_backup: warn
      rdkafka: info
  env:
    - name: RUST_LOG
      value: "kafka_backup=warn,rdkafka=info"
  schedule:
    cron: "0 2 * * *"    # Daily at 2 AM
    timezone: "Europe/London"
  retention:
    maxBackups: 30
    maxAge: "90d"
    pruneOnSchedule: true

Restore from a Backup

apiVersion: kafkabackup.com/v1
kind: KafkaRestore
metadata:
  name: my-cluster-restore
  namespace: kafka
spec:
  strimziClusterRef:
    name: my-cluster
  backupRef:
    name: my-cluster-backup
  topics:
    include:
      - "orders-*"          # glob, or "~orders-\d+" for regex
    exclude:
      - "*-internal"
  pointInTime:
    timestamp: "2026-02-12T14:30:00.000Z"
  logging:
    level: info
    format: json
  topicMapping:
    - sourceTopic: orders
      targetTopic: orders-restored
  consumerGroups:
    restore: true
  restore:
    stripOffsetHeaders: false   # true = restore header-for-header identical to the source (kafka-backup >= v0.19.0)

Offset headers and record fidelity

kafka-backup restores keys, values, timestamps and headers verbatim (null and empty values stay distinct — job images >= v0.18.0). The one deliberate addition is a pair of offset-tracking headers, controlled by three fields:

Field kafka-backup key Default Effect
spec.backup.includeOffsetHeaders backup.include_offset_headers true Adds x-original-offset / x-original-timestamp (little-endian i64) — and x-source-cluster when sourceClusterId is set — to every archived record. Needed for header-based consumer offset recovery.
spec.restore.includeOriginalOffsetHeader restore.include_original_offset_header false Adds x-original-offset, x-original-timestamp and x-source-partition to every restored record (also implied by the header-based strategy).
spec.restore.stripOffsetHeaders restore.strip_offset_headers false Removes all of the above from archived records before producing, so a restore of an archive taken with the default includeOffsetHeaders: true is header-for-header identical to the source. Requires a job image >= v0.19.0. Offset mapping is unaffected — the source offset is stored natively in the segment.

Any other native key can be passed through spec.backup.config / spec.restore.config. Known limitation: duplicate header keys on one record are collapsed to the last one (kafka-backup#156).

Custom Resource Definitions

CRD Short Name API Group Description
KafkaBackup kb kafkabackup.com/v1 (v1alpha1 served, deprecated) Defines a backup configuration with scheduling, retention, and storage
KafkaRestore kr kafkabackup.com/v1 (v1alpha1 served, deprecated) Defines a restore operation with PITR, topic mapping, and consumer group restore

Engine image (spec.image)

Both resources accept spec.image, the kafka-backup image the Job runs:

spec:
  image: osodevops/kafka-backup:v0.19.2   # example: any 0.x release newer than the default

Leave it unset to use the operator-wide default (Helm backupJobs.image, else the release's compiled-in engine). See Compatibility for which engines are supported and how the default is chosen; the image each Job actually ran with is recorded in status.lastBackup.image / status.restore.image.

Advanced options passthrough

The typed fields under spec.backup / spec.restore cover the common kafka-backup options with camelCase names (for example segmentSize maps to segment_max_bytes). Every other option in the kafka-backup config reference can be set through the free-form config map using kafka-backup's native snake_case key names — following the same pattern as Strimzi's spec.kafka.config:

spec:
  backup:
    compression: zstd
    config:
      fetch_max_bytes: 16777216       # native kafka-backup key names
      segment_max_records: 2000000

Keys set in config are passed through verbatim to the generated job config and take precedence over the typed fields (so config.segment_max_bytes wins over segmentSize). Keys the kafka-backup binary does not recognize are logged as warnings at job startup (kafka-backup >= v0.16.0) instead of being silently ignored. spec.restore.config works the same way for the restore: section.

Incremental backups (spec.offsetStorage)

Adding spec.offsetStorage to a KafkaBackup turns each scheduled run into an incremental one: the operator keeps backup_id equal to the resource name so every run resumes from the offsets the previous run saved, and the engine merges manifests across runs.

spec:
  offsetStorage:
    backend: sqlite            # the only implemented backend
    # dbPath: /data/offsets.db # optional; see below before setting it
    # syncIntervalSecs: 30     # how often the offset DB is synced to storage (engine >= 0.22.0)

How the offset database moves between runs:

  • The remote copy is authoritative. The engine uploads the SQLite offset database to <storage prefix>/<backup_id>/offsets.db alongside the manifest, and on start-up downloads it whenever the local database is empty.
  • Job pods are ephemeral. The operator mounts no volume for the database, so it lives in the container's writable layer and disappears with the pod. That is by design: every run starts from the remote copy, and losing a pod — or the node — is harmless.
  • dbPath on a persistent volume changes which copy wins. If you mount a PVC and point dbPath at it, the local database is no longer empty on the next run, so it is used even when the remote copy is newer (for example after a manual kafka-backup prune or a run from another pod). Leave dbPath unset unless you understand that trade-off.
  • syncIntervalSecs is honoured by kafka-backup 0.22.0 and later; older engines use backup.sync_interval_secs (30s default).
  • s3Key is deprecated and ignored by kafka-backup 0.22.0 and later: the remote key is always <prefix>/<backup_id>/offsets.db, which kafka-backup prune, status and this operator's retention all rely on. The operator logs a warning when it is set.

See the kafka-backup incremental backups guide for the engine-side details.

Pausing reconciliation

KafkaBackup and KafkaRestore support Strimzi's standard pause annotation. While its value is "true", the operator reports a ReconciliationPaused condition but does not add a finalizer, resolve dependencies, or create/update ConfigMaps, Jobs, or CronJobs. Removing the annotation or setting it to "false" resumes normal reconciliation.

kubectl annotate kafkabackup restore-source strimzi.io/pause-reconciliation="true"
kubectl annotate kafkabackup restore-source strimzi.io/pause-reconciliation-

Storage Configuration

Amazon S3

storage:
  type: s3
  s3:
    bucket: my-kafka-backups
    region: eu-west-1
    prefix: production/
    accessKeySecret:
      name: aws-credentials
      key: access-key-id
    secretKeySecret:
      name: aws-credentials
      key: secret-access-key

Azure Blob Storage

storage:
  type: azure
  azure:
    storageAccount: myaccount
    container: kafka-backups
    prefix: production/
    accountKeySecret:
      name: azure-credentials
      key: account-key

On AKS, prefer Workload Identity over an account key — no secret in the cluster, and the Job pod authenticates with a federated token: spec.storage.azure.useWorkloadIdentity: true on the resource, the Job's ServiceAccount annotated with azure.workload.identity/client-id, and the Job pod labelled azure.workload.identity/use: "true". A complete example, including the Azure-side federated credential, is in config/examples/kafka-backup-azure-workload-identity.yaml.

Google Cloud Storage

storage:
  type: gcs
  gcs:
    bucket: my-kafka-backups
    prefix: production/
    credentialsSecret:
      name: gcs-credentials
      key: service-account.json

S3-Compatible (MinIO)

storage:
  type: s3
  s3:
    bucket: kafka-backups
    endpoint: https://minio.example.com
    forcePathStyle: true
    allowHttp: false
    accessKeySecret:
      name: minio-credentials
      key: access-key-id
    secretKeySecret:
      name: minio-credentials
      key: secret-access-key

Authentication

The operator automatically discovers TLS certificates and authentication credentials from your Strimzi cluster. You can also reference KafkaUser CRs directly:

Automatic KafkaUser Resolution

spec:
  strimziClusterRef:
    name: my-cluster
  authentication:
    type: tls
    kafkaUserRef:
      name: backup-user    # References a Strimzi KafkaUser CR

Manual TLS Certificates

spec:
  authentication:
    type: tls
    certificateAndKey:
      secretName: my-tls-secret
      certificate: user.crt
      key: user.key

SCRAM-SHA-512

spec:
  authentication:
    type: scram-sha-512
    username: backup-user
    passwordSecret:
      name: backup-user-password
      key: password

Listener selection

Backup and restore jobs connect through the Kafka listener whose authentication.type matches the resource's spec.authentication — SCRAM credentials go to a scram-sha-512 listener, client certificates to a tls listener, and resources without authentication use an unauthenticated listener. Among matching listeners, in-cluster types (internal, cluster-ip) are preferred over external ones, and TLS-encrypted listeners over plaintext. If no listener matches, reconciliation fails with a condition listing the cluster's listeners.

To bypass the automatic selection, name a listener explicitly:

spec:
  strimziClusterRef:
    name: my-cluster
    listener: external    # connect via this listener, as declared in the Kafka CR

Helm Values

Parameter Description Default
image.repository Operator container image ghcr.io/osodevops/strimzi-backup-operator
image.tag Image tag Chart appVersion
image.pullPolicy Image pull policy Always
replicaCount Number of operator replicas (extra replicas are warm standbys; needs leaderElection.enabled) 1
updateStrategy Deployment update strategy; the default (RollingUpdate, maxSurge: 0, maxUnavailable: 1) deletes the outgoing pod before creating its replacement see values.yaml
watchNamespaces Namespaces to watch (empty = all) []
logging.level Rust log filter info,kafka_backup_operator=debug
logging.format Log output format json
serviceAccount.create Create a service account true
backupJobs.serviceAccountName Service account used by backup/restore job pods (empty = operator service account) ""
backupJobs.image Default kafka-backup image for job pods (empty = the release's compiled-in engine; spec.image overrides per resource) — see Compatibility ""
backupJobs.imagePullPolicy imagePullPolicy for job pods (empty = Kubernetes default) ""
azureWorkloadIdentity.enabled Enable Azure Workload Identity false
azureWorkloadIdentity.clientId Azure Managed Identity client ID ""
metrics.enabled Enable Prometheus metrics true
metrics.serviceMonitor.enabled Create a Prometheus ServiceMonitor false
metrics.jobPodMonitor.enabled Create a PodMonitor for backup/restore job metrics false
leaderElection.enabled Only the replica holding the <release>-leader Lease reconciles true
leaderElection.leaseDuration Time a standby waits without seeing a lease change before taking over 15s
leaderElection.renewDeadline Time the leader keeps retrying renewals before it exits 10s
leaderElection.retryPeriod Interval between acquire/renew attempts 2s
resources.requests.cpu CPU request 100m
resources.requests.memory Memory request 128Mi
resources.limits.cpu CPU limit 500m
resources.limits.memory Memory limit 512Mi

Job service accounts across namespaces

Backup and restore Jobs run in the namespace of the KafkaBackup/KafkaRestore resource, and by default reference the service account named by backupJobs.serviceAccountName (falling back to the operator's own service account). Service accounts are namespace-scoped, so for resources created outside the operator's namespace, set spec.template.pod.serviceAccountName to a service account that exists in that namespace:

spec:
  template:
    pod:
      serviceAccountName: kafka-backup-jobs

The service account only needs to exist (job pods don't call the Kubernetes API), but it should carry any workload-identity annotations (IRSA, Azure Workload Identity) your storage backend requires.

Compatibility

Each operator release ships with a default kafka-backup engine image. It is compiled in (DEFAULT_BACKUP_IMAGE in src/engine.rs), named in the header of this README, logged at start-up (default_job_image), exposed as the strimzi_backup_operator_engine_image_info metric, and is what every Job runs unless told otherwise. The operator generates the engine's YAML config and runs the engine in a Job; backup and restore behaviour lives in the engine.

Choosing the engine. Per resource: spec.image on KafkaBackup / KafkaRestore. For the whole installation: Helm backupJobs.image (env BACKUP_JOB_IMAGE). Precedence: spec.image → backupJobs.image → compiled-in default. The image a Job actually used is recorded in status.lastBackup.image / status.restore.image.

Policy.

  • Default engine — tested in CI on every commit (the configs the operator generates are run through it) and the only combination we guarantee.
  • Newer engine — any kafka-backup 0.x release newer than the default is supported: pin it with spec.image or backupJobs.image to pick up engine fixes without waiting for an operator release. Since kafka-backup 0.16.0 an unknown config key is warned about instead of failing the run, and spec.backup.config / spec.restore.config are passed through verbatim, so new engine options are usable immediately. A nightly CI leg runs the generated configs against the latest engine release.
  • Older engine — supported down to the minimum in the table, with degraded behaviour: options the older engine does not know are ignored with a warning (see the feature table). Below the minimum the operator still runs the Job but sets EngineVersionSupported=False on the resource (reason EngineOlderThanMinimum) and counts it in strimzi_backup_operator_engine_version_unsupported_total. Images whose tag is not a release (latest, a digest, a custom tag) get EngineVersionSupported=True with reason EngineVersionUnknown.
  • Behaviour changes — an engine bump can change data semantics (0.18.0 stopped flattening null header values to empty; 0.19.0 added strip_offset_headers). Every default-image bump has a CHANGELOG entry saying what changed and whether existing archives need re-taking. Pin spec.image to keep an older engine.
  • Engine 1.x will require an operator release; 0.x does not support it.
Operator Default engine Minimum engine Notes
0.3.0 – 0.3.1 v0.22.0 v0.16.0 kafkabackup.com/v1 API, templated CRDs; engine 0.22: prune/backup.retention, on_missing_topic, syncIntervalSecs honoured
0.2.25 v0.19.1 v0.16.0 backupJobs.image, EngineVersionSupported condition, status.*.image
0.2.22 – 0.2.24 v0.19.1 v0.16.0
0.2.21 v0.19.0 v0.16.0 stripOffsetHeaders needs ≥ v0.19.0
0.2.20 v0.16.0 v0.16.0 spec.backup.config passthrough needs ≥ v0.16.0
≤ 0.2.19 v0.15.x — not supported
Feature Needs engine
spec.backup.config / spec.restore.config passthrough ≥ v0.16.0
null header values preserved (not flattened to empty) ≥ v0.18.0
spec.restore.stripOffsetHeaders ≥ v0.19.0
per-run incremental progress gauges (kafka_backup_snapshot_records_*) ≥ v0.19.1

Bumping the default is a scripted, gated step — see RELEASING.md.

High availability and upgrades

The operator is a single writer: exactly one replica may reconcile at a time, otherwise two versions can race on the same child resources (issue #62 — a scheduled backup's CronJob kept the previous version's job image after helm upgrade). Two mechanisms enforce that:

  • updateStrategy (default RollingUpdate with maxSurge: 0, maxUnavailable: 1) — the outgoing pod is deleted before its replacement is created, so at most one operator pod is scheduled at any time; the outgoing pod may still be draining (finishing in-flight reconciles, up to terminationGracePeriodSeconds) while the new one starts, which is what the lease below covers. type: Recreate additionally waits for the old pod to be fully gone, but an existing release managed with server-side apply (Helm 4) cannot switch to it in place — Kubernetes forbids the API-defaulted rollingUpdate block together with Recreate and SSA cannot clear an unowned default. Use it on fresh installs, or remove the block first: kubectl patch deploy <release> --type=json -p '[{"op":"remove","path":"/spec/strategy/rollingUpdate"}]'.
  • Leader election (leaderElection.enabled, default true) — replicas compete for the Lease <release fullname>-leader in the operator namespace (kubectl get lease -n <ns>). Only the holder runs the controllers; the others stand by. On shutdown the leader drains its reconciles first and then releases the lease, so the successor takes over within about one retryPeriod; after a crash the standby waits leaseDuration (15s). A leader that cannot renew within renewDeadline exits and is restarted by the kubelet as a candidate. Lease expiry is judged on the observing pod's own clock, so clock skew between nodes does not cause premature takeovers.

For a warm standby run replicaCount: 2 (optionally with maxSurge: 1 so a rollout keeps two pods up): the incoming pod becomes Ready as a standby (/readyz returns standby) and acquires the lease as soon as the outgoing leader releases it. /readyz returns 503 leader election pending until a replica has observed the lease at least once, so an install whose ServiceAccount lacks the coordination.k8s.io/leases rule (rendered by the chart when leader election is enabled) fails helm upgrade --wait instead of running silently. Set leaderElection.enabled=false to opt out; the maxSurge: 0 rollout and the CronJob watch below still cover plain upgrades. The gauge strimzi_backup_operator_leader{identity} is 1 on the leader.

On start-up (and on every out-of-band change to an owned CronJob) the backup controller re-applies the desired CronJob, and it reconciles every resource again 5s and 60s after start, so a stale CronJob is corrected within seconds even if the mechanisms above are disabled.

CRD upgrades

Since 0.3.0 the chart renders the CRDs as templates (crds.install, default true), so helm upgrade applies CRD changes — including the addition of the v1 API version — and crds.keep (default true) marks them helm.sh/resource-policy: keep so helm uninstall never deletes your KafkaBackup/KafkaRestore resources. Installations of 0.2.x used Helm's static crds/ directory, which installs CRDs without release ownership metadata, so a plain helm upgrade to 0.3.0 stops with CustomResourceDefinition "kafkabackups.kafkabackup.com" … cannot be imported into the current release: invalid ownership metadata. Let Helm adopt them — once, on the upgrade from 0.2.x:

# Helm 3.17+ / Helm 4
helm upgrade <release> oso/strimzi-backup-operator -n <namespace> --take-ownership

# Older Helm: label the two CRDs for adoption first, then upgrade as usual
for crd in kafkabackups.kafkabackup.com kafkarestores.kafkabackup.com; do
  kubectl label crd "$crd" app.kubernetes.io/managed-by=Helm --overwrite
  kubectl annotate crd "$crd" meta.helm.sh/release-name=<release> meta.helm.sh/release-namespace=<namespace> --overwrite
done

Prefer not to let Helm manage the CRDs at all? Install with --set crds.install=false and apply them yourself:

kubectl apply --server-side -f https://github.com/osodevops/strimzi-backup-operator/releases/download/v0.3.0/crds.yaml

Existing v1alpha1 objects keep working; they are read back as v1 (identical schema) and every v1alpha1 request prints a deprecation warning.

Logging

The operator deployment and the backup/restore job pods are configured separately.

Configure the operator deployment log level with Helm:

logging:
  level: "info,kafka_backup_operator=debug"
  format: json

Configure kafka-backup job logging on each KafkaBackup or KafkaRestore:

apiVersion: kafkabackup.com/v1
kind: KafkaBackup
metadata:
  name: debug-backup
  namespace: kafka
spec:
  strimziClusterRef:
    name: my-kafka-cluster
  storage:
    type: s3
    s3:
      bucket: my-kafka-backups
      region: eu-west-1
  logging:
    level: warn
    format: json
    output: stderr
    modules:
      kafka_backup: warn
      rdkafka: info

For environment-based logging, use top-level spec.env; these entries are added to backup and restore job containers:

spec:
  env:
    - name: RUST_LOG
      value: "kafka_backup=warn,rdkafka=info"

Monitoring

There are two independent metrics endpoints:

  • The operator exposes controller health metrics on port 9090 at /metrics.
  • Each kafka-backup backup/restore pod exposes operation and progress metrics on port 8080 by default when spec.metrics.enabled is not false.

The operator does not proxy or copy job metrics. A ServiceMonitor only scrapes the operator Service; enable the chart's PodMonitor to discover job pods directly across namespaces.

Operator metrics include:

Metric Type Description
strimzi_backup_operator_build_info Gauge Running operator version
strimzi_backup_operator_reconciliations_total Counter Reconciliations by controller and result
strimzi_backup_operator_reconciliation_duration_seconds Histogram Reconciliation latency by controller and result

Job metrics include kafka_backup_lag_records, the low-cardinality kafka_backup_lag_records_sum, snapshot progress gauges kafka_backup_snapshot_records_target and kafka_backup_snapshot_records_remaining, kafka_backup_records_total, kafka_backup_bytes_total, and the runtime's storage, throughput, compression, error, and restore metric families.

Prometheus ServiceMonitor

# values.yaml
metrics:
  enabled: true
  serviceMonitor:
    enabled: true
    interval: 30s
    scrapeTimeout: 10s
  jobPodMonitor:
    enabled: true
    interval: 30s
    scrapeTimeout: 10s

The PodMonitor requires the Prometheus Operator CRDs and defaults to namespaceSelector.any: true because backup and restore resources may live outside the Helm release namespace. Its endpoint path defaults to /metrics; if a CR customizes spec.metrics.path, provide a matching custom PodMonitor.

For a one-shot backup or restore, keep the metrics endpoint alive long enough for at least one scrape. A practical minimum is twice the PodMonitor interval:

spec:
  metrics:
    enabled: true
    keepAliveSeconds: 60
    maxPartitionLabels: 100

This setting is supported by the default job image (see Compatibility). maxPartitionLabels limits unique topic/partition series; set it to 0 only when unlimited per-partition cardinality is intentional. Durable last-success reporting should still come from the CR status or a service-level batch metric store rather than an operator proxy.

Disaster Recovery Workflow

1. SCHEDULE         2. BACKUP          3. DISASTER        4. RESTORE

+-----------+      +-----------+      +-----------+      +-----------+
|  CronJob  |----->|   Kafka   |----->|   Data    |      |   PITR    |
|  triggers |reads |   Backup  |stores|   Safe    |apply |  Restore  |
|  backup   |data  |   Job     |to S3 |   in S3   |----->|   Job     |
+-----------+      +-----------+      +-----------+      +-----------+
  1. Schedule — The operator creates a CronJob based on your KafkaBackup schedule
  2. Backup — The Job reads data from Kafka and writes it to object storage with compression
  3. Disaster — Your data is safe in durable, versioned object storage
  4. Restore — Apply a KafkaRestore CR with an optional point-in-time timestamp to recover

Development

Prerequisites

  • Rust 1.88+ (stable)
  • A Kubernetes cluster with Strimzi installed (for integration testing)

Build from Source

# Build the operator
cargo build --release

# Run tests
cargo test --all-features

# Run clippy
cargo clippy --all-features -- -D warnings

# Generate CRDs
cargo run --release --bin crdgen

# Build Docker image
docker build -t strimzi-backup-operator .

Local Development

# Run the operator locally against your kubeconfig
RUST_LOG=debug cargo run

# Install CRDs
kubectl apply -f deploy/crds/

API stability, support and maintenance

  • API versions — kafkabackup.com/v1 is the stable API (operator 0.3.0+); v1alpha1 is still served with an identical schema but deprecated. What v1 guarantees, the deprecation window for v1alpha1, and how to migrate (change apiVersion, nothing else) are in docs/api-stability.md.
  • Support — what is covered, severity levels and response targets, the supported-version window and the security-fix policy are in SUPPORT.md. Support is included with the kafka-backup Enterprise licence.
  • Security — how to report a vulnerability and how advisories are handled: SECURITY.md.
  • Maintenance and continuity — the release process, who can cut a release, and the source-availability commitment are in SUPPORT.md#maintenance-and-continuity.

Contributing

Contributions are welcome. Please open an issue or submit a pull request.

  1. Fork the repository
  2. Create a feature branch (git checkout -b feature/my-feature)
  3. Commit your changes
  4. Push to the branch (git push origin feature/my-feature)
  5. Open a pull request

Releases, including how the default kafka-backup engine image is bumped, are described in RELEASING.md.

License

Apache License 2.0 — see LICENSE for details.

Related Projects

  • Strimzi — Kafka on Kubernetes
  • kafka-backup — The backup engine used by this operator
  • OSO DevOps — Enterprise Kafka and Kubernetes consulting

About

Strimzi-native Kubernetes operator for Kafka backup and disaster recovery. Scheduled backups, point-in-time recovery, multi-cloud storage (S3/Azure/GCS).

Topics

Resources

Security policy

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages