Generate a Debug Bundle with rpk in Kubernetes

Use rpk to generate a debug bundle from the Redpanda brokers running in Kubernetes, or from the Redpanda Operators of a Stretch Cluster.

Which bundle to generate depends on how Redpanda is deployed:

Deployment Generate Command

Single Kubernetes cluster (Redpanda Operator or Helm chart)

Redpanda debug bundle

rpk debug bundle inside a broker Pod, or rpk debug remote-bundle from your machine. See Single Kubernetes cluster.

Stretch Cluster across several Kubernetes clusters

Redpanda debug bundle, plus the multicluster operator debug bundle when the problem involves the operators

rpk debug bundle in a broker Pod in each Kubernetes cluster, or rpk debug remote-bundle with brokers from every cluster. rpk k8s multicluster bundle for the operators, run from your machine with a kubeconfig for any one cluster. See Stretch Cluster.

Collect the multicluster operator debug bundle when a Stretch Cluster shows operator symptoms such as a RedpandaBrokerPool that isn’t reconciled, operators that don’t form a Raft group, or TLS errors between operators. For broker symptoms such as under-replicated partitions or slow produce and consume, the Redpanda debug bundle is enough in both deployment types.

Single Kubernetes cluster

In a single-cluster deployment, the only debug bundle is the one collected from the Redpanda brokers.

To generate a debug bundle with rpk, you have two options depending on your needs:

  • Use rpk debug bundle: Run this command directly on each broker in the cluster. This method requires access to the nodes your brokers are running on.

  • Use rpk debug remote-bundle: Run this command from a remote machine to collect diagnostics data from all brokers in the cluster. This method is ideal when you want to gather data without logging into each node individually.

Prerequisites

You must have rpk installed on your host machine.

Use rpk debug bundle

To generate a debug bundle with rpk, you can run the rpk debug bundle command on each broker in the cluster.

  1. Create a ClusterRole to allow Redpanda to collect information from the Kubernetes API:

    • Operator

    • Helm

    redpanda-cluster.yaml
    apiVersion: cluster.redpanda.com/v1alpha2
    kind: Redpanda
    metadata:
      name: redpanda
    spec:
      chartRef: {}
      clusterSpec:
        serviceAccount:
          create: true
        rbac:
          enabled: true
    kubectl apply -f redpanda-cluster.yaml --namespace <namespace>
    • --values

    • --set

    serviceaccount.yaml
    serviceAccount:
      create: true
    rbac:
      enabled: true
    helm upgrade --install redpanda redpanda/redpanda --namespace redpanda --create-namespace \
      --values serviceaccount.yaml --reuse-values
    helm upgrade --install redpanda redpanda/redpanda --namespace <namespace> --create-namespace \
      --set serviceAccount.create=true \
      --set rbac.enabled=true

    If you aren’t using the Helm chart, you can create the ClusterRole manually:

    kubectl create clusterrolebinding redpanda --clusterrole=view --serviceaccount=redpanda:default
  2. Execute the rpk debug bundle command on a broker:

    kubectl exec -it --namespace <namespace> redpanda-0 -c redpanda -- rpk debug bundle --namespace <namespace>

    If you have an upload URL from the Redpanda support team, provide it in the --upload-url flag to upload your debug bundle to Redpanda.

    kubectl exec -it --namespace <namespace> redpanda-0 -c redpanda -- rpk debug bundle \
      --upload-url <url> \
      --namespace <namespace>

    Example output:

    Creating bundle file...
    
    Debug bundle saved to "/var/lib/redpanda/1675440652-bundle.zip"
  3. On your host machine, make a directory in which to save the debug bundle:

    mkdir debug-bundle
  4. Copy the debug bundle ZIP file to the debug-bundle directory on your host machine.

    Replace <bundle-name> with the name of your ZIP file.

    kubectl cp <namespace>/redpanda-0:/var/lib/redpanda/<bundle-name> debug-bundle/<bundle-name>.zip
  5. Unzip the file on your host machine and use it to debug your cluster.

    cd debug-bundle
    unzip <bundle-name>.zip

    For guidance on reading the debug bundle, see Inspect a Debug Bundle.

  6. Remove the debug bundle from the Redpanda broker:

    kubectl exec redpanda-0 -c redpanda --namespace <namespace> -- rm /var/lib/redpanda/<bundle-name>.zip

When you’ve finished troubleshooting, remove the debug bundle from your host machine:

rm -r debug-bundle

Use rpk debug remote-bundle

The rpk debug remote-bundle command allows you to remotely generate a consolidated debug bundle from all brokers in your cluster. This command is useful when you do not want to log into each broker’s node individually using rpk debug bundle.

  1. Create an rpk profile to connect to your cluster. Include all brokers you want to collect data from in the admin.hosts configuration.

    rpk profile create <profile-name> --set admin.hosts=<brokers>
  2. Check the configured addresses in your profile:

    rpk profile print

    Example output:

    profile.yml
    admin:
      hosts:
        - broker1:9644
        - broker2:9644
        - broker3:9644
  3. Start the debug bundle process:

    rpk debug remote-bundle start --namespace <namespace>

    Replace <namespace> with the Kubernetes namespace in which your Redpanda cluster is running.

    To skip the confirmation steps, use the --no-confirm flag.

    To generate a bundle for a subset of brokers, use the -X admin.hosts flag. For example, rpk debug remote-bundle start -X admin.hosts target-broker:9644.

  4. Check the status:

    rpk debug remote-bundle status

    Example output:

    BROKER           STATUS   JOB-ID
    localhost:29644  running  7f93fd6e-fc5e-46a5-8842-717542f89e59
    localhost:19644  running  7f93fd6e-fc5e-46a5-8842-717542f89e59

    When the status changes to success, the process is complete.

    To cancel a debug bundle process while it’s running, use rpk debug remote-bundle cancel.
  5. When the process is complete, download the debug bundle:

    rpk debug remote-bundle download

    By default the compressed file is downloaded to your current working directory. To choose a different location or filename, use the --output flag. For example:

    rpk debug remote-bundle download --output ~/redpanda/debug-bundles/cluster1

    This command results in cluster1.zip downloaded to the ~/redpanda/debug-bundles/ directory.

Unzip the file and use the contents to debug your cluster. For guidance on reading the debug bundle, see Inspect a Debug Bundle.

Configure the debug_bundle_auto_removal_seconds property to automatically remove debug bundles after a period of time. See Automatically remove debug bundles.

Stretch Cluster

A Stretch Cluster has two sources of diagnostics: the Redpanda brokers, and the Redpanda Operator that runs in each Kubernetes cluster. For an operator problem, Redpanda Support usually needs a bundle from both.

Redpanda debug bundle

The brokers in a Stretch Cluster form one Redpanda cluster, so the steps are the same as for a single Kubernetes cluster, but you must collect from brokers in every Kubernetes cluster:

  • For rpk debug bundle, run the command in a broker Pod in each Kubernetes cluster, for example by passing --context to kubectl.

  • For rpk debug remote-bundle, list brokers from every Kubernetes cluster in admin.hosts. The command collects only from the brokers that you list.

Multicluster operator debug bundle

In a Stretch Cluster, one Redpanda Operator runs in each Kubernetes cluster and the operators coordinate through Raft. The rpk k8s multicluster bundle command collects diagnostics from every operator in one run and writes a single ZIP file to your machine. It is the operator-side counterpart to rpk debug bundle, which collects data from the brokers. For an operator problem, Redpanda Support usually needs both.

Prerequisites

  • Redpanda Operator version 26.2.1 or later on every cluster.

  • rpk version 26.2 or later with the k8s plugin. rpk installs the plugin the first time you run an rpk k8s command, or you can install it in advance with rpk k8s install.

  • A kubeconfig with access to at least one of the Kubernetes clusters in the Stretch Cluster.

  • Permissions in the operator namespace of each cluster to read the operator’s Pods, Deployment, logs, and Secrets, to port-forward to the operator Pod (pods/portforward), and to create ServiceAccount tokens (serviceaccounts/token). The command uses the token to scrape the operator’s /metrics endpoint.

Generate the bundle

  1. Run the command against any one cluster. The operator keeps a kubeconfig for each peer in a labeled Secret in its namespace, so the command discovers the remaining clusters from the one you point it at:

    rpk k8s multicluster bundle \
      --kubeconfig <path-to-kubeconfig> \
      --context <cluster-context> \
      --namespace <operator-namespace>

    If you omit --namespace, the command uses redpanda.

    If the cached peer kubeconfigs use server addresses that aren’t reachable from your machine, name every cluster instead. With more than one --context, discovery is skipped:

    rpk k8s multicluster bundle \
      --context <cluster-a-context> \
      --context <cluster-b-context> \
      --context <cluster-c-context> \
      --namespace <operator-namespace>

    By default, the command takes two /metrics samples 10 seconds apart on each cluster, one cluster after another, so expect it to run for at least 10 seconds per cluster.

  2. Find the ZIP file in the current directory, named operator-bundle-<timestamp>.zip, or at the path you passed with -o.

  3. Check status.txt inside the bundle for the list of clusters that were collected. If anything failed, such as a missing cluster or a failed metrics scrape, the bundle also contains errors.txt listing it. These failures don’t stop the rest of the bundle from being collected.

The bundle contains a manifest.json file at the root and a clusters/<cluster-name>/ directory for each cluster. A cluster you name with --context uses the context name, and a discovered peer uses its peer name from the operator. Each directory holds the manifests of one running operator Pod and its Deployment, the operator container arguments, the logs for each container, including the previous log after a restart, /metrics samples, Raft status, the TLS certificates, and the results of the per-cluster checks. A cross-cluster/checks.json file holds checks that span clusters, such as Raft quorum. managedFields are removed from all Kubernetes objects, and TLS private keys and cached kubeconfigs are redacted unless you pass --include-private-keys.

For the options that control log size, metrics sampling, and redaction, see Configure the multicluster operator debug bundle.