Generate a Debug Bundle with rpk in Kubernetes
Use rpk to generate a debug bundle from the Redpanda brokers running in Kubernetes, or from the Redpanda Operators of a Stretch Cluster.
Which bundle to generate depends on how Redpanda is deployed:
| Deployment | Generate | Command |
|---|---|---|
Single Kubernetes cluster (Redpanda Operator or Helm chart) |
Redpanda debug bundle |
|
Stretch Cluster across several Kubernetes clusters |
Redpanda debug bundle, plus the multicluster operator debug bundle when the problem involves the operators |
|
Collect the multicluster operator debug bundle when a Stretch Cluster shows operator symptoms such as a RedpandaBrokerPool that isn’t reconciled, operators that don’t form a Raft group, or TLS errors between operators. For broker symptoms such as under-replicated partitions or slow produce and consume, the Redpanda debug bundle is enough in both deployment types.
Single Kubernetes cluster
In a single-cluster deployment, the only debug bundle is the one collected from the Redpanda brokers.
To generate a debug bundle with rpk, you have two options depending on your needs:
-
Use
rpk debug bundle: Run this command directly on each broker in the cluster. This method requires access to the nodes your brokers are running on. -
Use
rpk debug remote-bundle: Run this command from a remote machine to collect diagnostics data from all brokers in the cluster. This method is ideal when you want to gather data without logging into each node individually.
Prerequisites
You must have rpk installed on your host machine.
Use rpk debug bundle
To generate a debug bundle with rpk, you can run the rpk debug bundle command on each broker in the cluster.
-
Create a ClusterRole to allow Redpanda to collect information from the Kubernetes API:
-
Operator
-
Helm
redpanda-cluster.yamlapiVersion: cluster.redpanda.com/v1alpha2 kind: Redpanda metadata: name: redpanda spec: chartRef: {} clusterSpec: serviceAccount: create: true rbac: enabled: truekubectl apply -f redpanda-cluster.yaml --namespace <namespace>-
--values
-
--set
serviceaccount.yamlserviceAccount: create: true rbac: enabled: truehelm upgrade --install redpanda redpanda/redpanda --namespace redpanda --create-namespace \ --values serviceaccount.yaml --reuse-valueshelm upgrade --install redpanda redpanda/redpanda --namespace <namespace> --create-namespace \ --set serviceAccount.create=true \ --set rbac.enabled=trueIf you aren’t using the Helm chart, you can create the ClusterRole manually:
kubectl create clusterrolebinding redpanda --clusterrole=view --serviceaccount=redpanda:default -
-
Execute the
rpk debug bundlecommand on a broker:kubectl exec -it --namespace <namespace> redpanda-0 -c redpanda -- rpk debug bundle --namespace <namespace>If you have an upload URL from the Redpanda support team, provide it in the
--upload-urlflag to upload your debug bundle to Redpanda.kubectl exec -it --namespace <namespace> redpanda-0 -c redpanda -- rpk debug bundle \ --upload-url <url> \ --namespace <namespace>Example output:
Creating bundle file... Debug bundle saved to "/var/lib/redpanda/1675440652-bundle.zip"
-
On your host machine, make a directory in which to save the debug bundle:
mkdir debug-bundle -
Copy the debug bundle ZIP file to the
debug-bundledirectory on your host machine.Replace
<bundle-name>with the name of your ZIP file.kubectl cp <namespace>/redpanda-0:/var/lib/redpanda/<bundle-name> debug-bundle/<bundle-name>.zip -
Unzip the file on your host machine and use it to debug your cluster.
cd debug-bundle unzip <bundle-name>.zipFor guidance on reading the debug bundle, see Inspect a Debug Bundle.
-
Remove the debug bundle from the Redpanda broker:
kubectl exec redpanda-0 -c redpanda --namespace <namespace> -- rm /var/lib/redpanda/<bundle-name>.zip
When you’ve finished troubleshooting, remove the debug bundle from your host machine:
rm -r debug-bundle
Use rpk debug remote-bundle
The rpk debug remote-bundle command allows you to remotely generate a consolidated debug bundle from all brokers in your cluster. This command is useful when you do not want to log into each broker’s node individually using rpk debug bundle.
-
Create an
rpkprofile to connect to your cluster. Include all brokers you want to collect data from in theadmin.hostsconfiguration.rpk profile create <profile-name> --set admin.hosts=<brokers> -
Check the configured addresses in your profile:
rpk profile printExample output:
profile.ymladmin: hosts: - broker1:9644 - broker2:9644 - broker3:9644 -
Start the debug bundle process:
rpk debug remote-bundle start --namespace <namespace>Replace
<namespace>with the Kubernetes namespace in which your Redpanda cluster is running.To skip the confirmation steps, use the
--no-confirmflag.To generate a bundle for a subset of brokers, use the
-X admin.hostsflag. For example,rpk debug remote-bundle start -X admin.hosts target-broker:9644. -
Check the status:
rpk debug remote-bundle statusExample output:
BROKER STATUS JOB-ID localhost:29644 running 7f93fd6e-fc5e-46a5-8842-717542f89e59 localhost:19644 running 7f93fd6e-fc5e-46a5-8842-717542f89e59
When the status changes to
success, the process is complete.To cancel a debug bundle process while it’s running, use rpk debug remote-bundle cancel. -
When the process is complete, download the debug bundle:
rpk debug remote-bundle downloadBy default the compressed file is downloaded to your current working directory. To choose a different location or filename, use the
--outputflag. For example:rpk debug remote-bundle download --output ~/redpanda/debug-bundles/cluster1This command results in
cluster1.zipdownloaded to the~/redpanda/debug-bundles/directory.
Unzip the file and use the contents to debug your cluster. For guidance on reading the debug bundle, see Inspect a Debug Bundle.
Configure the debug_bundle_auto_removal_seconds property to automatically remove debug bundles after a period of time. See Automatically remove debug bundles.
|
Stretch Cluster
A Stretch Cluster has two sources of diagnostics: the Redpanda brokers, and the Redpanda Operator that runs in each Kubernetes cluster. For an operator problem, Redpanda Support usually needs a bundle from both.
Redpanda debug bundle
The brokers in a Stretch Cluster form one Redpanda cluster, so the steps are the same as for a single Kubernetes cluster, but you must collect from brokers in every Kubernetes cluster:
-
For
rpk debug bundle, run the command in a broker Pod in each Kubernetes cluster, for example by passing--contexttokubectl. -
For
rpk debug remote-bundle, list brokers from every Kubernetes cluster inadmin.hosts. The command collects only from the brokers that you list.
Multicluster operator debug bundle
In a Stretch Cluster, one Redpanda Operator runs in each Kubernetes cluster and the operators coordinate through Raft. The rpk k8s multicluster bundle command collects diagnostics from every operator in one run and writes a single ZIP file to your machine. It is the operator-side counterpart to rpk debug bundle, which collects data from the brokers. For an operator problem, Redpanda Support usually needs both.
Prerequisites
-
Redpanda Operator version 26.2.1 or later on every cluster.
-
rpkversion 26.2 or later with thek8splugin.rpkinstalls the plugin the first time you run anrpk k8scommand, or you can install it in advance withrpk k8s install. -
A kubeconfig with access to at least one of the Kubernetes clusters in the Stretch Cluster.
-
Permissions in the operator namespace of each cluster to read the operator’s Pods, Deployment, logs, and Secrets, to port-forward to the operator Pod (
pods/portforward), and to create ServiceAccount tokens (serviceaccounts/token). The command uses the token to scrape the operator’s/metricsendpoint.
Generate the bundle
-
Run the command against any one cluster. The operator keeps a kubeconfig for each peer in a labeled Secret in its namespace, so the command discovers the remaining clusters from the one you point it at:
rpk k8s multicluster bundle \ --kubeconfig <path-to-kubeconfig> \ --context <cluster-context> \ --namespace <operator-namespace>If you omit
--namespace, the command usesredpanda.If the cached peer kubeconfigs use server addresses that aren’t reachable from your machine, name every cluster instead. With more than one
--context, discovery is skipped:rpk k8s multicluster bundle \ --context <cluster-a-context> \ --context <cluster-b-context> \ --context <cluster-c-context> \ --namespace <operator-namespace>By default, the command takes two
/metricssamples 10 seconds apart on each cluster, one cluster after another, so expect it to run for at least 10 seconds per cluster. -
Find the ZIP file in the current directory, named
operator-bundle-<timestamp>.zip, or at the path you passed with-o. -
Check
status.txtinside the bundle for the list of clusters that were collected. If anything failed, such as a missing cluster or a failed metrics scrape, the bundle also containserrors.txtlisting it. These failures don’t stop the rest of the bundle from being collected.
The bundle contains a manifest.json file at the root and a clusters/<cluster-name>/ directory for each cluster. A cluster you name with --context uses the context name, and a discovered peer uses its peer name from the operator. Each directory holds the manifests of one running operator Pod and its Deployment, the operator container arguments, the logs for each container, including the previous log after a restart, /metrics samples, Raft status, the TLS certificates, and the results of the per-cluster checks. A cross-cluster/checks.json file holds checks that span clusters, such as Raft quorum. managedFields are removed from all Kubernetes objects, and TLS private keys and cached kubeconfigs are redacted unless you pass --include-private-keys.
For the options that control log size, metrics sampling, and redaction, see Configure the multicluster operator debug bundle.