cassandra
Runs a query against a Cassandra database for each message in order to insert data.
Query arguments can be set using a bloblang array for the fields using the args_mapping field.
When populating timestamp columns the value must either be a string in ISO 8601 format (2006-01-02T15:04:05Z07:00), or an integer representing unix time in seconds.
Performance
This output benefits from sending multiple messages in flight in parallel for improved performance. You can tune the max number of in flight messages (or message batches) with the field max_in_flight.
This output benefits from sending messages as a batch for improved performance. Batches can be formed at both the input and output level. You can find out more in this doc.
-
Common
-
Advanced
output:
label: ""
cassandra:
addresses: [] # No default (required)
timeout: 600ms
reconnect_interval: 60s
query: "" # No default (required)
args_mapping: "" # No default (optional)
max_in_flight: 64
batching:
count: 0
byte_size: 0
period: ""
check: ""
output:
label: ""
cassandra:
addresses: [] # No default (required)
tls:
enabled: false
skip_cert_verify: false
enable_renegotiation: false
root_cas: ""
root_cas_file: ""
client_certs: []
password_authenticator:
enabled: false
username: ""
password: ""
disable_initial_host_lookup: false
max_retries: 3
backoff:
initial_interval: 1s
max_interval: 5s
timeout: 600ms
host_selection_policy:
local_dc: "" # No default (optional)
local_rack: "" # No default (optional)
reconnect_interval: 60s
exponential_reconnection:
max_retries: 0 # No default (required)
initial_interval: "" # No default (required)
max_interval: "" # No default (required)
query: "" # No default (required)
args_mapping: "" # No default (optional)
consistency: QUORUM
logged_batch: true
max_in_flight: 64
batching:
count: 0
byte_size: 0
period: ""
check: ""
processors: [] # No default (optional)
Examples
Basic Inserts
If we were to create a table with some basic columns with CREATE TABLE foo.bar (id int primary key, content text, created_at timestamp);, and were processing JSON documents of the form {"id":"342354354","content":"hello world","timestamp":1605219406} using logged batches, we could populate our table with the following config:
output:
cassandra:
addresses:
- localhost:9042
query: 'INSERT INTO foo.bar (id, content, created_at) VALUES (?, ?, ?)'
args_mapping: |
root = [
this.id,
this.content,
this.timestamp
]
batching:
count: 500
period: 1s
Insert JSON Documents
The following example inserts JSON documents into the table footable of the keyspace foospace using INSERT JSON (https://cassandra.apache.org/doc/latest/cassandra/developing/cql/json.html#insert-json).
output:
cassandra:
addresses:
- localhost:9042
query: 'INSERT INTO foospace.footable JSON ?'
args_mapping: 'root = [ this ]'
batching:
count: 500
period: 1s
Fields
addresses[]
A list of Cassandra nodes to connect to. Multiple comma separated addresses can be specified on a single line.
Type: array<string>
# Examples:
addresses:
- "localhost:9042"
# ---
addresses:
- "foo:9042"
- "bar:9042"
# ---
addresses:
- "foo:9042,bar:9042"
args_mapping
A Bloblang mapping that can be used to provide arguments to Cassandra queries. The result of the query must be an array containing a matching number of elements to the query arguments.
Requires version 3.55.0 or later.
Type: string
backoff.initial_interval
The period to wait before the first retry of a failed query. Each later retry doubles the previous wait, with random jitter of up to half this value, until backoff.max_interval is reached.
Type: string
Default: 1s
batching
Configure a batching policy.
Type: object
# Examples:
batching:
byte_size: 5000
count: 0
period: 1s
# ---
batching:
count: 10
period: 1s
# ---
batching:
check: this.contains("END BATCH")
count: 0
period: 1m
batching.byte_size
The maximum total size (in bytes) that a batch can reach before it is flushed. When the combined size of all messages in the batch reaches or exceeds this limit, the batch is immediately sent to the next stage (such as a processor or output).
Set to 0 to disable size-based batching. When disabled, messages are flushed based on other conditions (such as count or period).
Type: int
Default: 0
batching.check
A Bloblang query that returns a boolean value indicating whether a message should end a batch.
Type: string
Default: ""
# Examples:
check: this.type == "end_of_transaction"
batching.count
The number of messages at which the batch is flushed. Set to 0 to disable count-based batching.
Type: int
Default: 0
batching.period
The length of time after which an incomplete batch is flushed regardless of its size. This field accepts Go duration format strings such as 100ms, 1s, or 5s. Supported time units are ns, us, ms, s, m, and h.
Type: string
Default: ""
# Examples:
period: 1s
# ---
period: 1m
# ---
period: 500ms
batching.processors[]
A list of processors to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. All resulting messages are flushed as a single batch, so splitting the batch into smaller batches with these processors has no effect.
Type: array<processor>
# Examples:
processors:
- archive:
format: concatenate
# ---
processors:
- archive:
format: lines
# ---
processors:
- archive:
format: json_array
consistency
The consistency level to use.
Type: string
Default: QUORUM
Options: ANY, ONE, TWO, THREE, QUORUM, ALL, LOCAL_QUORUM, EACH_QUORUM, LOCAL_ONE
disable_initial_host_lookup
If enabled the driver will not attempt to get host info from the system.peers table. This can speed up queries but will mean that data_centre, rack and token information will not be available.
Type: bool
Default: false
exponential_reconnection
Configure exponential backoff for reconnection attempts to DOWN nodes. When enabled, this replaces the driver’s default constant reconnection policy with an exponential backoff strategy that gradually increases the delay between reconnection attempts. This reduces connection storm scenarios during widespread outages while ensuring eventual recovery.
Requires version 4.66.0 or later.
Type: object
exponential_reconnection.initial_interval
The initial period to wait between retry attempts.
Type: string
exponential_reconnection.max_interval
The longest wait between reconnection attempts to a node marked as DOWN.
Type: string
host_selection_policy
Advanced host selection policy settings for Cassandra clusters, highly recommended in multi-datacenter (DC) deployments. Use these options to optimize query routing in multi-DC and multi-rack deployments. By specifying a local DC and rack, you can use the DC-aware and rack-aware policies to direct queries to the closest nodes, reducing latency and improving fault tolerance. If not set, the default policy is round-robin across all available nodes. Host selection is always token-aware if the token can be calculated from the query.
Requires version 4.61.0 or later.
Type: object
# Examples:
host_selection_policy:
local_dc: dc-east
local_rack: rack1
host_selection_policy.local_dc
The name of the local datacenter to prioritize for query routing. Enables DC-aware host selection, ensuring queries are sent to nodes within this datacenter whenever possible. Recommended for clusters spanning multiple datacenters to minimize cross-DC traffic.
Type: string
host_selection_policy.local_rack
The name of the local rack to prioritize for query routing. Requires local_dc to be set. Enables rack-aware host selection, further optimizing query placement within the specified datacenter. Useful for deployments with multiple racks per datacenter to improve resilience and reduce intra-DC latency.
Type: string
logged_batch
If enabled the driver will perform a logged batch. Disabling this prompts unlogged batches to be used instead, which are less efficient but necessary for alternative storages that do not support logged batches.
Requires version 4.10.0 or later.
Type: bool
Default: true
max_in_flight
The maximum number of messages to have in flight at a given time. For outputs that send messages in batches, this limit applies to message batches. Increase this value to improve throughput.
Type: int
Default: 64
password_authenticator.password
The password to authenticate with.
|
This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see Secrets. |
Type: string
Default: ""
reconnect_interval
The interval at which Redpanda Connect attempts to reconnect to Cassandra nodes that are marked as DOWN. This setting helps maintain connectivity in unstable network conditions or during node maintenance. Use Go duration format such as 30s, 1m, or 5m. Setting this too low may create unnecessary connection attempts, while setting it too high may delay recovery from network issues.
Requires version 4.66.0 or later.
Type: string
Default: 60s
tls
Configure Transport Layer Security (TLS) settings to secure network connections. This includes options for standard TLS as well as mutual TLS (mTLS) authentication where both client and server authenticate each other using certificates. Key configuration options include enabled to enable TLS, client_certs for mTLS authentication, root_cas/root_cas_file for custom certificate authorities, and skip_cert_verify for development environments.
Type: object
tls.client_certs[]
A list of client certificates for mutual TLS (mTLS) authentication. Configure this field to enable mTLS, authenticating the client to the server with these certificates.
You must set tls.enabled: true for the client certificates to take effect.
Certificate pairing rules: For each certificate item, provide either:
-
Inline PEM data using both
certandkeyor -
File paths using both
cert_fileandkey_file.
Mixing inline and file-based values within the same item is not supported.
Type: array<object>
Default: []
# Examples:
client_certs:
- cert: foo
key: bar
# ---
client_certs:
- cert_file: ./example.pem
key_file: ./example.key
tls.client_certs[].cert
The plaintext certificate to use for TLS authentication. Must be paired with the corresponding private key in the key field when using inline PEM data for mTLS client certificates.
Type: string
Default: ""
tls.client_certs[].cert_file
The path to a file containing the certificate to use for TLS authentication. Must be paired with the corresponding private key file in the key_file field when using file-based configuration for mTLS client certificates.
Type: string
Default: ""
tls.client_certs[].key
Private key for mTLS client certificate as inline PEM data. Must correspond to the client certificate specified in the cert field. Use this field together with cert when providing certificate data inline rather than through files.
|
This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see Secrets. |
Type: string
Default: ""
tls.client_certs[].key_file
Path to private key file for mTLS client certificate in PEM format. Must correspond to the client certificate specified in the cert_file field. Use this field together with cert_file when loading certificate data from files.
Type: string
Default: ""
tls.client_certs[].password
The password to use for the private key (specified in the key or key_file fields), if it is password-protected. The PKCS#1 and PKCS#8 formats are supported. Supports environment variable interpolation for secure password management.
The pbeWithMD5AndDES-CBC algorithm is obsolete and not supported for the PKCS#8 format. This algorithm does not authenticate the ciphertext, making it vulnerable to padding oracle attacks that can let an attacker recover the plaintext.
|
This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see Secrets. |
Requires version 4.3.0 or later.
Type: string
Default: ""
# Examples:
password: foo
# ---
password: ${KEY_PASSWORD}
tls.enable_renegotiation
Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message local error: tls: no renegotiation.
Requires version 3.45.0 or later.
Type: bool
Default: false
tls.enabled
Whether to enable TLS for secure connections. Set to true to enable TLS encryption. Required to be true for other TLS options (like client_certs, root_cas, etc.) to take effect.
Type: bool
Default: false
tls.root_cas
Specify a root certificate authority to use (optional). This is a string that represents a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for inline certificate data or root_cas_file for file-based certificate loading.
|
This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see Secrets. |
Type: string
Default: ""
# Examples:
root_cas: |-
-----BEGIN CERTIFICATE-----
...
-----END CERTIFICATE-----
tls.root_cas_file
Specify the path to a root certificate authority file (optional). This is a file, often with a .pem extension, which contains a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for file-based certificate loading or root_cas for inline certificate data.
Type: string
Default: ""
# Examples:
root_cas_file: ./root_cas.pem
tls.skip_cert_verify
Whether to skip server-side certificate verification. Set to true only for testing environments as this reduces security by disabling certificate validation. When using self-signed certificates or in development, this may be necessary, but should never be used in production. Consider using root_cas or root_cas_file to specify trusted certificates instead of disabling verification entirely.
Type: bool
Default: false