kafka_franz
|
Deprecated in 4.68.0
This component is deprecated and will be removed in the next major version release. Please consider moving onto the unified |
A Kafka input using the Franz Kafka client library.
-
Common
-
Advanced
input:
label: ""
kafka_franz:
seed_brokers: [] # No default (required)
topics: [] # No default (optional)
regexp_topics_include: [] # No default (optional)
regexp_topics_exclude: [] # No default (optional)
transaction_isolation_level: read_uncommitted
consumer_group: "" # No default (optional)
auto_replay_nacks: true
input:
label: ""
kafka_franz:
seed_brokers: [] # No default (required)
client_id: redpanda-connect
tls:
enabled: false
skip_cert_verify: false
enable_renegotiation: false
root_cas: ""
root_cas_file: ""
client_certs: []
sasl: [] # No default (optional)
metadata_max_age: 1m
request_timeout_overhead: 10s
conn_idle_timeout: 20s
tcp:
connect_timeout: 0s
keep_alive:
idle: 15s
interval: 15s
count: 9
tcp_user_timeout: 0s
topics: [] # No default (optional)
regexp_topics_include: [] # No default (optional)
regexp_topics_exclude: [] # No default (optional)
rack_id: ""
instance_id: ""
rebalance_timeout: 45s
session_timeout: 1m
heartbeat_interval: 3s
start_offset: earliest
fetch_max_bytes: 50MiB
fetch_max_wait: 5s
fetch_min_bytes: 1B
fetch_max_partition_bytes: 1MiB
transaction_isolation_level: read_uncommitted
consumer_group: "" # No default (optional)
checkpoint_limit: 1024
commit_period: 5s
multi_header: false
batching:
count: 0
byte_size: 0
period: ""
check: ""
processors: [] # No default (optional)
topic_lag_refresh_period: 5s
auto_replay_nacks: true
timely_nacks_maximum_wait: "" # No default (optional)
When you specify a consumer group in your configuration, this input consumes one or more topics and automatically balances the topic partitions across any other connected clients with the same consumer group. Otherwise, topics are consumed in their entirety or with explicit partitions.
This input often out-performs the traditional kafka input and provides more useful logs and error messages.
Metadata
This input adds the following metadata fields to each message:
-
kafka_key -
kafka_topic -
kafka_partition -
kafka_offset -
kafka_lag -
kafka_timestamp_ms -
kafka_timestamp_unix -
kafka_tombstone_message -
All record headers
Fields
auto_replay_nacks
Whether to automatically replay rejected messages (negative acknowledgements, or nacks) at the output level. If the cause of rejections persists, leaving this option enabled can result in back pressure.
Set auto_replay_nacks to false to delete rejected messages. Disabling auto replays can greatly improve memory efficiency of high throughput streams, as the original shape of the data is discarded immediately upon consumption and mutation.
Type: bool
Default: true
batching
Configure a batching policy that applies to individual topic partitions, so that the input batches messages together before flushing them for processing. Batching can improve performance and is useful for windowed processing. Batching this way preserves the ordering of topic partitions.
Type: object
# Examples:
batching:
byte_size: 5000
count: 0
period: 1s
# ---
batching:
count: 10
period: 1s
# ---
batching:
check: this.contains("END BATCH")
count: 0
period: 1m
batching.byte_size
The maximum total size (in bytes) that a batch can reach before it is flushed. When the combined size of all messages in the batch reaches or exceeds this limit, the batch is immediately sent to the next stage (such as a processor or output).
Set to 0 to disable size-based batching. When disabled, messages are flushed based on other conditions (such as count or period).
Type: int
Default: 0
batching.check
A Bloblang query that returns a boolean value indicating whether a message should end a batch.
Type: string
Default: ""
# Examples:
check: this.type == "end_of_transaction"
batching.count
The number of messages at which the batch is flushed. Set to 0 to disable count-based batching.
Type: int
Default: 0
batching.period
The length of time after which an incomplete batch is flushed regardless of its size. This field accepts Go duration format strings such as 100ms, 1s, or 5s. Supported time units are ns, us, ms, s, m, and h.
Type: string
Default: ""
# Examples:
period: 1s
# ---
period: 1m
# ---
period: 500ms
batching.processors[]
A list of processors to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. All resulting messages are flushed as a single batch, so splitting the batch into smaller batches with these processors has no effect.
Type: array<processor>
# Examples:
processors:
- archive:
format: concatenate
# ---
processors:
- archive:
format: lines
# ---
processors:
- archive:
format: json_array
checkpoint_limit
The maximum number of messages that are processed in parallel inside the same partition before back pressure is applied.
When a message with a specific offset is delivered to the output, the offset is only committed when all messages of previous offsets have also been delivered. This behavior ensures at-least-once delivery guarantees. However, in the event of crashes or server faults, it also increases the likelihood of duplicates. To decrease this risk, reduce the checkpoint_limit value.
Type: int
Default: 1024
client_id
An identifier for the client connection. This identifier appears in broker logs and metrics, which helps you identify the Redpanda Connect instance that is connecting.
Type: string
Default: redpanda-connect
commit_period
The period of time between each commit of the current partition offsets to the consumer group. Offsets are always committed during shutdown.
Type: string
Default: 5s
conn_idle_timeout
The approximate amount of time that connections can remain idle before they are closed. In the worst case, a connection can stay idle for up to twice this value. This field accepts Go duration format strings such as 100ms, 1s, or 5s.
Type: string
Default: 20s
consumer_group
An optional consumer group. When you specify this value:
-
The partitions of any topics specified in the
topicsfield are automatically distributed across consumers that share the consumer group. -
Partition offsets are automatically committed and resumed under this name.
Consumer groups are not supported when you specify explicit partitions to consume from in the topics field.
Type: string
fetch_max_bytes
The maximum number of bytes that a broker tries to send during a fetch.
If individual records are larger than the fetch_max_bytes value, brokers still send them.
This field is equivalent to the Java setting fetch.max.bytes.
Type: string
Default: 50MiB
fetch_max_partition_bytes
The maximum number of bytes that are consumed from a single partition in a fetch request. This field is equivalent to the Java setting fetch.max.partition.bytes.
If a single batch is larger than the fetch_max_partition_bytes value, the batch is still sent so that the client can make progress.
Type: string
Default: 1MiB
fetch_max_wait
The maximum period of time a broker can wait for a fetch response to reach the required minimum number of bytes (fetch_min_bytes). This field is equivalent to the Java setting fetch.max.wait.ms.
Type: string
Default: 5s
fetch_min_bytes
The minimum number of bytes that a broker tries to send during a fetch. This field is equivalent to the Java setting fetch.min.bytes.
Type: string
Default: 1B
heartbeat_interval
When you specify a consumer_group, heartbeat_interval sets how frequently a consumer group member should send heartbeats to Apache Kafka. Apache Kafka uses heartbeats to make sure that a group member’s session is active.
This value must be lower than session_timeout, and should be no higher than one-third of session_timeout.
This field is equivalent to the Java heartbeat.interval.ms setting and accepts Go duration format strings such as 10s or 2m.
Type: string
Default: 3s
instance_id
When you specify a consumer_group, set instance_id to define the group’s static membership, which can prevent unnecessary rebalances during reconnections. The value must be unique per consumer within the group.
When you assign an instance ID, the client does not leave the consumer group when it closes. To remove the client from the group, you must use an external admin command on behalf of the instance ID.
Type: string
Default: ""
metadata_max_age
The maximum period of time after which metadata is refreshed. This field accepts Go duration format strings such as 100ms, 1s, or 5s.
Lower values provide more responsive topic and partition discovery but may increase broker load. Higher values reduce broker queries but can delay detection of topology changes.
This interval also controls how frequently regex topic patterns are re-evaluated to discover new matching topics.
Type: string
Default: 1m
multi_header
Decode headers into lists to allow the handling of multiple values with the same key.
Type: bool
Default: false
rack_id
A rack specifies where the client is physically located, and changes fetch requests to consume from the closest replica as opposed to the leader replica.
Type: string
Default: ""
rebalance_timeout
When you specify a consumer_group, rebalance_timeout sets how long consumer group members can take to complete their work and commit offsets after a rebalance has begun. The time that a member takes to detect the rebalance (from a heartbeat) counts against this timeout. This field accepts Go duration format strings such as 100ms, 1s, or 5s.
Type: string
Default: 45s
regexp_topics_exclude[]
A list of regular expression patterns for excluding topics when regex mode is enabled (using regexp_topics_include or the deprecated regexp_topics boolean). Topics matching any of these patterns will be excluded from consumption, even if they match include patterns.
Each pattern is a full regular expression evaluated against the complete topic name. Patterns are not anchored by default, so use ^ and $ for exact matching. Exclude patterns are applied after include patterns, providing fine-grained control over topic selection.
Example: regexp_topics_exclude: ["^_", ".-temp$", ".-test.*"] excludes topics starting with underscore, ending with -temp, or containing -test.
Type: array<string>
regexp_topics_include[]
A list of regular expression patterns for matching topics to consume from. When specified, the client will periodically refresh the list of matching topics based on the metadata_max_age interval.
Each pattern is a full regular expression evaluated against the complete topic name. Patterns are not anchored by default, so logs_. matches my-logs_events and logs_errors. Use ^logs_.$ to match only topics starting with logs_.
This field enables regex mode (replacing the deprecated regexp_topics boolean) and cannot be used together with explicit topics lists. Use regexp_topics_exclude to filter out specific patterns from the matched topics.
Example: regexp_topics_include: ["events_.", "logs_."] consumes from all topics starting with events_ or logs_.
Type: array<string>
# Examples:
regexp_topics_include:
- logs_.*
- metrics_.*
# ---
regexp_topics_include:
- "events_[0-9]+"
request_timeout_overhead
Additional time to apply as overhead when calculating request deadlines. For most requests, the deadline is this overhead alone. For requests that define their own timeout field, the overhead is added on top of that timeout, which helps prevent premature timeouts.
This field is roughly equivalent to Apache Kafka’s request.timeout.ms parameter, but grants extra time to requests that have timeout fields.
Type: string
Default: 10s
sasl[]
Specify one or more methods or mechanisms of SASL authentication. They are tried in order. If the broker supports the first SASL mechanism, all connections use it. If the first mechanism fails, the client picks the first supported mechanism. If the broker does not support any client mechanisms, all connections fail.
Type: array<object>
# Examples:
sasl:
- mechanism: SCRAM-SHA-512
password: bar
username: foo
sasl[].aws.credentials
Manually configure the AWS credentials to use (optional). For more information, see the Amazon Web Services guide.
Type: object
sasl[].aws.credentials.from_ec2_role
Use the credentials of a host EC2 machine configured to assume an IAM role associated with the instance.
Type: bool
sasl[].aws.credentials.secret
The secret for the AWS credentials in use.
|
This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see Manage Secrets before adding it to your configuration. |
Type: string
sasl[].aws.credentials.token
The token for the AWS credentials in use. Required only when using short-term credentials.
Type: string
sasl[].aws.endpoint
A custom endpoint URL for AWS API requests. Use this to connect to AWS-compatible services or local testing environments instead of the standard AWS endpoints.
Type: string
sasl[].aws.tcp
Configure TCP socket-level settings to optimize network performance and reliability. These low-level controls are useful for:
-
Unresponsive hosts: Set
connect_timeoutto limit how long a connection attempt can take (the default0ssets no limit) -
Long-lived connections: Configure
keep_alivesettings to detect and recover from stale connections -
Unstable networks: Tune keep-alive probes to balance between quick failure detection and avoiding false positives
-
Linux systems with specific requirements: Use
tcp_user_timeout(Linux 2.6.37+) to control data acknowledgment timeouts
Most users should keep the default values. Only modify these settings if you’re experiencing connection stability issues or have specific network requirements.
Type: object
sasl[].aws.tcp.connect_timeout
Maximum amount of time a dial will wait for a connect to complete. Zero disables.
Type: string
Default: 0s
sasl[].aws.tcp.keep_alive.count
Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9.
Type: int
Default: 9
sasl[].aws.tcp.keep_alive.idle
Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes.
Type: string
Default: 15s
sasl[].aws.tcp.keep_alive.interval
Duration between keep-alive probes. Zero defaults to 15s.
Type: string
Default: 15s
sasl[].aws.tcp.tcp_user_timeout
Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep_alive.idle must be greater than this value per RFC 5482. Zero disables.
Type: string
Default: 0s
sasl[].extensions
Key/value pairs to add to OAUTHBEARER authentication requests.
Type: object<string>
sasl[].mechanism
The SASL mechanism to use for authentication.
Type: string
| Option | Summary |
|---|---|
|
AWS IAM based authentication as specified by the 'aws-msk-iam-auth' java library. |
|
OAuth Bearer based authentication. |
|
Plain text authentication. |
|
Redpanda Cloud Service Account authentication when running in Redpanda Cloud. |
|
SCRAM based authentication as specified in RFC5802. |
|
SCRAM based authentication as specified in RFC5802. |
|
Disable sasl authentication |
sasl[].password
The password to use for PLAIN or SCRAM-* authentication.
|
This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see Manage Secrets before adding it to your configuration. |
Type: string
Default: ""
sasl[].token
The token to use for a single session’s OAUTHBEARER authentication.
Type: string
Default: ""
seed_brokers[]
A list of broker addresses used to establish connections. If an item of the list contains commas, it is expanded into multiple addresses.
Type: array<string>
# Examples:
seed_brokers:
- "localhost:9092"
# ---
seed_brokers:
- "foo:9092"
- "bar:9092"
# ---
seed_brokers:
- "foo:9092,bar:9092"
session_timeout
When you specify a consumer_group, session_timeout sets the maximum interval between heartbeats sent by a consumer group member to the broker. If a broker doesn’t receive a heartbeat from a group member before the timeout expires, it removes the member from the consumer group and initiates a rebalance. This field accepts Go duration format strings such as 100ms, 1s, or 5s.
Type: string
Default: 1m
start_offset
Specify the offset from which this input starts or restarts consuming messages. Restarts occur when the OffsetOutOfRange error is seen during a fetch.
Type: string
Default: earliest
| Option | Summary |
|---|---|
|
Prevents consuming a partition in a group if the partition has no prior commits. Corresponds to Kafka’s |
|
Start from the earliest offset. Corresponds to Kafka’s |
|
Start from the latest offset. Corresponds to Kafka’s |
tcp
Configure TCP socket-level settings to optimize network performance and reliability. These low-level controls are useful for:
-
Unresponsive hosts: Set
connect_timeoutto limit how long a connection attempt can take (the default0ssets no limit) -
Long-lived connections: Configure
keep_alivesettings to detect and recover from stale connections -
Unstable networks: Tune keep-alive probes to balance between quick failure detection and avoiding false positives
-
Linux systems with specific requirements: Use
tcp_user_timeout(Linux 2.6.37+) to control data acknowledgment timeouts
Most users should keep the default values. Only modify these settings if you’re experiencing connection stability issues or have specific network requirements.
Type: object
tcp.connect_timeout
Maximum amount of time a dial will wait for a connect to complete. Zero disables.
Type: string
Default: 0s
tcp.keep_alive.count
Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9.
Type: int
Default: 9
tcp.keep_alive.idle
Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes.
Type: string
Default: 15s
tcp.keep_alive.interval
Duration between keep-alive probes. Zero defaults to 15s.
Type: string
Default: 15s
tcp.tcp_user_timeout
Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep_alive.idle must be greater than this value per RFC 5482. Zero disables.
Type: string
Default: 0s
timely_nacks_maximum_wait
EXPERIMENTAL: Specify a maximum period of time in which each message can be consumed and awaiting either acknowledgement or rejection before rejection is instead forced. This can be useful for avoiding situations where certain downstream components can result in blocked confirmation of delivery that exceeds SLAs. Accepts Go duration format strings such as 100ms, 1s, or 5s.
Type: string
tls
Configure Transport Layer Security (TLS) settings to secure network connections. This includes options for standard TLS as well as mutual TLS (mTLS) authentication where both client and server authenticate each other using certificates. Key configuration options include enabled to enable TLS, client_certs for mTLS authentication, root_cas/root_cas_file for custom certificate authorities, and skip_cert_verify for development environments.
Type: object
tls.client_certs[]
A list of client certificates for mutual TLS (mTLS) authentication. Configure this field to enable mTLS, authenticating the client to the server with these certificates.
You must set tls.enabled: true for the client certificates to take effect.
Certificate pairing rules: For each certificate item, provide either:
-
Inline PEM data using both
certandkeyor -
File paths using both
cert_fileandkey_file.
Mixing inline and file-based values within the same item is not supported.
Type: array<object>
Default: []
# Examples:
client_certs:
- cert: foo
key: bar
# ---
client_certs:
- cert_file: ./example.pem
key_file: ./example.key
tls.client_certs[].cert
The plaintext certificate to use for TLS authentication. Must be paired with the corresponding private key in the key field when using inline PEM data for mTLS client certificates.
Type: string
Default: ""
tls.client_certs[].cert_file
The path to a file containing the certificate to use for TLS authentication. Must be paired with the corresponding private key file in the key_file field when using file-based configuration for mTLS client certificates.
Type: string
Default: ""
tls.client_certs[].key
Private key for mTLS client certificate as inline PEM data. Must correspond to the client certificate specified in the cert field. Use this field together with cert when providing certificate data inline rather than through files.
|
This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see Manage Secrets before adding it to your configuration. |
Type: string
Default: ""
tls.client_certs[].key_file
Path to private key file for mTLS client certificate in PEM format. Must correspond to the client certificate specified in the cert_file field. Use this field together with cert_file when loading certificate data from files.
Type: string
Default: ""
tls.client_certs[].password
The password to use for the private key (specified in the key or key_file fields), if it is password-protected. The PKCS#1 and PKCS#8 formats are supported. Supports environment variable interpolation for secure password management.
The pbeWithMD5AndDES-CBC algorithm is obsolete and not supported for the PKCS#8 format. This algorithm does not authenticate the ciphertext, making it vulnerable to padding oracle attacks that can let an attacker recover the plaintext.
|
This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see Manage Secrets before adding it to your configuration. |
Type: string
Default: ""
# Examples:
password: foo
# ---
password: ${KEY_PASSWORD}
tls.enable_renegotiation
Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message local error: tls: no renegotiation.
Type: bool
Default: false
tls.enabled
Whether to enable TLS for secure connections. Set to true to enable TLS encryption. Required to be true for other TLS options (like client_certs, root_cas, etc.) to take effect.
Type: bool
Default: false
tls.root_cas
Specify a root certificate authority to use (optional). This is a string that represents a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for inline certificate data or root_cas_file for file-based certificate loading.
|
This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see Manage Secrets before adding it to your configuration. |
Type: string
Default: ""
# Examples:
root_cas: |-
-----BEGIN CERTIFICATE-----
...
-----END CERTIFICATE-----
tls.root_cas_file
Specify the path to a root certificate authority file (optional). This is a file, often with a .pem extension, which contains a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for file-based certificate loading or root_cas for inline certificate data.
Type: string
Default: ""
# Examples:
root_cas_file: ./root_cas.pem
tls.skip_cert_verify
Whether to skip server-side certificate verification. Set to true only for testing environments as this reduces security by disabling certificate validation. When using self-signed certificates or in development, this may be necessary, but should never be used in production. Consider using root_cas or root_cas_file to specify trusted certificates instead of disabling verification entirely.
Type: bool
Default: false
topic_lag_refresh_period
The interval between consumer lag refreshes. During each cycle, this input asks the brokers for the consumer group’s committed offsets and the partition end offsets, and records the difference (the number of unread messages) for each topic partition in the consumer lag metric and the kafka_lag metadata field. Lag is only refreshed when consumer_group is set. This field accepts Go duration format strings such as 100ms, 1s, or 5s.
Type: string
Default: 5s
topics[]
A list of topics to consume from. You can list multiple comma-separated topics in a single element.
If you specify a consumer_group, partitions are automatically distributed across consumers of a topic. Otherwise, all partitions are consumed.
Alternatively, add a colon after the topic name to set the explicit partitions to consume. For example, foo:0 consumes the partition 0 of the topic foo. This syntax also supports ranges. For example, foo:0-10 consumes all partitions from 0 through to 10 inclusive.
Finally, add another colon after the partition to set an explicit offset to consume from. For example, foo:0:10 consumes the partition 0 of the topic foo starting from the offset 10. If the offset is not present (or remains unspecified) then the field start_offset determines which offset to start from.
Type: array<string>
# Examples:
topics:
- foo
- bar
# ---
topics:
- things.*
# ---
topics:
- "foo,bar"
# ---
topics:
- "foo:0"
- "bar:1"
- "bar:3"
# ---
topics:
- "foo:0,bar:1,bar:3"
# ---
topics:
- "foo:0-5"
transaction_isolation_level
The isolation level for handling transactional messages. This setting determines how transactions are processed and affects data consistency guarantees.
Type: string
Default: read_uncommitted
| Option | Summary |
|---|---|
|
If set, only committed transactional records are processed. |
|
If set, then uncommitted records are processed. |