Cloud

postgres_cdc

Streams data changes from a PostgreSQL database using logical replication. There is also a configuration option to stream all existing data from the database.

  • Common

  • Advanced

input:
  label: ""
  postgres_cdc:
    dsn: "" # No default (required)
    include_transaction_markers: false
    stream_snapshot: false
    snapshot_batch_size: 1000
    schema: "" # No default (required)
    tables: [] # No default (required)
    checkpoint_limit: 1024
    temporary_slot: false
    slot_name: "" # No default (required)
    pg_standby_timeout: 10s
    pg_wal_monitor_interval: 3s
    max_parallel_snapshot_tables: 1
    auto_replay_nacks: true
    batching:
      count: 0
      byte_size: 0
      period: ""
      check: ""
input:
  label: ""
  postgres_cdc:
    dsn: "" # No default (required)
    include_transaction_markers: false
    stream_snapshot: false
    snapshot_batch_size: 1000
    schema: "" # No default (required)
    tables: [] # No default (required)
    checkpoint_limit: 1024
    temporary_slot: false
    slot_name: "" # No default (required)
    pg_standby_timeout: 10s
    pg_wal_monitor_interval: 3s
    max_parallel_snapshot_tables: 1
    unchanged_toast_value: null
    heartbeat_interval: 1h
    tls:
      skip_cert_verify: false
      enable_renegotiation: false
      root_cas: ""
      root_cas_file: ""
      client_certs: []
    aws:
      enabled: false
      region: "" # No default (optional)
      endpoint: "" # No default (required)
      id: "" # No default (optional)
      secret: "" # No default (optional)
      token: "" # No default (optional)
      role: "" # No default (optional)
      role_external_id: "" # No default (optional)
      roles: [] # No default (optional)
    signal_table_name: ""
    incremental_snapshot:
      enabled: false
      chunk_size: 1024
      heartbeat_interval: 1s
      retry_cooldown: 30s
      checkpoint_cache: "" # No default (optional)
      checkpoint_cache_key: postgres_cdc_incremental_snapshot
    auto_replay_nacks: true
    batching:
      count: 0
      byte_size: 0
      period: ""
      check: ""
      processors: [] # No default (optional)

The postgres_cdc input uses logical replication to capture changes made to a PostgreSQL database in real time and streams them to Redpanda Connect. Redpanda Connect uses this replication method to allow you to choose which database tables in your source database to receive changes from. There are also two replication modes to choose from, and an option to receive TOAST and deleted values in your data updates.

Prerequisites

  • PostgreSQL version 14 or later

  • Network access from the cluster where your Redpanda Connect pipeline is running to the source database environment. For detailed networking information, including how to set up a VPC peering connection, see Redpanda Cloud Networking.

  • Logical replication enabled on your PostgreSQL cluster

    To check whether logical replication is already enabled, run the following query:

    SHOW wal_level;

    If the wal_level value is logical, you can start to use this connector. Otherwise, choose from the following sets of instructions to update your replication settings.

  • Cloud platforms

  • Self-Hosted PostgreSQL

Use an account with sufficient permissions (superuser) to update your replication settings.

  1. Open the postgresql.conf file.

  2. Find the wal_level parameter.

  3. Update the parameter value to wal_level = logical. If you already use replication slots, you may need to increase the limit on replication slots (max_replication_slots). The max_wal_senders parameter value must also be greater than or equal to max_replication_slots.

  4. Restart the PostgreSQL server.

For this input to make a successful connection to your database, also make sure that it allows replication connections.

  1. Open the pg_hba.conf file.

  2. Update this line.

    host    replication     <replication-username>   <connector-ip>/32    md5

    Replace the following placeholders with your own values:

    • <replication-username>: The username from an account with superuser privileges.

    • <connector-ip>: The IP address of the server where you are running Redpanda Connect.

  3. Restart the PostgreSQL server.

Choose a replication mode

When you run a pipeline that uses the postgres_cdc input, Redpanda Connect connects to your PostgreSQL database and creates a replication slot. The replication slot uses a copy of the Write-Ahead Log (WAL) file to subscribe to changes in your database records as they are applied to the database. There are two replication modes you can choose from: snapshot mode and streaming mode.

In snapshot mode, Redpanda Connect first takes a snapshot of the database and streams the contents before processing changes from the WAL. In streaming mode, Redpanda Connect directly processes changes from the WAL starting from the most recent changes without taking a snapshot first.

For local testing, you can use the example pipeline on this page, which runs in snapshot mode.

Snapshot mode

If you set the stream_snapshot field to true, Redpanda Connect:

  1. Creates a snapshot of your database.

  2. Streams the contents of the tables specified in the postgres_cdc input.

  3. Starts processing changes in the WAL that occurred since the snapshot was taken, and streams them to Redpanda Connect.

Once the initial replication process is complete, the snapshot is removed and the input keeps a connection open to the database so that it can receive data updates.

If the pipeline restarts during the replication process, Redpanda Connect resumes processing data changes from where it left off. If there are other interruptions while the snapshot is taken, you may need to restart the snapshot process. For more information, see Troubleshoot replication failures.

Streaming mode

If you set the stream_snapshot field to false, Redpanda Connect starts processing data changes from the end of the WAL. If the pipeline restarts, Redpanda Connect resumes processing data changes from the last acknowledged position in the WAL.

Monitor the replication process

You can monitor the initial replication of data using the following metrics:

Metric name Description

replication_lag_bytes

Indicates how far the connector is lagging behind the source database when processing the transaction log.

postgres_snapshot_progress

Shows the progress of snapshot processing for each table.

Troubleshoot replication failures

If the database snapshot fails, the replication slot has only an incomplete record of the existing data in your database. To maintain data integrity, you must drop the replication slot manually in your source database and run the Redpanda Connect pipeline again.

SELECT pg_drop_replication_slot(SLOT_NAME);

Receive TOAST and deleted values

For full visibility of all data updates, you can also choose to stream TOAST and deleted values. To enable this option, run the following query on your source database:

ALTER TABLE large_data REPLICA IDENTITY FULL;

Data mappings

The following table shows how selected PostgreSQL data types are mapped to data types supported in Redpanda Connect. All other data types are mapped to string values.

PostgreSQL data type Bloblang value

TEXT, TIMESTAMP, UUID, VARCHAR

JSON strings, for example: this data

BOOL

Boolean JSON fields, for example: true or false

Numeric types (INT4)

JSON number types, for example: 1.

JSONB

JSON objects, for example: { "message": "message text" }

INTEGER[]

An array of integer values, for example: [1,2,3]

TEXT[]

An array of string values, for example: ["value1", "value2", "value3"]

INET

A string that contains an IP address, for example: "192.168.1.1"

POINT

A string that represents a point in a two-dimensional plane, for example: (x, y)

TSRANGE

A string that includes range bounds, for example: [2010-01-01 14:30, 2010-01-01 15:30)

TSVECTOR

A string that includes vector data, for example: "'the':2 'question':3 'is':4"

Metadata

This input adds the following metadata fields to each message:

  • table: Name of the table that the message originated from

  • operation: Type of operation that generated the message: "read", "insert", "update", or "delete". "read" is from messages that are read in the initial snapshot phase. This will also be "begin" and "commit" if include_transaction_markers is enabled

  • lsn: the log sequence number in postgres

  • schema: The table schema in benthos common schema format, compatible with processors like parquet_encode

  • commit_ts_ms: The commit timestamp of the transaction as a Unix millisecond timestamp. Not set for snapshot reads.

  • before: The pre-change state of the row for update and delete operations, in benthos common schema format. For updates, availability depends on the table’s REPLICA IDENTITY setting - with the default identity only key columns are present, with REPLICA IDENTITY FULL all columns are present.

Fields

auto_replay_nacks

Whether to automatically replay rejected messages (negative acknowledgements, or nacks) at the output level. If the cause of rejections persists, leaving this option enabled can result in back pressure.

Set auto_replay_nacks to false to delete rejected messages. Disabling auto replays can greatly improve memory efficiency of high throughput streams, as the original shape of the data is discarded immediately upon consumption and mutation.

Type: bool

Default: true

aws

AWS IAM authentication configuration for PostgreSQL instances. When enabled, IAM credentials are used to generate temporary authentication tokens instead of a static password.

This is useful for connecting to Amazon RDS or Aurora PostgreSQL instances with IAM database authentication enabled. The generated tokens are valid for 15 minutes and are automatically refreshed.

For more information about AWS credentials configuration, see the credentials for AWS guide.

Type: object

aws.enabled

Enable AWS IAM authentication for PostgreSQL. When enabled, an IAM authentication token is generated and used as the password.

Type: bool

Default: false

aws.endpoint

The PostgreSQL endpoint hostname (for example, mydb.abc123.us-east-1.rds.amazonaws.com).

Type: string

aws.id

The AWS access key ID to authenticate with. When empty, the default AWS credential chain is used.

Type: string

aws.region

The AWS region where the PostgreSQL instance is located. The region is used to sign the IAM authentication token and for STS calls when assuming roles. If no region is specified, the environment default is used.

Type: string

aws.role

Optional AWS IAM role ARN to assume for authentication. When roles is also set, this role is assumed first and the roles entries are assumed after it.

Type: string

aws.role_external_id

Optional external ID to use when assuming the role set in role. Each entry in roles sets its own external ID.

Type: string

aws.roles[]

Optional array of AWS IAM roles to assume for authentication. Roles are assumed in sequence, each using the credentials of the previous one, enabling chaining for purposes such as cross-account access. Each role can optionally specify an external ID. When role is also set, it is assumed before the first entry.

Type: array<object>

aws.roles[].role

AWS IAM role ARN to assume.

Type: string

Default: ""

aws.roles[].role_external_id

Optional external ID for the role assumption.

Type: string

Default: ""

aws.secret

The AWS secret access key that pairs with id.

This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see Manage Secrets before adding it to your configuration.

Type: string

aws.token

The AWS session token to use with id and secret. Required only when using short-term credentials.

Type: string

batching

Configure a batching policy.

Type: object

# Examples:
batching:
  byte_size: 5000
  count: 0
  period: 1s

# ---

batching:
  count: 10
  period: 1s

# ---

batching:
  check: this.contains("END BATCH")
  count: 0
  period: 1m

batching.byte_size

The maximum total size (in bytes) that a batch can reach before it is flushed. When the combined size of all messages in the batch reaches or exceeds this limit, the batch is immediately sent to the next stage (such as a processor or output).

Set to 0 to disable size-based batching. When disabled, messages are flushed based on other conditions (such as count or period).

Type: int

Default: 0

batching.check

A Bloblang query that returns a boolean value indicating whether a message should end a batch.

Type: string

Default: ""

# Examples:
check: this.type == "end_of_transaction"

batching.count

The number of messages at which the batch is flushed. Set to 0 to disable count-based batching.

Type: int

Default: 0

batching.period

The length of time after which an incomplete batch is flushed regardless of its size. This field accepts Go duration format strings such as 100ms, 1s, or 5s. Supported time units are ns, us, ms, s, m, and h.

Type: string

Default: ""

# Examples:
period: 1s

# ---

period: 1m

# ---

period: 500ms

batching.processors[]

A list of processors to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. All resulting messages are flushed as a single batch, so splitting the batch into smaller batches with these processors has no effect.

Type: array<processor>

# Examples:
processors:
  - archive:
      format: concatenate


# ---

processors:
  - archive:
      format: lines


# ---

processors:
  - archive:
      format: json_array

checkpoint_limit

The maximum number of messages that this input can process at a given time. Increasing this limit enables parallel processing, and batching at the output level. To preserve at-least-once guarantees, any given log sequence number (LSN) is not acknowledged until all messages up to that position are delivered.

Type: int

Default: 1024

dsn

The data source name (DSN) of the PostgreSQL database from which you want to stream updates. Use the format postgres://[user[:password]@][netloc][:port][/dbname][?param1=value1&…​]. PostgreSQL enforces SSL by default. To disable SSL, for example in a secure environment, add sslmode=disable to the connection string.

Type: string

# Examples:
dsn: postgres://foouser:foopass@localhost:5432/foodb?sslmode=disable

heartbeat_interval

The interval between heartbeat messages, which Redpanda Connect writes to the write-ahead log (WAL) using the pg_logical_emit_message function.

Heartbeat messages are useful when you subscribe to data changes from tables with low activity, while other tables in the database have higher-frequency updates. Without new messages to acknowledge, PostgreSQL cannot reclaim the WAL, which can exhaust the local disk. Heartbeat messages allow Redpanda Connect to periodically acknowledge new messages even when no data updates occur. Each acknowledgement advances the committed point in the WAL, which ensures that PostgreSQL can safely reclaim older log segments.

Set heartbeat_interval to 0s to disable heartbeats.

Heartbeats also pace incremental_snapshot. On a quiet table they’re the only thing that advances the snapshot. This interval keeps the slot current, and incremental_snapshot.heartbeat_interval applies alongside it: the more frequent of the two wins for the life of the input. A non-zero value is still required here.

Type: string

Default: 1h

# Examples:
heartbeat_interval: 0s

# ---

heartbeat_interval: 24h

include_transaction_markers

When set to true, creates empty messages for the BEGIN and COMMIT operations that start and complete each transaction. Messages with the operation metadata field set to begin or commit have null message payloads.

Type: bool

Default: false

incremental_snapshot

Configures chunked snapshotting that runs alongside replication streaming.

Type: object

incremental_snapshot.checkpoint_cache

A cache resource storing the snapshot’s progress, so a restart resumes instead of starting over. Required when enabled is true.

Type: string

incremental_snapshot.checkpoint_cache_key

The key used to store the incremental snapshot progress in checkpoint_cache. Use a different key if multiple incremental snapshots share the same cache.

Changing or clearing this key discards the record of which tables have been backfilled, so a snapshot-execute signal reads a table again.

Type: string

Default: postgres_cdc_incremental_snapshot

incremental_snapshot.chunk_size

The number of rows to read per chunk while incrementally snapshotting a table.

Type: int

Default: 1024

incremental_snapshot.enabled

Backfills tables in chunks alongside replication, on request. Tables are not configured here: insert a snapshot-execute row into signal_table_name to ask for one, so a backfill can be started at any time without a config change. A signal table is therefore required. Unlike stream_snapshot this needs no up-front snapshot phase and does not delay replication. The two are mutually exclusive: both read the same rows, so enabling either alongside the other would deliver everything twice.

Progress is driven by the replication stream: each streamed transaction releases a buffered chunk, and several more follow immediately if the database was idle during the read. Quiet tables therefore advance in bursts on each heartbeat, paced by heartbeat_interval.

A row can arrive twice, once from replication and once from the backfill: when a primary key reuses or fills a gap below the table’s current maximum, or, whatever the key type, when a row is inserted after replication starts but before the snapshot reaches its table. Treat rows as idempotent upserts keyed by primary key, as is standard CDC practice.

The following primary keys are currently not supported and a signal naming such a table is rejected and logged rather than started: bytea, interval, bit, bit varying, and any range or multirange type. Every other key type is supported, composite keys included.

Type: bool

Default: false

incremental_snapshot.heartbeat_interval

How often to heartbeat while enabled is true. The snapshot only advances on a streamed transaction, so on quiet tables this paces it. Raise it to reduce write load at the cost of a slower backfill.

Whichever of this and the top-level heartbeat_interval is more frequent wins, and applies for the life of the input: it is fixed at startup, and stays in force between backfills as well as during them. Heartbeats are transactional only while a backfill is in progress, since that is the only time the snapshot needs a transaction id from one; between backfills they cost nothing extra.

Type: string

Default: 1s

incremental_snapshot.retry_cooldown

How long the backfill waits before retrying a chunk read that failed transiently, usually because of a conflicting lock on the table being backfilled, such as VACUUM FULL or most ALTER TABLE statements.

The snapshot’s chunk read shares the goroutine that reads the replication stream, so each retry that blocks on the lock delays replication for every table. Retrying on every streamed transaction would repeat that delay for as long as the lock is held. Raise it to protect replication latency during a long migration, or lower it to resume the backfill sooner after a brief lock. Setting it to 0s disables the cooldown, retrying on the next streamed transaction.

Type: string

Default: 30s

max_parallel_snapshot_tables

Specify the maximum number of tables that are processed in parallel when the initial snapshot of the source database is taken.

Type: int

Default: 1

pg_standby_timeout

Specify the standby timeout after which an idle connection is refreshed to keep the connection alive.

Type: string

Default: 10s

# Examples:
pg_standby_timeout: 30s

pg_wal_monitor_interval

How often to report changes to the replication lag and write them to Redpanda Connect metrics.

Type: string

Default: 3s

# Examples:
pg_wal_monitor_interval: 6s

schema

The PostgreSQL schema from which to replicate data.

Type: string

# Examples:
schema: public

# ---

schema: '"MyCaseSensitiveSchemaNeedingQuotes"'

signal_table_name

The name of the table used to send control signals to the connector, excluding the schema. The table must exist in the schema configured via the schema field, and must not also appear in tables, since the signal table is implicitly added to the publication and excluded from snapshot scans, so listing it in both places is rejected at startup. It must have at least these columns; startup validation checks column names only, not types, so a wrong column type (for example data JSONB instead of TEXT) is only caught at runtime, on the first signal row read:

  • id: any type representable as a string (for example SERIAL, BIGSERIAL, UUID, VARCHAR)

  • type: should be VARCHAR or another string type (the signal type, see supported signals below)

  • data: should be TEXT, a JSON object containing signal parameters

Create the table with:

CREATE TABLE <schema>.<signal_table_name> (
    id   SERIAL PRIMARY KEY,
    type VARCHAR(32),
    data TEXT
);

Signal rows are published as regular output messages (operation=insert, table=<signal_table_name>). To exclude them from downstream processing, filter on the table metadata field using a mapping processor:

pipeline:
  processors:
    - mapping: |
        root = if @table == "rpcn_signal_table" { deleted() } else { this }

Supported signals

log: recognized and logged when received. The data column must contain a JSON object with a message key, whose value is written to the connector’s log output.

INSERT INTO <schema>.<signal_table_name> (type, data) VALUES ('log', '{"message": "Signal message"}');

snapshot-execute: backfills the named tables incrementally, alongside streaming. Requires incremental_snapshot.enabled. The data column must contain a JSON object with a tables key listing table names in the configured schema, excluding the schema itself:

INSERT INTO <schema>.<signal_table_name> (type, data) VALUES ('snapshot-execute', '{"tables": ["orders", "customers"]}');

Each table must appear in tables (or that list must be empty, replicating everything): an unreplicated table has no live changes to deduplicate its backfill against, so a write landing after its chunk is read would be lost. Each must also have a primary key, which the backfill pages by. A table replicated under REPLICA IDENTITY FULL without one cannot be snapshotted. That key may not be bytea: its value is read back and bound as the next chunk’s bound, and raw bytes survive neither that nor the checkpoint.

Partitioned tables

Partitioned tables are not yet supported in incremental snapshotting unless their publication sets publish_via_partition_root = true. PostgreSQL otherwise publishes their changes under the individual partitions' names while the snapshot backfill reads the parent, so updates to already-read snapshot rows cannot be detected.

A signal naming a table with partitions is rejected and logged, and it is left to replication alone. If a check cannot be run at all (for example, because the connection resets), the stream restarts and the signal is read again, so the request is not lost.

Each table joins the back of the backfill queue. A table this run already covers is skipped and logged, so a repeated signal does not re-read it. To read one again, point incremental_snapshot.checkpoint_cache_key at a fresh key.

Tables with large (TOASTed) column values

Set REPLICA IDENTITY FULL on a table with large (TOASTed) column values before backfilling using incremental snapshotting. PostgreSQL omits an unchanged TOAST value from an UPDATE, sending a marker instead, and under the default replica identity there is nothing in the message to recover it from, so unchanged_toast_value is emitted for that column. For a row updated while its chunk is buffered, the backfilled copy that held the real value is dropped as a duplicate, leaving the placeholder as the only value the destination ever receives for it. REPLICA IDENTITY FULL makes PostgreSQL send the previous row, which the connector reads the real value from. It can be set for the backfill and reverted afterwards; it takes a brief lock but rewrites nothing.

Type: string

Default: ""

# Examples:
signal_table_name: rpcn_signal_table

slot_name

The name of the PostgreSQL logical replication slot to use. If the slot does not exist, the input creates it. You can also create the slot manually before starting replication.

To avoid granting the replication user permission to create publications, you can create the publications manually ahead of time. This input uses the naming pattern pglog_stream_<replication_slot_name>, so create publications using this convention.

Starting with version 4.48.0, this input no longer adds the prefix rs_ to the names of the replication slots it creates. To keep using a replication slot that an earlier version created, add the rs_ prefix to this field yourself.

Type: string

# Examples:
slot_name: my_test_slot

snapshot_batch_size

The number of table rows to fetch in each batch when querying the snapshot.

This option is only available when stream_snapshot is set to true.

Type: int

Default: 1000

# Examples:
snapshot_batch_size: 10000

stream_snapshot

When set to true, this input streams a snapshot of all existing data in the source database before streaming data changes. To use this setting, all database tables that you want to replicate must have a primary key, which allows the input to read each table in parallel. This setting has no effect if tables is left empty, since the snapshot is only planned for tables listed there.

Type: bool

Default: false

# Examples:
stream_snapshot: true

tables[]

A list of database table names to include in the snapshot and logical replication. Specify each table name as a separate item.

If left empty, the underlying PostgreSQL publication is created FOR ALL TABLES, which replicates every table in every schema of the database, ignoring schema. This also disables stream_snapshot, since the initial snapshot is only planned for tables listed here.

Type: array<string>

# Examples:
tables:
  - my_table_1
  - '"MyCaseSensitiveTableNeedingQuotes"'

temporary_slot

If set to true, the input creates a temporary replication slot that is automatically dropped when the connection to your source database is closed. You might use this option to:

  • Avoid data accumulating in the replication slot when a pipeline is paused or stopped.

  • Test the connector.

If the pipeline is restarted and stream_snapshot is enabled, another data snapshot is taken before data updates are streamed.

Type: bool

Default: false

tls

Custom TLS settings for the PostgreSQL connection. When enabled is true, these settings replace the TLS settings derived from the dsn and from PG* environment variables such as PGSSLMODE, and the server name is set to the host from the DSN.

Type: object

tls.client_certs[]

A list of client certificates for mutual TLS (mTLS) authentication. Configure this field to enable mTLS, authenticating the client to the server with these certificates.

Certificate pairing rules: For each certificate item, provide either:

  • Inline PEM data using both cert and key or

  • File paths using both cert_file and key_file.

Mixing inline and file-based values within the same item is not supported.

Type: array<object>

Default: []

# Examples:
client_certs:
  - cert: foo
    key: bar


# ---

client_certs:
  - cert_file: ./example.pem
    key_file: ./example.key

tls.client_certs[].cert

The plaintext certificate to use for TLS authentication. Must be paired with the corresponding private key in the key field when using inline PEM data for mTLS client certificates.

Type: string

Default: ""

tls.client_certs[].cert_file

The path to a file containing the certificate to use for TLS authentication. Must be paired with the corresponding private key file in the key_file field when using file-based configuration for mTLS client certificates.

Type: string

Default: ""

tls.client_certs[].key

Private key for mTLS client certificate as inline PEM data. Must correspond to the client certificate specified in the cert field. Use this field together with cert when providing certificate data inline rather than through files.

This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see Manage Secrets before adding it to your configuration.

Type: string

Default: ""

tls.client_certs[].key_file

Path to private key file for mTLS client certificate in PEM format. Must correspond to the client certificate specified in the cert_file field. Use this field together with cert_file when loading certificate data from files.

Type: string

Default: ""

tls.client_certs[].password

The password to use for the private key (specified in the key or key_file fields), if it is password-protected. The PKCS#1 and PKCS#8 formats are supported. Supports environment variable interpolation for secure password management.

The pbeWithMD5AndDES-CBC algorithm is obsolete and not supported for the PKCS#8 format. This algorithm does not authenticate the ciphertext, making it vulnerable to padding oracle attacks that can let an attacker recover the plaintext.

This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see Manage Secrets before adding it to your configuration.

Type: string

Default: ""

# Examples:
password: foo

# ---

password: ${KEY_PASSWORD}

tls.enable_renegotiation

Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message local error: tls: no renegotiation.

Type: bool

Default: false

tls.root_cas

Specify a root certificate authority to use (optional). This is a string that represents a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for inline certificate data or root_cas_file for file-based certificate loading.

This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see Manage Secrets before adding it to your configuration.

Type: string

Default: ""

# Examples:
root_cas: |-
  -----BEGIN CERTIFICATE-----
  ...
  -----END CERTIFICATE-----

tls.root_cas_file

Specify the path to a root certificate authority file (optional). This is a file, often with a .pem extension, which contains a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for file-based certificate loading or root_cas for inline certificate data.

Type: string

Default: ""

# Examples:
root_cas_file: ./root_cas.pem

tls.skip_cert_verify

Whether to skip server-side certificate verification. Set to true only for testing environments as this reduces security by disabling certificate validation. When using self-signed certificates or in development, this may be necessary, but should never be used in production. Consider using root_cas or root_cas_file to specify trusted certificates instead of disabling verification entirely.

Type: bool

Default: false

unchanged_toast_value

Specify the value to emit when unchanged TOAST values appear in the message stream. Unchanged values occur for data updates and deletes when REPLICA IDENTITY is not set to FULL.

Prefer a distinctive sentinel over the null default. A null can’t be told apart from a column that is genuinely null, so a consumer can’t skip the field instead of overwriting a good value with it. This matters most alongside incremental_snapshot, where a backfilled row can be the only delivery that carries a large column’s real value. See signal_table_name.

Type: unknown

Default: null

# Examples:
unchanged_toast_value: __redpanda_connect_unchanged_toast_value__

Example pipeline

You can run the following pipeline locally to check that data updates are streamed from your source database to Redpanda Connect. All transactions are written to stdout.

input:
  label: "postgres_cdc"
  postgres_cdc:
    dsn: postgres://user:password@host:port/dbname
    include_transaction_markers: false
    slot_name: test_slot_native_decoder
    snapshot_batch_size: 100000
    stream_snapshot: true
    temporary_slot: true
    schema: schema_name
    tables:
      - table_name

cache_resources:
  - label: data_caching
    file:
      directory: /tmp/cache

output:
  label: main
  stdout: {}