Connect

azure_blob_storage

Downloads objects within an Azure Blob Storage container, optionally filtered by a prefix.

Introduced in version 3.36.0.

Supports multiple authentication methods but only one of the following is required:

  • storage_connection_string

  • storage_account and storage_access_key

  • storage_account and storage_sas_token

  • storage_account to access via DefaultAzureCredential

If multiple are set then the storage_connection_string is given priority.

If the storage_connection_string does not contain the AccountName parameter, please specify it in the storage_account field.

Download large files

When downloading large files it’s often necessary to process it in streamed parts in order to avoid loading the entire file in memory at a given time. In order to do this a scanner can be specified that determines how to break the input into smaller individual messages.

Stream new files

By default this input will consume all files found within the target container and will then gracefully terminate. This is referred to as a "batch" mode of operation. However, it’s possible to instead configure a container as an Event Grid source and then use this as a targets_input, in which case new files are consumed as they’re uploaded and Redpanda Connect will continue listening for and downloading files as they arrive. This is referred to as a "streamed" mode of operation.

Metadata

This input adds the following metadata fields to each message:

  • blob_storage_key

  • blob_storage_container

  • blob_storage_last_modified

  • blob_storage_last_modified_unix

  • blob_storage_content_type

  • blob_storage_content_encoding

  • All user defined metadata

You can access these metadata fields using function interpolation.

  • Common

  • Advanced

input:
  label: ""
  azure_blob_storage:
    storage_account: ""
    storage_access_key: ""
    storage_connection_string: ""
    storage_sas_token: ""
    container: "" # No default (required)
    prefix: ""
    scanner:
      to_the_end: {}
    targets_input: {} # No default (optional)
input:
  label: ""
  azure_blob_storage:
    storage_account: ""
    storage_access_key: ""
    storage_connection_string: ""
    storage_sas_token: ""
    container: "" # No default (required)
    prefix: ""
    scanner:
      to_the_end: {}
    delete_objects: false
    targets_input: {} # No default (optional)

Fields

container

The name of the container from which to download blobs.

This field supports interpolation functions.

Type: string

delete_objects

Whether to delete downloaded blobs from the container once they are processed.

Type: bool

Default: false

prefix

An optional path prefix, if set only objects with the prefix are consumed.

Type: string

Default: ""

scanner

The scanner used to split the stream of bytes into individual messages. Scanners are useful for processing large data sources efficiently without holding the entire data set in memory. For example, the csv scanner processes individual rows in a CSV file without loading the entire file in memory.

Requires version 4.25.0 or later.

Type: scanner

Default:

to_the_end: {}

storage_access_key

The access key for the storage account. Use this field along with storage_account for authentication. This field is ignored when the storage_connection_string field is populated.

Type: string

Default: ""

storage_account

The storage account to access. This field is ignored when the storage_connection_string field is populated, unless the connection string does not contain the AccountName parameter.

Type: string

Default: ""

storage_connection_string

The connection string for the storage account. This field is required if storage_account is not set.

If the storage_connection_string field does not contain the AccountName parameter value, specify it in the storage_account field.

Type: string

Default: ""

storage_sas_token

The SAS token for the storage account. Use this field along with storage_account for authentication. This field is ignored when either the storage_connection_string or storage_access_key fields are populated.

Type: string

Default: ""

targets_input

This is an experimental field that provides an optional source of download targets, configured as a regular Redpanda Connect input. Each message yielded by this input should be a single structured object containing a field name, which represents the blob to be downloaded.

This requires setting up Azure Blob Storage as an Event Grid source and an associated event handler that a Redpanda Connect input can read from. For example, use either one of the following:

Requires version 4.27.0 or later.

Type: input

# Examples:
targets_input:
  mqtt:
    topics:
      - some-topic
    urls:
      - example.westeurope-1.ts.eventgrid.azure.net:8883
  processors:
    - unarchive:
        format: json_array
    - mapping: |-
        if this.eventType == "Microsoft.Storage.BlobCreated" {
          root.name = this.data.url.parse_url().path.trim_prefix("/foocontainer/")
        } else {
          root = deleted()
        }