Connect

file

Consumes data from files on disk, emitting messages according to a chosen codec.

Metadata

This input adds the following metadata fields to each message:

  • path

  • mod_time_unix

  • mod_time (RFC3339)

You can access these metadata fields using function interpolation.

  • Common

  • Advanced

input:
  label: ""
  file:
    paths: [] # No default (required)
    scanner:
      lines: {}
    auto_replay_nacks: true
input:
  label: ""
  file:
    paths: [] # No default (required)
    scanner:
      lines: {}
    delete_on_finish: false
    auto_replay_nacks: true

Fields

auto_replay_nacks

Whether to automatically replay rejected messages (negative acknowledgements, or nacks) at the output level. If the cause of rejections persists, leaving this option enabled can result in back pressure.

Set auto_replay_nacks to false to delete rejected messages. Disabling auto replays can greatly improve memory efficiency of high throughput streams, as the original shape of the data is discarded immediately upon consumption and mutation.

Requires version 4.27.0 or later.

Type: bool

Default: true

delete_on_finish

Whether to delete input files from the disk once they are fully consumed.

Type: bool

Default: false

paths[]

A list of paths to consume sequentially. Glob patterns are supported, including super globs (double star).

Type: array<string>

scanner

The scanner used to split the stream of bytes into individual messages. Scanners are useful for processing large data sources efficiently without holding the entire data set in memory. For example, the csv scanner processes individual rows in a CSV file without loading the entire file in memory.

Requires version 4.25.0 or later.

Type: scanner

Default:

lines: {}

Examples

Read a Bunch of CSVs

If we wished to consume a directory of CSV files as structured documents we can use a glob pattern and the csv scanner:

input:
  file:
    paths: [ ./data/*.csv ]
    scanner:
      csv: {}