Connect

ollama_embeddings

Generates vector embeddings from text, using the Ollama API.

Introduced in version 4.32.0.

This processor sends text to your chosen Ollama large language model (LLM) and creates vector embeddings, using the Ollama API. Vector embeddings are long arrays of numbers that represent values or objects, in this case text.

By default, the processor starts and runs a locally installed Ollama server. Alternatively, to use an already running Ollama server, add your server details to the server_address field. You can download and install Ollama from the Ollama website.

For more information, see the Ollama documentation.

  • Common

  • Advanced

processor:
  label: ""
  ollama_embeddings:
    model: "" # No default (required)
    text: "" # No default (optional)
    runner:
      context_size: 0 # No default (optional)
      batch_size: 0 # No default (optional)
    server_address: "" # No default (optional)
processor:
  label: ""
  ollama_embeddings:
    model: "" # No default (required)
    text: "" # No default (optional)
    runner:
      context_size: 0 # No default (optional)
      batch_size: 0 # No default (optional)
      gpu_layers: 0 # No default (optional)
      threads: 0 # No default (optional)
      use_mmap: false # No default (optional)
    server_address: "" # No default (optional)
    cache_directory: "" # No default (optional)
    download_url: "" # No default (optional)

Fields

cache_directory

If server_address is not set, download the Ollama binary to this directory and use it as a model cache.

Requires version 4.34.0 or later.

Type: string

# Examples:
cache_directory: /opt/cache/connect/ollama

download_url

If server_address is not set, download the Ollama binary from this URL. The default value is the official Ollama GitHub release for this platform.

Requires version 4.34.0 or later.

Type: string

model

The name of the Ollama model to use. For a full list of models, see the Ollama website.

Type: string

# Examples:
model: nomic-embed-text

# ---

model: mxbai-embed-large

# ---

model: snowflake-artic-embed

# ---

model: all-minilm

runner

Options for the model runner that are used when the model is first loaded into memory.

Requires version 4.34.0 or later.

Type: object

runner.batch_size

The maximum number of requests to process in parallel.

Type: int

runner.context_size

Sets the size of the context window used to generate the next token. Using a larger context window uses more memory and takes longer to process.

Type: int

runner.gpu_layers

Sets the number of layers to offload to the GPU for computation. This generally results in increased performance. By default, the runtime decides the number of layers dynamically.

Type: int

runner.threads

Sets the number of threads to use during response generation. For optimal performance, set this value to the number of physical CPU cores your system has. By default, the runtime decides the optimal number of threads.

Type: int

runner.use_mmap

Map the model into memory. Set to true to load only the necessary parts of the model into memory. This setting is only supported on Unix systems.

Type: bool

server_address

The address of the Ollama server to use. Leave this field blank and the processor starts and runs a local Ollama server, or specify the address of your own local or remote server.

Type: string

# Examples:
server_address: http://127.0.0.1:11434

text

The text you want to generate vector embeddings for. By default, the processor submits the entire payload of each message as a string.

This field supports interpolation functions.

Type: string

Examples

Store embedding vectors in Qdrant

Computes embeddings for some generated data and stores them in Qdrant.

input:
  generate:
    interval: 1s
    mapping: |
      root = {"text": fake("paragraph")}
pipeline:
  processors:
  - ollama_embeddings:
      model: snowflake-artic-embed
      text: "${!this.text}"
output:
  qdrant:
    grpc_host: localhost:6334
    collection_name: "example_collection"
    id: "root = uuid_v4()"
    vector_mapping: "root = this"

Store embedding vectors in CyborgDB

Computes embeddings for some generated data and stores them in CyborgDB.

input:
  generate:
    interval: 1s
    mapping: |
      root = {"text": fake("paragraph")}
pipeline:
  processors:
  - ollama_embeddings:
      model: snowflake-artic-embed
      text: "${!this.text}"
output:
  cyborgdb:
    host: "${CYBORGDB_HOST}"
    api_key: "${CYBORGDB_API_KEY}"
    index_key: "${CYBORGDB_INDEX_KEY}"
    index_name: "my_encrypted_index"
    operation: "upsert"
    id: "root = uuid_v4()"
    vector_mapping: "root = this"

Store embedding vectors in Clickhouse

Compute embeddings for some generated data and store it within Clickhouse

input:
  generate:
    interval: 1s
    mapping: |
      root = {"text": fake("paragraph")}
pipeline:
  processors:
  - branch:
      processors:
      - ollama_embeddings:
          model: snowflake-artic-embed
          text: "${!this.text}"
      result_map: |
        root.embeddings = this
output:
  sql_insert:
    driver: clickhouse
    dsn: "clickhouse://localhost:9000"
    table: searchable_text
    columns: ["id", "text", "vector"]
    args_mapping: "root = [uuid_v4(), this.text, this.embeddings]"