Connect

cohere_embeddings

Generates vector embeddings to represent input text, using the Cohere API.

Introduced in version 4.37.0.

  • Common

  • Advanced

processor:
  label: ""
  cohere_embeddings:
    base_url: https://api.cohere.com
    api_key: "" # No default (required)
    model: "" # No default (required)
    text_mapping: "" # No default (optional)
    input_type: search_document
    dimensions: 0 # No default (optional)
processor:
  label: ""
  cohere_embeddings:
    base_url: https://api.cohere.com
    api_key: "" # No default (required)
    model: "" # No default (required)
    text_mapping: "" # No default (optional)
    input_type: search_document
    dimensions: 0 # No default (optional)

This processor sends text strings to your chosen large language model (LLM), which generates vector embeddings for them using the Cohere API. By default, the processor submits the entire payload of each message as a string, unless you use the text_mapping field to customize it.

To learn more about vector embeddings, see the Cohere API documentation.

Examples

Store embedding vectors in Qdrant

Computes embeddings for some generated data and stores them in Qdrant.

input:
  generate:
    interval: 1s
    mapping: |
      root = {"text": fake("paragraph")}
pipeline:
  processors:
  - cohere_embeddings:
      model: embed-english-v3
      api_key: "${COHERE_API_KEY}"
      text_mapping: "root = this.text"
output:
  qdrant:
    grpc_host: localhost:6334
    collection_name: "example_collection"
    id: "root = uuid_v4()"
    vector_mapping: "root = this"

Fields

api_key

Your API key for the Cohere API.

This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see Secrets.

Type: string

base_url

The base URL to use for API requests.

Type: string

dimensions

The number of dimensions (numerical values) in each vector embedding generated by this processor. This parameter only supports embed-v4.0 and newer models. Possible values are 256, 512, 1024, and 1536.

Type: int

input_type

The type of text input passed to the model.

Requires version 4.53.0 or later.

Type: string

Default: search_document

Option Summary

classification

Used for embeddings passed through a text classifier.

clustering

Used for the embeddings run through a clustering algorithm.

search_document

Used for embeddings stored in a vector database for search use-cases.

search_query

Used for embeddings of search queries run against a vector DB to find relevant documents.

model

The name of the Cohere model you want to use.

Type: string

# Examples:
model: embed-english-v3.0

# ---

model: embed-english-light-v3.0

# ---

model: embed-multilingual-v3.0

# ---

model: embed-multilingual-light-v3.0

text_mapping

A Bloblang mapping that returns the text you want to generate vector embeddings for. By default, the processor submits the entire payload of each message as a string.

Type: string