Connect

openai_embeddings

Generates vector embeddings to represent input text, using the OpenAI API.

Introduced in version 4.32.0.

This processor sends text strings to the OpenAI API, which generates vector embeddings. By default, the processor submits the entire payload of each message as a string, unless you use the text_mapping configuration field to customize it.

To learn more about vector embeddings, see the OpenAI API documentation.

  • Common

  • Advanced

processor:
  label: ""
  openai_embeddings:
    server_address: https://api.openai.com/v1
    api_key: "" # No default (required)
    model: "" # No default (required)
    text_mapping: "" # No default (optional)
    dimensions: 0 # No default (optional)
processor:
  label: ""
  openai_embeddings:
    server_address: https://api.openai.com/v1
    api_key: "" # No default (required)
    model: "" # No default (required)
    text_mapping: "" # No default (optional)
    dimensions: 0 # No default (optional)

Fields

api_key

The API secret key for OpenAI API.

This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see Secrets.

Type: string

dimensions

The number of dimensions the resulting output embeddings should have. Only supported in text-embedding-3 and later models.

Requires version 4.34.0 or later.

Type: int

model

The name of the OpenAI model to use.

Type: string

# Examples:
model: text-embedding-3-large

# ---

model: text-embedding-3-small

# ---

model: text-embedding-ada-002

server_address

The OpenAI API endpoint to which the processor sends requests. Update the default value to use a different OpenAI-compatible service.

Type: string

text_mapping

A Bloblang mapping that returns the text you want to generate vector embeddings for. By default, the processor submits the entire payload of each message as a string.

Type: string

Examples

Store embedding vectors in Pinecone

Computes embeddings for some generated data and stores them in Pinecone.

input:
  generate:
    interval: 1s
    mapping: |
      root = {"text": fake("paragraph")}
pipeline:
  processors:
  - openai_embeddings:
      model: text-embedding-3-large
      api_key: "${OPENAI_API_KEY}"
      text_mapping: "root = this.text"
output:
  pinecone:
    host: "${PINECONE_HOST}"
    api_key: "${PINECONE_API_KEY}"
    id: "root = uuid_v4()"
    vector_mapping: "root = this"

Store embedding vectors in CyborgDB

Computes embeddings for some generated data and stores them in CyborgDB.

input:
  generate:
    interval: 1s
    mapping: |
      root = {"text": fake("paragraph")}
pipeline:
  processors:
  - openai_embeddings:
      model: text-embedding-3-large
      api_key: "${OPENAI_API_KEY}"
      text_mapping: "root = this.text"
output:
  cyborgdb:
    host: "${CYBORGDB_HOST}"
    api_key: "${CYBORGDB_API_KEY}"
    index_key: "${CYBORGDB_INDEX_KEY}"
    index_name: "my_encrypted_index"
    operation: "upsert"
    id: "root = uuid_v4()"
    vector_mapping: "root = this"