openai_transcription
Generates a transcription of spoken audio in the input language, using the OpenAI API.
Introduced in version 4.32.0.
This processor sends an audio file object along with the input language to OpenAI API to generate a transcription. By default, the processor submits the entire payload of each message as a string, unless you use the file configuration field to customize it.
To learn more about audio transcription, see the: OpenAI API documentation.
-
Common
-
Advanced
processor:
label: ""
openai_transcription:
server_address: https://api.openai.com/v1
api_key: "" # No default (required)
model: "" # No default (required)
file: "" # No default (required)
processor:
label: ""
openai_transcription:
server_address: https://api.openai.com/v1
api_key: "" # No default (required)
model: "" # No default (required)
file: "" # No default (required)
language: "" # No default (optional)
prompt: "" # No default (optional)
Fields
api_key
The API secret key for OpenAI API.
|
This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see Secrets. |
Type: string
file
The audio file object (not file name) to transcribe, in one of the following formats: flac, mp3, mp4, mpeg, mpga, m4a, ogg, wav, or webm.
Type: string
language
The language of the input audio. Supplying the input language in ISO-639-1 format improves accuracy and latency.
This field supports interpolation functions.
Type: string
# Examples:
language: en
# ---
language: fr
# ---
language: de
# ---
language: zh
prompt
Optional text to guide the model’s style or continue a previous audio segment. The prompt should match the audio language.
This field supports interpolation functions.
Type: string
server_address
The OpenAI API endpoint to which the processor sends requests. Update the default value to use a different OpenAI-compatible service.
Type: string
Default: https://api.openai.com/v1