Skip to contents

Transcribes audio files with a speech-to-text model and returns the transcripts as a character vector that carries the provenance of each one: the file it came from and its hash, the model, the language and prompt given, the time, and the usage the provider reported. Passed to qlm_code() with a text codebook, the transcripts are coded as ordinary text and that provenance is recorded with the run, so qlm_trail() documents the transcription as part of the measurement instrument and the same text can be coded again by qlm_replicate() or qlm_backfill() without another transcription request.

Usage

qlm_transcribe(
  x,
  model = "openai/gpt-4o-mini-transcribe",
  language = NULL,
  prompt = NULL,
  api_key = NULL,
  base_url = NULL,
  max_active = 10,
  rpm = 60,
  on_error = c("continue", "return", "stop"),
  ...
)

Arguments

x

character; paths of audio files, or http(s) URLs of them, optionally named. See the section "Names and identifiers".

model

character; the transcription model in "provider/model" form. See the section "Routes".

language

character; the ISO 639-1 code of the language spoken, such as "en" or "zh", or NULL to let the model detect it. Sent to the OpenAI endpoint as its language field; added to the instruction for a Gemini model.

prompt

character; optional text to guide the transcription, such as names and terms the recording contains or the style of punctuation wanted. Sent to the OpenAI endpoint as its prompt field; added to the instruction for a Gemini model.

api_key

character; the API key. NULL reads the environment variable the provider uses, OPENAI_API_KEY, the variable a registered provider was given, or the one ellmer reads for a chat provider. On the chat route the value is passed to ellmer as the credential itself, which is what Gemini and Anthropic take. The key is never recorded.

base_url

character; the endpoint to send the requests to. On the endpoint route, the prefix before /audio/transcriptions; NULL is the provider's own host. On the chat route, passed to ellmer's chat constructor. Recorded with any credential it carries redacted.

max_active

integer; the number of requests in flight at once, as in ellmer::parallel_chat().

rpm

integer; the request rate in requests per minute. The default is below ellmer's because transcription endpoints have their own, lower, rate limits. A rate-limited request is retried after the delay the provider asks for.

on_error

character; what to do when a transcription fails. See the section "Failures".

...

Reserved; must be empty.

Value

A named character vector of class qlm_transcript, one element per element of x in the same order, with attribute provenance, a data frame with one row per element:

.id

the element's name.

status

"ok", "failed" or "unsubmitted".

source

the basename of a local file, or the URL with any credential it carried redacted.

.error

the failure message, or NA.

size, sha256

the bytes transcribed and their hash; NA when a download failed.

model

as given.

language, prompt

as given, or NA.

base_url

the host the requests went to, redacted: on the endpoint route always, on the chat route when given.

timestamp

when the response arrived, or NA.

usage

a list column holding what the provider reported, or on the chat route ellmer's tokens, cost and version.

Subsetting with [, renaming with names<- and concatenating with c() keep the table aligned with the elements. Assigning a qlm_transcript with [<- or [[<- replaces rows of the table too, so a retried transcription replaces the failure it retries; assigning plain text records an edit. as.character() drops the table.

Details

This is the two-stage route to audio: transcribe once, then code the text with any provider. The single-pass route, a codebook with input_type = "audio", sends the recording itself to a model that can hear it; see the "Audio input" section of qlm_code() for the providers that accept it.

Routes

The route is chosen from the provider prefix of model, not from a list of models known to transcribe. Whether a model can is for the provider to say, and asking costs nothing: an upload is free and a refused request is not billed.

  • The transcription endpoint, /audio/transcriptions, for openai/ and for any provider registered with qlm_register_provider(), at the host it was registered with. OpenAI's models are gpt-4o-mini-transcribe (the default), gpt-4o-transcribe and whisper-1; a registered host serves whatever it serves, such as Whisper on Groq. Each file must be at most 25 MB and one of flac, mp3, mp4, mpeg, mpga, m4a, ogg, wav or webm. OpenAI reports usage as audio and text tokens for the gpt-4o models and as seconds of audio for whisper-1.

  • A chat model that hears the recording, for every other provider ellmer reaches: the file is uploaded through ellmer's file upload and the model is asked for a verbatim transcript. Known to work: Google Gemini's pro, flash and flash-lite models (google_gemini/). Anthropic takes no audio, and OpenAI's chat models refuse it; a provider that cannot take the recording says so in the failure recorded for each unit. Usage is the token count and the cost ellmer computes, with the qualification that ellmer prices audio tokens at the text rate. Gemini's dedicated transcription model, gemini-3.5-transcribe, cannot be reached through ellmer yet.

No dollar cost is computed on the endpoint route: ellmer has no rates for transcription models, and per-minute pricing does not fit a per-token table. The usage is recorded as reported so it can be costed by hand.

Names and identifiers

The names of the result become the .id of each unit when it is coded, and the document names when the vector is made a corpus. A supplied name is kept exactly. An unnamed local file is named by its basename; an unnamed URL is named text1, text2, ... by its position in x. The resolved names must be unique and non-empty: two files that share a basename need names supplied, and the error says so.

Failures

Requests run in parallel. Under on_error = "continue" every file is attempted and the result has an element for each, NA where the transcription failed, with the provider's message in the .error column of the provenance table. "return" stops submitting after the first failure and marks the files it never sent as such; "stop" raises the first error. A failed download, a failed upload on the chat route, a refused request and a response with no transcript in it are all failures of the unit, under the same policy. The one limit is on the chat route, where an empty answer is known only after every request has returned, so "return" cannot withhold submissions on its account. Validation of the arguments, the files and the model all happen before anything is downloaded or sent, and abort whatever on_error says.

A missing transcript passed to qlm_code() is never sent to the model: its unit is recorded as failed with the transcription's reason, and qlm_backfill() leaves it alone. Transcribe the file again and assign the result at that position, transcripts[failed] <- qlm_transcribe(files[failed]), which replaces the record with it; or concatenate independent runs with c().

URLs

An element of x that is an http:// or https:// URL is downloaded to a temporary file, which is removed when the function returns. The hash and size recorded are those of the downloaded bytes, and the URL is recorded, with any credential it carried redacted, as the source. The format is read from the URL's path, so a URL with no file extension is refused before anything is fetched.

See also

qlm_code() for coding the transcripts, and its "Audio input" section for the single-pass route; qlm_trail() for the record a coded transcript leaves.

Examples

if (FALSE) { # \dontrun{
files <- list.files("recordings", pattern = "\\.wav$", full.names = TRUE)
transcripts <- qlm_transcribe(files)
transcripts
attr(transcripts, "provenance")

# Code the transcripts with any provider; the run records the transcription
coded <- qlm_code(transcripts, codebook_sentiment, model = "anthropic/claude-sonnet-5")
qlm_trail(coded, path = "sentiment_trail")

# A Gemini chat model as the transcriber, with a language hint
transcripts <- qlm_transcribe(files, model = "google_gemini/gemini-2.5-flash",
                              language = "fr")

# Whisper on a registered OpenAI-compatible host
qlm_register_provider("groq", "https://api.groq.com/openai/v1", "GROQ_API_KEY")
transcripts <- qlm_transcribe(files, model = "groq/whisper-large-v3")

# A recording on the web, named so the name becomes its .id
url <- c(harvard = "https://www.voiptroubleshooter.com/open_speech/american/OSR_us_000_0010_8k.wav")
qlm_transcribe(url)
} # }