
Code qualitative data with an LLM
qlm_code.RdApplies a codebook to input data using a large language model, returning a rich object that includes the codebook, execution settings, results, and metadata for reproducibility.
Arguments
- x
character; the input data: texts for a text codebook, file paths or URLs for an image codebook (see the section on image input), or file paths for an audio codebook (see the section on audio input), or file paths, YouTube links and URLs of video files for a video codebook (see the section on video input). Named vectors will use names as identifiers in the output; unnamed vectors will use sequential integers. The identifiers become the
.idcolumn, on which every later operation keys, so names must be unique.- codebook
qlm_codebook; a codebook created with
qlm_codebook(). Also accepts deprecatedtask()objects for backward compatibility.- model
character; the provider (and optionally model) name in the form
"provider/model"or"provider"(which will use the default model for that provider). Native prefixes are passed toellmer::chat(). Registered prefixes, such as"dashscope/qwen-plus", resolve throughqlm_register_provider()and require an explicit model name. Examples:"openai/gpt-4o-mini","anthropic/claude-3-5-sonnet-20241022","ollama/llama3.2","openai"(uses default OpenAI model).- ...
Additional arguments passed to
ellmer::chat(),ellmer::parallel_chat_structured(), orellmer::batch_chat_structured(). Arguments recognized byellmer::parallel_chat_structured()orellmer::batch_chat_structured()are routed there; all other arguments (including provider-specific arguments likebase_url,credentials, orapi_argsfor OpenAI-compatible endpoints) are passed toellmer::chat().- batch
logical; if
TRUE, usesellmer::batch_chat_structured()instead ofellmer::parallel_chat_structured(). Batch processing is more cost-effective for large jobs but may have longer turnaround times. Default isFALSE. Seeellmer::batch_chat_structured()for details.- tools
Optional list of ellmer tool objects to register on the chat before coding: a provider's hosted web-search tool (
ellmer::openai_tool_web_search(),ellmer::claude_tool_web_search(),ellmer::google_tool_web_search()) or a custom tool fromellmer::tool(). A single tool may be passed directly. Default isNULL, no tools.Tools change the instrument: with a hosted web search the model draws on live sources rather than its training data, so they are recorded on the object, disclosed by
print()andqlm_trail(), carried to backfill passes and to a replication on the same endpoint (a provider's hosted tool belongs to that provider), and kept in the trail as name, type, description and configuration rather than as objects.Three limits. A hosted tool takes effect on both coding paths, but a custom tool only on the JSON path: the structured transport sends a custom tool's definition and runs no tool-calling loop, so the model may request it and get no result. Tools cannot be used with
batch = TRUE, which does not send them. And a hosted tool's calls are billed by the provider outside token pricing, so the run's cost is tokens only, as its cost note says.- structured
character; how the output schema is obtained.
"structured"sends it through the provider's structured-output mechanism."json"asks for JSON, puts the schema in the system prompt, and re-prompts a unit whose response does not conform."auto"(the default) attempts the structured call and falls back to"json"if it fails, or if no response it completed matched the schema. On either path every response is validated against the codebook locally before it is tabulated. The JSON path asks for JSON syntax in the field the transport takes:text.formatfor native OpenAI's Responses API andresponse_formatfor every provider ellmer reaches through Chat Completions, including registered andopenai_compatible/endpoints. Providers with neither field, Anthropic among them, use prompted JSON with the same parsing, validation and repair. See Details for which to use.- json_retries
Integer; the number of additional requests quallmer may make for a unit on the JSON path after an unusable response. Default is 2, giving at most three JSON-path requests per unit. This is implemented by quallmer and is not passed to ellmer. Each request separately uses ellmer's transport retry policy, controlled by
options(ellmer_max_tries = ). Applies on the JSON path only, so setting it alongsidestructured = "structured"is an error. What is still unusable after the run is left forbackfill.- on_error
character; what a failed request does to the rest of a parallel call, passed to
ellmer::parallel_chat_structured()or, on the JSON path,ellmer::parallel_chat()."continue"(the default) attempts every unit and records each failure in the.errorcolumn, forqlm_failures()to list andqlm_backfill()to re-code."return"stops submitting new requests after the first failure, waits for those in flight, and returns what the call has. On the structured path that call is the run: the units never sent are recorded in.erroras not completed, forqlm_backfill()to send. On the JSON path the call is one wave: a unit the wave did not reach counts as unanswered, sojson_retriessends it again in a later wave, each stopped in turn at its first failure, and whatever is still unsent at the end is recorded in.error;json_retries = 0stops after the first wave."stop"raises the first failure as an error. Applies to parallel runs only: the batch API has no equivalent, so it cannot be set withbatch = TRUE.- backfill
Logical, integer or
NULL; whether to complete the run before it is returned, by re-coding the units still failed withqlm_backfill(), using the same model and settings.FALSEor0(default) leaves the run as it came back;TRUEmakes the default number of passes, currently two; a positive integer makes at most that many;NULLmeansFALSEhere, since a fresh run has no parent whose passes could be replayed, which is whatNULLasksqlm_replicate()for. Each pass is recorded in the object's metadata, and a pass that recovers nothing ends the backfill early. Seeqlm_backfill()for what is retried and what is left alone.- prices
Optional. Rates for costing the run when ellmer cannot: a named numeric vector or list with
inputandoutput, and optionallycached_input, in US dollars per million tokens, for examplec(input = 0.435, output = 0.87, cached_input = 0.0036). Where ellmer prices the model itself its figure stands and these are not used. Acached_inputrate that is not given is taken as theinputrate. See the section on cost. Default isNULL.- name
character or
NULL; a name identifying this coding run. Default isNULL.- notes
character or
NULL; descriptive notes about this coding run. Useful for documenting the purpose or rationale when viewing results inqlm_trail(). Default isNULL.
Value
A qlm_coded object (a tibble with additional attributes):
- Data columns
The coded results with a
.idcolumn for identifiers. Atype_enum()declared"ordinal"in the codebook is an ordered factor whose levels are the enum's values in the order written; see thelevelsargument ofqlm_codebook().- Attributes
data,input_type, andrun(list containing name, batch, call, codebook, chat_args, execution_args, metadata, parent).
The object prints as a tibble and can be used directly in data manipulation workflows.
The batch flag in the run attribute indicates whether batch processing was used.
The execution_args contains all non-chat execution arguments (for either parallel or batch processing).
Details
Arguments in ... are dynamically routed to either ellmer::chat(),
ellmer::parallel_chat_structured(), or ellmer::batch_chat_structured()
based on their names.
Progress indicators and error handling are provided by the underlying
ellmer::parallel_chat_structured() or ellmer::batch_chat_structured()
function. Set verbose = TRUE to see progress messages during coding.
Retry logic for API failures should be configured through ellmer's options;
what a failure does to the rest of a parallel run is on_error.
Image input
An image codebook codes one image per element of x. A file path is read
and sent inline, after being resized as the codebook's image_file_resize
says: "high" by default, which fits the image within 2000x768 or 768x2000
pixels, "low" for 512x512, "none" to send the file as it is, or a
magick geometry string. Anything but "none" needs the magick
package, which is checked here before any request is sent. The resolution
is part of the codebook because it is part of the measurement: a poster
whose small print is legible at one size is not at another, and a
replication should read the image the original run read. Codebooks saved
before the setting existed are read as "low", which is what they were
coded at. See qlm_codebook().
A URL is passed to the provider as it is, through
ellmer::content_image_url(), so image_file_resize does not apply to
it; what the provider does with a remote image is its own affair, and not
every provider fetches URLs. The codebook's image_url_detail asks the
provider for "low" or "high" detail on such an image, where the
provider reads that field: OpenAI and OpenAI-compatible providers do,
others ignore it, and ellmer forwards it only from the version that
includes https://github.com/tidyverse/ellmer/pull/1133. When a value
other than "auto" cannot take effect, qlm_code() says so before the
run rather than recording a setting that was not applied. A path that
does not exist is refused before anything is sent, so a URL typed without
its scheme fails here with the path named, not inside the request.
Provider-specific parameters
params and api_args are forwarded to ellmer::chat() unchanged.
quallmer does not inspect or rewrite either, so which of the two a setting
belongs in is determined by ellmer and the provider, not here.
The distinction matters. ellmer::params() carries provider-agnostic
settings that ellmer translates per provider; api_args goes into the raw
request body untouched. A setting placed in the wrong one is not
necessarily rejected. For OpenAI-compatible providers ellmer maps top_k
onto the OpenAI field top_logprobs, which asks for log-probabilities per
token and has nothing to do with top-k sampling — so
params(top_k = 20) is rejected by Alibaba Model Studio
(Range of top_logprobs should be [0, 5]), while params(top_k = 3) is
accepted and silently applies no sampling setting at all. Non-OpenAI
sampling controls therefore belong in api_args:
# Qwen through Alibaba Model Studio
qlm_code(
x, codebook,
model = "openai_compatible/qwen3-max",
base_url = "https://dashscope-intl.aliyuncs.com/compatible-mode/v1",
credentials = function() {
list(Authorization = paste("Bearer", Sys.getenv("DASHSCOPE_API_KEY")))
},
params = ellmer::params(temperature = 0.6, top_p = 0.95),
api_args = list(top_k = 20, min_p = 0, enable_thinking = TRUE)
)
# Kimi K3 through Moonshot, whose temperature and top_p are fixed by the
# provider and documented as needing to be omitted rather than set
qlm_code(
x, codebook,
model = "openai_compatible/kimi-k3",
base_url = "https://api.moonshot.ai/v1",
credentials = function() {
list(Authorization = paste("Bearer", Sys.getenv("MOONSHOT_API_KEY")))
},
api_args = list(reasoning_effort = "max")
)Passing a model parameter such as temperature or max_tokens at the top
level does not work: those reach ellmer::chat(), which has no such
argument. Use params.
Cost
include_tokens = TRUE and include_cost = TRUE are forwarded to ellmer,
which adds per-unit token counts and a cost column in US dollars. ellmer
prices from a table fixed at its release, matched exactly on provider and
model, and returns NA on any miss. Some providers are absent from that
table altogether, DeepSeek among them, so no model of theirs is ever priced;
a model newer than the installed ellmer is missed on a provider it otherwise
prices, which upgrading fixes; and local endpoints such as ollama have no
per-token charge. In each case qlm_code() says so once before the run,
and the reason is kept with the object and shown when it is printed. With
include_tokens = TRUE the token counts are recorded, from which such a
run can be costed at the provider's published rates.
prices does that costing, at rates you supply from the provider's
published price list. Supplying them implies include_tokens = TRUE and
include_cost = TRUE. Only rows ellmer left NA are filled, by the same
sum ellmer applies to its own table: uncached input tokens at the input
rate, cache hits at the cached_input rate, output at the output rate,
each per million. Where ellmer priced every row itself the rates are not
used, and you are told so. The rates are kept in the run's metadata, shown
by print() and in the trail report, and reused by qlm_replicate() when
the model, the endpoint, the batch setting and the service tier are
unchanged, so a cost that rests on entered figures is always labelled as
such. quallmer bundles no prices of its own.
Schema enforcement and validation
Some providers accept a JSON Schema without enforcing it, so a response
can come back parseable but non-conforming: a number as a string, a
required property missing, an extra one added. Providers reached through
ellmer's generic OpenAI-compatible request path are all in this position:
strict = TRUE is sent and may simply be ignored. Converted straight to
a table, such a response would arrive silently as NA, or as an empty
list-column cell that a valid empty answer also produces.
So every response is validated against codebook$schema before it is
converted, on either path and whatever the provider: required properties
present and not null, scalars of the declared type without coercion,
enum values from the declared set, arrays and nested objects of the
declared shape, no undeclared properties unless the schema allows them.
A response that fails is a failed unit: its row is NA, its .error
names the offending JSON path ($.claims[2].score must be a number), and
qlm_failures() lists it for qlm_backfill() to re-code. The other
units keep their coded values. The response's usage is recorded with the
failure, since the request was billed.
structured chooses how the schema reaches the model, and what to do
when the provider ignores it:
"structured"Send the schema through the provider's structured-output mechanism. Fails loudly if the call fails; a response that does not conform is a failed unit.
"json"Ask for JSON, put the schema in the system prompt, and re-prompt with the specific validation error when a response does not conform, up to
json_retriestimes. The reliable choice for an endpoint known not to enforce."auto"Attempt the structured call; fall back to
"json"for the whole run if it errors, or if every response the provider completed fails validation, which is what an endpoint that ignored the schema produces. Requests the provider refused, and responses it cut off or filtered, are left out of that judgement, since neither says anything about the schema. The fallback re-codes the units JSON mode can help: a response cut off at the output limit, or an input rejected as longer than the context window, fails the same way on any path, so those units keep the row, reason and usage the structured attempt gave them. A run in which some responses conform keeps them, and records the rest as failed units: the intermittent kind of non-enforcement is caught per unit, not by re-coding the corpus.
Where every completed response fails validation and nothing can fall
back, under "structured", batch = TRUE or a file input, the run is
returned with every unit failed and its reason recorded, and a warning
says why: the failed rows, their reasons and their usage are what the run
has, and qlm_backfill() can retry them with another model.
On either path, qlm_failures() lists the units that produced no usable
coding, with the reason for each, and print() reports how many there
were. Most such failures are transient, and qlm_backfill() re-codes
just those units and merges them back; backfill does that before
returning. Batch processing and file inputs (image and audio codebooks)
are not supported on the JSON path, so "auto" will not fall back under
batch = TRUE or for a file input: a failed structured call then stops
with the provider's own error, and structured = "json" is refused up
front. The path actually taken is recorded in the run metadata as
backend, and a run validated this way carries validation = "local".
The schema itself must be a type_object() at the root, whose properties
become the columns of the result, built from type_string(),
type_boolean(), type_integer(), type_number(), type_enum(),
type_array() and type_object(). Any other type is refused before a
request is sent, since there would be nothing to validate a response
against.
Audio input
A codebook with input_type = "audio" codes recordings in one pass: each
file in x is uploaded to the provider through ellmer's file upload,
and the model receives a reference to it with the codebook's
instructions, so the schema can ask for a transcript, a language, a
summary or any coding of the content. Accepted formats are mp3, wav, ogg,
m4a, flac and aac.
Which providers accept audio this way is not checked in advance: the
recordings are uploaded and the model is asked, and a provider that
cannot take them refuses with its own message, to which qlm_code() adds
what is known. As of this version only Google Gemini (google_gemini/)
is known to accept an uploaded recording alongside a schema-constrained
request; OpenAI's and Anthropic's endpoints refuse, and Vertex has no
file upload. For those providers, transcribe the recordings first and
code the transcripts with a text codebook.
Every upload completes before the first coding request is sent, so either
all the inputs are ready or nothing is spent; a failed upload stops the
run with the provider's message, which says whether the failure was
transient or the file itself. Uploads expire after 48 hours and storage
is free, so qlm_replicate() and qlm_backfill() upload the files again
from the paths in x. Before they do, they check the files against the
SHA-256 hashes the run recorded for each unit, and refuse to continue if
a path now points at different bytes. The hashes are taken before
anything is uploaded, so they are of the bytes the model received even
if a file is replaced while the requests run; they are kept in the run's
metadata as input_files and reported by qlm_trail(). A backfill pass
records its own hashes for the units it re-coded, so a run coded by an
earlier version that recorded none gains them unit by unit; units still
without one are reported as unverifiable, with a notice, rather than as
changed.
batch = TRUE is not supported for audio: ellmer's batch cache is keyed
on the prompts, and an upload gets a new reference every time, so a
batch run could not be resumed.
With include_cost = TRUE, or prices, the cost of an audio run is
potentially underestimated: providers charge more per audio token than
per text token, and the figure is computed at the text rate from the
total. The run's cost note says so, in print(), the trail, and any
backfill pass.
Video input
A codebook with input_type = "video" codes picture and sound in one
pass. Each element of x is one of: the path of a local video file
(mp4, mov, avi, wmv or webm), which is uploaded to the
provider; a YouTube link, which the provider fetches itself, so nothing
is uploaded; or the URL of a video file, which is downloaded here and then
uploaded like a file, and must end in one of those extensions. The model
receives a reference to each with the codebook's instructions, so the
schema can ask for a transcript, a description of what is shown, or any
coding of the content. A video carries its audio track, so speech and
picture are coded together.
As with audio, which providers accept video is not checked in advance; a
provider that cannot take it refuses with its own message, and
qlm_code() adds what is known. As of this version only Google Gemini
(google_gemini/) accepts video, and all its chat models do. Gemini
samples one frame a second and, at its default resolution, spends
roughly 300 input tokens per second of video, so a ten-minute clip is
about 175,000 tokens and a model with a million tokens of context takes
about an hour of video per request. Before uploading, qlm_code() says
how much it is about to send: the total duration and a token estimate
when the av package is installed, the total size otherwise. A file
over 2 GB, the upload's limit, is refused. Video tokens are charged at
the text rate, so a cost from include_cost or prices needs no
qualification here, unlike audio. A YouTube video must be public; the
free tier caps YouTube input at eight hours of video a day. Gemini's own
video settings (frame rate, clip offsets, media resolution) are not
exposed by ellmer, so whole clips are coded at the defaults.
Provenance and replication work as for audio. The run records each
file's size and SHA-256 hash; for a downloaded URL those are of the
downloaded bytes, and for a YouTube link the URL alone is recorded, since
nothing passed through this machine. qlm_replicate() and
qlm_backfill() download and upload again, checking the hashes first,
and batch = TRUE is refused as for audio.
Transcripts
A qlm_transcript from qlm_transcribe() is coded as ordinary text with
a text codebook, on any provider. The run records the provenance of every
transcript, the recording's hash, the transcription model, language,
prompt and usage, in its metadata as transcription, and qlm_trail()
reports it, so the trail documents the transcription as part of the
instrument rather than starting at the text. qlm_replicate() and
qlm_backfill() code the stored transcripts again without another
transcription request, and carry the record forward.
A text unit that is NA, whether a transcription that failed or a
missing value in any character vector, is never sent to the model. It is
recorded as a failed unit with the reason, the transcription's own
message where there is one, and qlm_backfill() leaves it alone, since
there is nothing to retry. Transcribe the recording again and assign the
text at that position before coding.
Rejected runs
When the provider rejects every request with a status that will not change
on retry (400, 401, 403, 404 or 422), qlm_code() stops rather than
returning a table of NAs. The most common cause is a model name the
provider does not have, and providers rarely say so; many answer with a
bare "HTTP 400 Bad Request". So before reporting, qlm_code() asks the
provider for its model list, through ellmer's models_<provider>(), and
says when the name is not on it, with the nearest names it does have. The
lookup runs only after a run has failed, once per failed run, and only for
providers whose listing is known to cover every name they will invoke
(Bedrock, for one, invokes inference profiles its listing omits). Where
it cannot run, or the name is on the list, the provider's own error is
reported unchanged.
Under the default on_error = "continue" the parallel call does not stop
at the first refusal, so every unit is sent once before the run comes
back and is diagnosed. Each such request is refused before anything is
generated, so it is cheap, but on a large corpus there are many of them,
paced by rpm. on_error = "return" stops the structured call after the
first wave, at the cost described under that argument; on the JSON path,
whose default has always been "continue", json_retries sends the
units a wave did not reach in later waves, so "return" stops after the
first wave there only with json_retries = 0.
Incomplete runs
A run over a corpus of any size rarely comes back complete, and an
incomplete run is not an error. Under the default on_error = "continue"
every unit is attempted, and a unit that produced no usable coding is
returned as a row of NA with the reason in its .error column;
on_error says what a failure does to the rest of the run.
qlm_failures() lists the failed units with their reasons, and print()
counts them. Trying again happens in layers. ellmer retries each request
at the transport level, on every path. On the JSON path, json_retries
sends a unit again during the run when its response was unusable. After
the run, backfill, or qlm_backfill() on the returned object, re-codes
what is still failed and merges it back, with the same model and
settings. A backfill leaves alone the two failures that re-sending the
same request cannot fix, a text rejected as longer than the context
window and a response cut off at max_tokens (see below); those need a
different model or a higher params(max_tokens = ). The section "When
a run comes back incomplete" of the workflow guide
(https://quallmer.github.io/quallmer/articles/pkgdown/getting-started/workflow.html)
walks through this on a run that ships with the package, and its table
matches each kind of failure to the mechanism that handles it.
Truncated responses
A response that runs into the provider's output limit (max_tokens) is cut
off mid-JSON, and the request is billed in full. The affected units are
systematically the longest and richest ones, and for a codebook where an
empty answer is a legitimate outcome, a cut-off answer that reads as empty
is the worst kind of silent failure.
The provider's finish reason travels with every response, on either path,
and is read before the response is parsed: a unit cut off this way is
recorded in .error with the token count, whether or not the fragment
happens to parse, is listed by qlm_failures(), and is not retried: a
repair prompt cannot supply what the limit withheld, and would only press
the model into a shorter answer. A backfill leaves it alone too; raise
params(max_tokens = ) and backfill again. A response the provider
withheld under a content filter, or finished for a reason it did not
recognise, is recorded the same way with that reason.
When batch = TRUE, the function uses ellmer::batch_chat_structured()
which submits jobs to the provider's batch API. This is typically more
cost-effective but has longer turnaround times. The path argument specifies
where batch results are cached, wait controls whether to wait for completion,
and ignore_hash can force reprocessing of cached results. on_error does
not apply: the batch API has no equivalent.
Registered providers
OpenAI-compatible endpoints can also be addressed by a registered prefix,
for example model = "dashscope/qwen-plus". See
qlm_register_provider() for built-in endpoints, credential sources,
and adding a private gateway. Native ellmer prefixes take precedence.
See also
qlm_codebook() for creating codebooks, qlm_replicate() for replicating
coding runs, qlm_compare() and qlm_validate() for assessing reliability.
Examples
# Requires API credentials and internet access; not run in package checks.
if (FALSE) { # \dontrun{
# Basic sentiment analysis
texts <- c("I love this product!", "Terrible experience.", "It's okay.")
coded <- qlm_code(texts, data_codebook_sentiment, model = "openai/gpt-4o-mini")
coded
# With named inputs (names become IDs in output)
texts_named <- c(review1 = "Great service!", review2 = "Very disappointing.")
coded2 <- qlm_code(texts_named, data_codebook_sentiment, model = "openai/gpt-4o-mini")
coded2
# Audio recordings, coded in one pass by a Gemini model; see the section
# "Audio input" for which providers accept audio. The model hears the
# recording, so the schema can ask for the transcript as well as the codes
speech_codebook <- qlm_codebook(
"Speech", "Transcribe the recording, identify the language and summarise what is said.",
ellmer::type_object(
transcript = ellmer::type_string("Verbatim transcript, in the language spoken"),
language = ellmer::type_string("Language spoken"),
summary = ellmer::type_string("One-sentence summary in English")
),
input_type = "audio"
)
coded_audio <- qlm_code(
c(interview1 = "interview1.mp3", interview2 = "interview2.wav"),
speech_codebook, model = "google_gemini/gemini-2.5-flash"
)
# A video codebook takes local files, YouTube links and URLs of video
# files in one vector; see "Video input" for what is uploaded and what
# the provider fetches itself
codebook_video <- qlm_codebook(
name = "Video description",
instructions = "Describe what is shown and transcribe what is said.",
schema = ellmer::type_object(
transcript = ellmer::type_string("Verbatim transcript of the speech"),
setting = ellmer::type_string("Where the video is filmed, in a few words")
),
input_type = "video"
)
coded_video <- qlm_code(
c(clip = "clip.mp4", zoo = "https://www.youtube.com/watch?v=jNQXAC9IVRw"),
codebook = codebook_video,
model = "google_gemini/gemini-2.5-flash"
)
} # }