Skip to contents

Creates a codebook definition for use with qlm_code(). A codebook specifies what information to extract from input data, including the instructions that guide the LLM and the structured output schema.

Usage

qlm_codebook(
  name,
  instructions,
  schema,
  role = NULL,
  input_type = input_types(),
  levels = NULL,
  image_file_resize = NULL,
  image_url_detail = NULL
)

Arguments

name

Name of the codebook (character).

instructions

Instructions to guide the model in performing the coding task.

schema

Structured output definition, e.g., created by type_object(), type_array(), or type_enum().

role

Optional role description for the model (e.g., "You are an expert annotator"). If provided, this will be prepended to the instructions when creating the system prompt.

input_type

Type of input data: "text" (default), "image", "audio" or "video". For the other three the elements of x in qlm_code() are file paths; for "image" they may also be URLs, and for "video" YouTube links or URLs of video files. See the sections "Image input", "Audio input" and "Video input" of qlm_code() for how the files are handled and what is known about which providers accept them.

levels

Optional named list specifying measurement levels for each variable in the schema. Names should match schema property names. Values should be one of "nominal", "ordinal", "interval", or "ratio". If NULL (default), levels are auto-detected from schema types using the following mapping: type_boolean and type_enum = nominal, type_string = nominal, type_integer = ordinal, type_number = interval.

A type_enum() is nominal unless declared "ordinal" here. For an ordinal enum the values, in the order they are written, are the scale order: type_enum(c("low", "medium", "high")) with levels = list(severity = "ordinal") ranks low below medium below high. qlm_code() stores such a column as an ordered factor with those levels, and as_qlm_coded() does the same for human-coded data given this codebook, so qlm_compare() and qlm_validate() rank the categories by that order rather than alphabetically.

Names may refer to properties at any depth of the schema, including those inside a nested type_object() or the items of a type_array(). A level declared for a nested variable describes that variable in a table unnested to one row per item. qlm_compare() and qlm_validate() merge on .id, so such a table needs an identifier that is unique per item (a document-item key, not the document's .id alone) and must be re-wrapped with as_qlm_coded(), passing this codebook as codebook, for the declared levels to be found. Because levels is a flat list, a name that occurs at more than one place in the schema cannot be declared and is an error; rename the properties to make them distinct.

image_file_resize

How image files are resized before they are sent, for input_type = "image" only; setting it on any other codebook is an error. One of "high" (the default: fit within 2000x768 or 768x2000 pixels, whichever suits the orientation), "low" (fit within 512x512), "none" (send the file as it is), or a magick geometry string such as "1024x1024>". Anything other than "none" needs the magick package, and qlm_code() says so if it is missing. The value is stored on the codebook, so replications and backfills code at the same resolution and qlm_trail() records it. It applies to file paths only: an element of x that is a URL is fetched by the provider as it is. Codebooks saved before this field existed, including task() objects, are read as "low", the resolution they were coded at, rather than the current default. See ellmer::content_image_file().

image_url_detail

How much detail the provider should read from an image given as a URL, for input_type = "image" only: "auto" (the default, the provider chooses), "low" or "high". Passed to ellmer::content_image_url() for every element of x that is a URL; files are governed by image_file_resize instead. OpenAI and OpenAI-compatible providers use it and other providers ignore it, and ellmer forwards it only from the version that includes https://github.com/tidyverse/ellmer/pull/1133; qlm_code() says so before the run when a value other than "auto" cannot take effect. Stored on the codebook and recorded with the run like image_file_resize; codebooks saved before the field existed are read as "auto".

Value

A codebook object (a list with class c("qlm_codebook", "task")) containing the codebook definition. Use with qlm_code() to apply the codebook to data.

Details

This function replaces task(), which is now deprecated. The returned object has dual class inheritance (c("qlm_codebook", "task")) to maintain backward compatibility.

See also

qlm_code() for applying codebooks to data, data_codebook_sentiment for a predefined codebook example, task() for the deprecated function.

Examples

# Define a custom codebook
my_codebook <- qlm_codebook(
  name = "Sentiment",
  instructions = "Rate the sentiment from -1 (negative) to 1 (positive).",
  schema = type_object(
    score = type_number("Sentiment score from -1 to 1"),
    explanation = type_string("Brief explanation")
  )
)

# With a role
my_codebook_role <- qlm_codebook(
  name = "Sentiment",
  instructions = "Rate the sentiment from -1 (negative) to 1 (positive).",
  schema = type_object(
    score = type_number("Sentiment score from -1 to 1"),
    explanation = type_string("Brief explanation")
  ),
  role = "You are an expert sentiment analyst."
)

# With explicit measurement levels
my_codebook_levels <- qlm_codebook(
  name = "Sentiment",
  instructions = "Rate the sentiment from -1 (negative) to 1 (positive).",
  schema = type_object(
    score = type_number("Sentiment score from -1 to 1"),
    explanation = type_string("Brief explanation")
  ),
  levels = list(score = "interval", explanation = "nominal")
)

if (FALSE) { # \dontrun{
# Use with qlm_code() (requires API key)
texts <- c("I love this!", "This is terrible.")
coded <- qlm_code(texts, my_codebook, model = "openai/gpt-4o-mini")
coded
} # }