Skip to contents

Applies a codebook to input texts to segment them into thematic or conceptual units, returning a quanteda::corpus() where each segment is a document. This is the LLM-powered analogue of quanteda::corpus_segment().

Usage

qlm_segment(x, codebook, model, ..., prices = NULL, name = NULL, notes = NULL)

Arguments

x

A character vector of texts or a quanteda::corpus() object. Named character vectors use names as document identifiers; unnamed vectors use sequential labels (text1, text2, ...).

codebook

A codebook object created with qlm_codebook(). The schema should be a type_object() whose fields become docvars in the output corpus. Do not include a field named text; it is reserved for the verbatim segment text and is added automatically.

model

character; the provider (and optionally model) name in the form "provider/model" or "provider" (which will use the default model for that provider). Native prefixes are passed to ellmer::chat(). Registered prefixes, such as "dashscope/qwen-plus", resolve through qlm_register_provider() and require an explicit model name. Examples: "openai/gpt-4o-mini", "anthropic/claude-3-5-sonnet-20241022", "ollama/llama3.2", "openai" (uses default OpenAI model).

...

Additional arguments passed to ellmer::chat() or ellmer::parallel_chat_structured(). Arguments recognized by ellmer::parallel_chat_structured() are routed there; all other arguments (including provider-specific arguments like base_url, credentials, or api_args for OpenAI-compatible endpoints) are passed to ellmer::chat().

prices

Optional. Rates for costing the run when ellmer cannot: a named numeric vector or list with input and output, and optionally cached_input, in US dollars per million tokens. As for qlm_code(); see the section on cost there. Supplying them implies include_tokens = TRUE and include_cost = TRUE. Default is NULL.

name

character or NULL; a name identifying this coding run. Default is NULL.

notes

Optional character string with descriptive notes about this segmentation run. Default is NULL.

Value

A quanteda::corpus() where each segment is a document. Document names follow the {source}.{i} convention of quanteda::corpus_segment(). Docvars include:

docid

Name of the source document.

segid

Integer segment index within the source document.

...

Any fields defined in the codebook schema.

input_tokens, output_tokens, cached_input_tokens, cost

With include_tokens = TRUE or include_cost = TRUE: the usage of the call made for the source document, repeated on each of its segments. See the section on cost.

...

Original docvars inherited from the input (if x is a corpus).

The corpus metadata (see quanteda::meta()) carries name, continuum_lengths, and, when usage was requested, usage, a data frame with one row per source document, plus cost_note and prices where they apply.

Details

The codebook schema defines additional document-level variables (docvars) for each segment. A text field (the verbatim segment text) is always added automatically and must not appear in the schema. Measurement levels defined in the codebook are not applicable to segmentation and are silently ignored.

Cost

Each source document is one request, so token counts and cost belong to the document, not to a segment. With include_tokens = TRUE or include_cost = TRUE they are recorded twice: in the corpus metadata as usage, one row per input document including documents that yielded no segments, and on each segment as docvars for convenience. Sum the metadata table for the run's total; summing the docvars counts a document once per segment, and a document that produced no segments has no docvars at all.

ellmer prices from a table fixed at its release and returns NA for a model it does not list; qlm_segment() says so once before the run, as qlm_code() does, and prices costs the run from the token counts at rates you supply. The four usage names are reserved when usage is requested: a codebook field or an inherited docvar of the same name is an error rather than silently overwritten.

See also

qlm_code() for document-level coding, qlm_codebook() for creating codebooks, quanteda::corpus_segment() for pattern-based segmentation.

Examples

if (FALSE) { # \dontrun{
# Aspect-based segmentation of a hotel review (character vector input
# returns a data.frame).
review <- paste(
  "The room was clean and tidy, despite being rather basic in its furnishings.",
  "The location of the hotel was really great, however.",
  "We loved the proximity to both public transport and to the city's main attractions."
)

cb_absa <- qlm_codebook(
  name = "Aspect-based segmentation",
  instructions = paste(
    "Segment the text according to the distinct aspects (topics or features).",
    "Each segment will continue as long as it is part of the same aspect.",
    "An aspect-based segment may be more than one sentence or may be just a",
    "part of a sentence.",
    "",
    "Aspects in hotel reviews include: cleanliness, features, location, service,",
    "and value. Return each aspect segment with its verbatim text and a short",
    "aspect label."
  ),
  schema = type_object(
    aspect    = type_string("Short aspect label"),
    sentiment = type_enum(c("negative", "neutral", "positive"),
                          "Sentiment toward this aspect")
  )
)

segs <- qlm_segment(review, cb_absa, model = "anthropic")
quanteda::docvars(segs)
#   docid segid      aspect sentiment
# 1 text1     1 cleanliness  positive
# 2 text1     2    features  negative
# 3 text1     3    location  positive

# Corpus input preserves existing docvars
reviews_corp <- quanteda::corpus(
  c(hotel_a = review),
  docvars = data.frame(city = "London", stars = 4L)
)
segs_corp <- qlm_segment(reviews_corp, cb_absa, model = "anthropic")
quanteda::docvars(segs_corp)
} # }