martech cookbookSearch recipes & patternsSign up
Recipe·Updated 23 September 2026

LLM intent extraction from unstructured surfaces

Turn search queries, chat, support tickets, and reviews into structured intent signals, without mistaking sentiment for intent

01Problem

Most of what customers tell a business is in text the business never structures: the words they type into site search, the questions in a support chat, the frustration in a ticket, the specifics in a review. Clickstream captures that someone looked at a product; the search query "waterproof jacket for hiking under 200" captures what they actually want, in their words. That intent is sitting in unstructured surfaces, and until recently it was too expensive to extract at scale, so it stayed unused while teams inferred intent from clicks that carry a fraction of the signal.

Language models make extraction cheap enough to do at scale, which is genuinely new and genuinely useful. It is also where teams get into trouble, because the easy version (run text through a model, get a label, act on it) skips the parts that make the signal trustworthy: distinguishing intent from sentiment, keeping extraction consistent across runs, joining the signal to a person, and processing sensitive transcripts under a defensible basis.

This recipe extracts structured intent from unstructured surfaces in a way the rest of the stack can actually rely on.

02Outcome

Intent signals derived from what customers said, available as structured attributes for personalization, routing, and measurement. The metric is intent signal coverage: the share of meaningful interactions from which a usable, identity-joined intent signal is extracted, versus the near-zero most stacks start at because the text is never processed. The value depends on how much the business's customers express in text (a considered-purchase category with rich search and support interaction yields far more than an impulse category with terse queries) and on whether the extracted signal is joined to identity, because intent captured in aggregate but not tied to a person is interesting and unusable.

The second outcome is honesty about a signal that was previously guessed: intent inferred from a stated query is closer to truth than intent inferred from a click, provided the extraction does not overreach.

03Ingredients
  • Unstructured interaction text
  • Extraction prompt definition
  • Customer identifier
  • Consent state flag
04Equipment

A language model (hosted API or self-hosted, with the choice driven heavily by data-residency and privacy posture rather than by benchmark scores) plus a pipeline that pulls text from the source surfaces, runs extraction against a canonical prompt, validates the output against a schema, and writes the structured signal to the warehouse joined to identity. Composable stacks suit this because the extracted signal lands beside the rest of the customer data where downstream recipes read it. The capability that decides whether the signal is trustworthy is evaluation: a held-out set of inputs with known correct extractions, run regularly, so prompt or model changes that shift the output are caught rather than silently changing what every downstream recipe consumes.

05Staff
  • Data science
    CriticalThe extraction model, the canonical prompt, output schema, consistency evaluation
  • Data engineering
    CriticalThe pipeline from source surfaces to structured signal joined to identity
  • Privacy ops
    CriticalThe basis for processing transcripts, the data-residency and retention of inputs
  • Marketing ops
    SupportingWhich downstream recipes consume the intent signal and how

Data science owns the extraction model, the canonical prompt, the output schema, and the consistency evaluation that keeps the signal stable. Data engineering builds the pipeline from source surfaces to identity-joined structured signal. Privacy ops is critical, because processing customer transcripts through a model, especially a third-party API, raises real questions of basis, data residency, and how long the raw inputs are retained, and getting this wrong in the EU is a meaningful exposure. Marketing ops decides which downstream recipes consume the intent and how, since an intent signal with no consumer is just a cost. The recipe takes months because the prompt and schema need iteration against real text to extract reliably, and because the privacy review for processing transcripts is substantive. Calling a model on some text is a day's work; a consistent, identity-joined, defensible intent pipeline is the part that takes time.

06Technique
INPUTSPROCESSACTIVATIONUnstructured textCustomeridentifierConsent state flagConsent gateExtraction promptLLM extractvalidatePrompt evalmetricsIntent attributesPersonalizationMIDRoutingEARLYMeasurementLATE

Work in this order. Fix the schema and the prompt before anything consumes the signal.

  1. Keep intent and sentiment as separate fields. A frustrated review of a product the customer uses daily is negative sentiment, not churn intent; collapsing the two produces confidently misdirected campaigns.
  2. Write one canonical, version-controlled extraction prompt.
  3. Build the pipeline from source surface to an identity-joined signal.
  4. Evaluate on a held-out set and check consistency across runs, so the same input yields the same signal as the model and prompt change underneath.
  5. Store the output as an inferred attribute, with a confidence and a refresh cycle, not a fact about the customer.
  6. Choose model hosting by data residency. A third-party API can move transcripts outside their jurisdiction.
  7. Settle the processing basis and raw-input retention. Running chat and support transcripts through a model is a processing decision, not an internal analytics step.
  8. Define downstream consumption: which recipes read the intent and how.

Intent signal capture from unstructured surfaces covers steps 1 to 4, inferred attribute generation steps 5 and 8, and consent-scoped signal collection steps 6 and 7; signal quality monitoring runs on the step 4 evaluation.

07Gotcha
Failure 01

The first failure, and the most common, is mistaking sentiment for intent. A model readily returns that a review is negative, and teams treat that as intent ("this customer is at risk") when sentiment and intent are correlated but distinct: a frustrated review of a product the customer loves and uses daily is not churn intent. Conflating the two produces confidently misdirected campaigns. The extraction has to target intent specifically, and the schema should keep sentiment and intent as separate fields rather than collapsing them.

Failure 02

The second is inconsistency across runs. Without a canonical prompt and evaluation, the same query processed today and next month yields different structured signals as the model updates or the prompt drifts, and every downstream recipe inherits the instability without knowing it. The version-controlled prompt and the held-out evaluation set are what make the signal a dependable input rather than a moving one.

Failure 03

The third is the privacy gap. Processing chat and support transcripts through a language model, particularly a third-party API, moves sensitive personal data outside its original collection context, sometimes outside the jurisdiction. Teams that treat this as an internal analytics step rather than as a processing decision with a basis and a data-residency question build exposure that surfaces in an audit. The basis, the residency, and the retention of the raw inputs are part of the recipe, not paperwork after it.

WorkshopFor your stack·The questions this recipe raises

Eight questions this recipe raises for your stack.

The Workshop works out with your team which of these matter for your stack right now, and what to do first: a 90-minute session with the people who own the decision.

  1. 01Output schema: intent and sentiment as separate fieldsThe structure that stops a negative review getting read as churn intent.
  2. 02Canonical, version-controlled extraction promptThe prompt and schema that make the same input produce the same signal across runs.
  3. 03Source-surface pipeline to identity-joined signalPulling search, chat, tickets, and reviews into the warehouse beside the rest of the customer data.
  4. 04Held-out evaluation set and consistency checksCatching prompt or model changes that quietly shift what every downstream recipe consumes.
  5. 05Inferred-attribute framing: confidence and refresh cycleTreating extracted intent as a model output, not a fact about the customer.
  6. 06Model hosting choice driven by data residencyWhen a third-party API moves personal data outside its jurisdiction, and the self-hosted alternative.
  7. 07Processing basis and raw-input retentionThe basis for running transcripts through a model, and how long the inputs are kept.
  8. 08Downstream consumption: which recipes read the intent and howWiring the signal to a consumer, since intent with no consumer is just a cost.

If your customers are telling you what they want in search and support text you never structure, the Workshop is where we build the extraction that turns it into usable signal.

The intent-versus-sentiment distinction that keeps the signal honest, the canonical prompt and evaluation that keep it stable, and the privacy basis for processing transcripts in the EU: those are the decisions that turn a model call on some text into an intent signal the rest of the stack can trust.

take this to the martech workshop→

Did this recipe match your situation?Anonymous response. Sign up to leave a longer note tied to your account.

Related recipes