martech cookbookSearch recipes & patternsSign up
Recipe·Updated 23 September 2026

LLM-assisted next-best-action with deterministic fallback

Let a model propose the next action when it can, and serve a known-good rule when it cannot, so the experience never depends on the model being up

01Problem

A model that picks the next best action for each customer is an appealing idea and a fragile one if built naively. Language and recommendation models are probabilistic, occasionally slow, occasionally unavailable, and occasionally confidently wrong, and a customer experience that calls the model on every decision inherits all of that. The first time the model times out during a traffic spike, or returns a brand-unsafe suggestion, or simply goes down, the experience that depended on it breaks, and it breaks at exactly the moment volume is highest.

The pattern that makes model-driven decisioning shippable is not better models; it is the deterministic fallback. The model proposes the next action when it can, within latency, with sufficient confidence, and inside the brand's guardrails. When any of those conditions fails, a deterministic rule serves a known-good action instead. The fallback is not an error path bolted on at the end; it is the design that lets the model be used in production at all, because it converts every model failure from a broken experience into a slightly less personalized but still correct one.

This recipe is model-proposes, rule-catches, for the next-best-action decision.

02Outcome

A next-best-action capability that degrades gracefully rather than breaking, with the model improving the decision when it can and a deterministic rule guaranteeing a correct decision when it cannot. The metric is next-best-action acceptance rate (do customers respond to the proposed action), but the operational health metric that matters as much is the fallback rate: how often the deterministic path served instead of the model. A low, stable fallback rate means the model is genuinely contributing; a high or rising one means the experience is mostly running on rules while the team believes it is running on a model. The lift from the model over the deterministic baseline is the honest measure of whether the model is worth its cost and risk, and it should be measured rather than assumed.

The second outcome is resilience: the experience holds during the model outage, the latency spike, and the low-confidence case, because the fallback is the default the whole thing degrades to.

03Ingredients
  • Customer context
  • Action catalog
  • Deterministic fallback rule
  • Guardrail policy
04Equipment

A real-time decisioning path that calls the model within a strict latency budget, falls back to a deterministic rule on timeout, low confidence, or guardrail violation, and monitors which path served. Composable stacks suit this because the action catalog, the fallback, and the guardrails are things the team must own and inspect, and the context the model reasons over lives in the warehouse or context layer. The capability that decides whether this is production-grade is the fallback wiring: it has to be genuinely independent of the model, so that a model outage triggers the rule rather than taking the whole decision path down with it, and it has to be fast enough that the customer never waits on a failed model call.

05Staff
  • Data science
    CriticalThe model that proposes actions, confidence calibration, the fallback trigger
  • Marketing ops
    CriticalThe action catalog, the deterministic fallback, the brand-safety guardrails
  • Data engineering
    CriticalThe request-time decisioning path, latency budget, fallback wiring, monitoring
  • Legal
    SupportingAutomated-decision exposure where the action is consequential to the customer

Data science owns the model that proposes actions, the confidence calibration, and the threshold below which the system falls back. Marketing ops owns the action catalog, the deterministic fallback rule, and the brand-safety guardrails, which are the parts that keep the model inside acceptable bounds. Data engineering owns the request-time path, the latency budget, the fallback wiring, and the monitoring. Legal covers the automated-decision exposure where the next action is consequential to the customer rather than cosmetic, which under the AI Act and adjacent regimes is a real line. The recipe sits at high readiness and takes quarters because the model needs calibration to know when it is uncertain, the guardrails need real traffic to harden, and the fallback has to be proven to actually catch the failure cases under load. The naive version (call the model, use the answer) ships fast and breaks in production; the version with a real fallback is the one that lasts.

06Technique
INPUTSPROCESSACTIVATIONCustomer contextAction catalogGuardrail policyDeterministicfallback ruleInferredattributesLLM recommenderCandidate actionRule evaluator(hybrid)Signal & fallbackmonitoringFallback metricsFinal actionDeliver action

Work in this order. Guardrails define the allowed space, the model proposes within it, the fallback catches every failure, and monitoring watches the seam.

  1. Build the fallback independent of the model path. It is the precondition for using the model at all, and it must survive the model failing.
  2. Set the latency budget, with timeout going straight to the fallback.
  3. Calibrate confidence and set the fallback threshold.
  4. Evaluate the guardrail policy independently of the model. A probabilistic proposer eventually suggests something off-brand or non-compliant; rules decide what is permitted and win on conflict.
  5. Bound the proposals to the action catalog.
  6. Monitor the fallback rate against a baseline. It is the most skipped piece; without it the model can degrade for weeks while rules serve almost everything and the model gets the credit.
  7. Hold out a group to measure model lift over the deterministic baseline.
  8. Assess automated-decision exposure for consequential actions, where the regulatory exposure concentrates.

Real-time decisioning with deterministic fallback covers steps 1 to 3, rules-plus-model hybrid decisioning steps 4, 5 and 8, and signal quality monitoring steps 6 and 7. Inferred attribute generation supplies the context the model reasons over.

07Gotcha
Failure 01

The first failure is no real fallback. A system that calls the model and uses whatever comes back has no answer for the timeout, the outage, or the low-confidence case, so it breaks under exactly the conditions production guarantees will occur. The deterministic fallback is the precondition for using the model at all, and it has to be independent enough that the model failing does not take it down too.

Failure 02

The second is the unmonitored fallback rate. If nobody watches how often the deterministic path served, the model can degrade for weeks while the experience runs almost entirely on rules, and the team keeps crediting the model for outcomes the fallback produced. Monitoring the fallback rate against a baseline is what turns silent degradation into an alert, and it is the single most skipped piece.

Failure 03

The third is the guardrail gap. A probabilistic proposer will eventually suggest something off-brand, non-compliant, or simply wrong, and if the only thing between the suggestion and the customer is the model's own judgement, that suggestion ships. The guardrail rules that bound what is allowed, evaluated independently of the model, are what keep a confident-but-wrong proposal from reaching the customer, and where the action is consequential, that boundary is also where the regulatory exposure concentrates.

WorkshopFor your stack·The questions this recipe raises

Eight questions this recipe raises for your stack.

The Workshop works out with your team which of these matter for your stack right now, and what to do first: a 90-minute session with the people who own the decision.

  1. 01Fallback independence from the model pathWiring the deterministic rule so a model outage triggers it rather than taking it down.
  2. 02Latency budget and timeout-to-fallbackThe window after which the customer gets a rule instead of waiting on a failed call.
  3. 03Confidence calibration and the fallback thresholdTeaching the model to know when it is uncertain, and where the floor sits.
  4. 04Guardrail policy evaluated independently of the modelBounding what a probabilistic proposer may do before its suggestion reaches anyone.
  5. 05Action catalog as the bounded proposal spaceLetting the model select from known actions rather than invent one.
  6. 06Fallback-rate monitoring against baselineCatching the silent slide back to rules while the team still credits the model.
  7. 07Holdout to measure model lift over the deterministic baselineThe honest test of whether the model is worth its cost and risk.
  8. 08Automated-decision exposure for consequential actionsWhere the action crosses from cosmetic into a regulated decision line.

If you want model-driven next-best-action but cannot have the experience break when the model does, the Workshop is where we build the fallback-first version.

The deterministic rule that catches every model failure, the guardrails that bound what a probabilistic proposer may do, the latency budget that keeps the customer from waiting on a failed call, and the fallback-rate monitoring that catches silent degradation: those are the decisions that turn a fragile model call into a next-best-action capability you can run in production.

take this to the martech workshop→

Did this recipe match your situation?Anonymous response. Sign up to leave a longer note tied to your account.

Related recipes