Let a model propose the next action when it can, and serve a known-good rule when it cannot, so the experience never depends on the model being up
A model that picks the next best action for each customer is an appealing idea and a fragile one if built naively. Language and recommendation models are probabilistic, occasionally slow, occasionally unavailable, and occasionally confidently wrong, and a customer experience that calls the model on every decision inherits all of that. The first time the model times out during a traffic spike, or returns a brand-unsafe suggestion, or simply goes down, the experience that depended on it breaks, and it breaks at exactly the moment volume is highest.
The pattern that makes model-driven decisioning shippable is not better models; it is the deterministic fallback. The model proposes the next action when it can, within latency, with sufficient confidence, and inside the brand's guardrails. When any of those conditions fails, a deterministic rule serves a known-good action instead. The fallback is not an error path bolted on at the end; it is the design that lets the model be used in production at all, because it converts every model failure from a broken experience into a slightly less personalized but still correct one.
This recipe is model-proposes, rule-catches, for the next-best-action decision.
A next-best-action capability that degrades gracefully rather than breaking, with the model improving the decision when it can and a deterministic rule guaranteeing a correct decision when it cannot. The metric is next-best-action acceptance rate (do customers respond to the proposed action), but the operational health metric that matters as much is the fallback rate: how often the deterministic path served instead of the model. A low, stable fallback rate means the model is genuinely contributing; a high or rising one means the experience is mostly running on rules while the team believes it is running on a model. The lift from the model over the deterministic baseline is the honest measure of whether the model is worth its cost and risk, and it should be measured rather than assumed.
The second outcome is resilience: the experience holds during the model outage, the latency spike, and the low-confidence case, because the fallback is the default the whole thing degrades to.
A real-time decisioning path that calls the model within a strict latency budget, falls back to a deterministic rule on timeout, low confidence, or guardrail violation, and monitors which path served. Composable stacks suit this because the action catalog, the fallback, and the guardrails are things the team must own and inspect, and the context the model reasons over lives in the warehouse or context layer. The capability that decides whether this is production-grade is the fallback wiring: it has to be genuinely independent of the model, so that a model outage triggers the rule rather than taking the whole decision path down with it, and it has to be fast enough that the customer never waits on a failed model call.
Compare the tools on Martech Stack Builder
Data science owns the model that proposes actions, the confidence calibration, and the threshold below which the system falls back. Marketing ops owns the action catalog, the deterministic fallback rule, and the brand-safety guardrails, which are the parts that keep the model inside acceptable bounds. Data engineering owns the request-time path, the latency budget, the fallback wiring, and the monitoring. Legal covers the automated-decision exposure where the next action is consequential to the customer rather than cosmetic, which under the AI Act and adjacent regimes is a real line. The recipe sits at high readiness and takes quarters because the model needs calibration to know when it is uncertain, the guardrails need real traffic to harden, and the fallback has to be proven to actually catch the failure cases under load. The naive version (call the model, use the answer) ships fast and breaks in production; the version with a real fallback is the one that lasts.
Work in this order. Guardrails define the allowed space, the model proposes within it, the fallback catches every failure, and monitoring watches the seam.
Real-time decisioning with deterministic fallback covers steps 1 to 3, rules-plus-model hybrid decisioning steps 4, 5 and 8, and signal quality monitoring steps 6 and 7. Inferred attribute generation supplies the context the model reasons over.
The first failure is no real fallback. A system that calls the model and uses whatever comes back has no answer for the timeout, the outage, or the low-confidence case, so it breaks under exactly the conditions production guarantees will occur. The deterministic fallback is the precondition for using the model at all, and it has to be independent enough that the model failing does not take it down too.
The second is the unmonitored fallback rate. If nobody watches how often the deterministic path served, the model can degrade for weeks while the experience runs almost entirely on rules, and the team keeps crediting the model for outcomes the fallback produced. Monitoring the fallback rate against a baseline is what turns silent degradation into an alert, and it is the single most skipped piece.
The third is the guardrail gap. A probabilistic proposer will eventually suggest something off-brand, non-compliant, or simply wrong, and if the only thing between the suggestion and the customer is the model's own judgement, that suggestion ships. The guardrail rules that bound what is allowed, evaluated independently of the model, are what keep a confident-but-wrong proposal from reaching the customer, and where the action is consequential, that boundary is also where the regulatory exposure concentrates.
The Workshop works out with your team which of these matter for your stack right now, and what to do first: a 90-minute session with the people who own the decision.
The deterministic rule that catches every model failure, the guardrails that bound what a probabilistic proposer may do, the latency budget that keeps the customer from waiting on a failed call, and the fallback-rate monitoring that catches silent degradation: those are the decisions that turn a fragile model call into a next-best-action capability you can run in production.
Did this recipe match your situation?Anonymous response. Sign up to leave a longer note tied to your account.
Show account-relevant content to known and reverse-IP-identified B2B visitors, with explicit handling of the wrong-match case
Show returning visitors what they had going, last-viewed products, cart contents, loyalty progress, without making them sign in or start over
Give the AI a memory of each customer that is curated and decaying, not a raw event firehose it cannot use and should not keep