Uplifting conversion across the acquisition funnel with personalization using contextual bandits on AWS

Amazon Payments applied AI-based personalization to a product acquisition funnel using a multi-objective contextual multi-armed bandit on Amazon SageMaker AI, according to an AWS Machine Learning Blog post. In a seven-week online A/B test, one customer population showed a high single-digit percentage relative lift in final-funnel conversion, while another saw no improvement over the existing experience. The post attributes the second outcome to content rather than the model.
The funnel comprises three stages: application start, submission, and approval. One Linear UCB (LinUCB) model runs per stage, with their UCB scores combined through a weighted linear sum. The post describes a seesaw problem in which optimizing one stage can degrade another, and states that a single-stage policy would have risked gains upstream while hurting the downstream metric. Stage weights can follow business priorities or be learned by a separate calibration step; approximately equal weights were used.
Each customer is represented as a context vector of behavioral signals such as payment behavior and transaction mix, replacing fixed segments. The entity ID is used only to route recommendations back to visitors and is never a model input. Approval feedback lags by days, handled through an attribution window, while starts and submissions update the model immediately.
Content was assembled from industry-themed images and benefit-focused taglines; the arm space is the Cartesian product of these parts. Building blocks were vetted individually rather than reviewing every combination, and a design system kept variations consistent. The post notes that for the second population, broad exploration found no combination that beat the static page, and the approval regression was statistically significant.
The deployment uses a weekly SageMaker AI batch Processing job that reads outcomes from Amazon S3, updates the model, and writes per-customer recommendations to a low-latency key-value store such as Amazon DynamoDB. Models were warm-started from a period of randomized content assignment. Inference optimizations included precomputing arm matrix inversions and chunking the prospect set with Python multiprocessing. A fallback serves the default static experience when no recommendation exists. A code repository, quick start Jupyter notebook, command-line demo, and unit tests accompany the post.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Generative AI makes it cheap to produce personalized content at scale, but which variation do you show each customer? Amazon Payments used a multi-objective contextual bandit on Amazon SageMaker AI to personalize an acquisition funnel, achieving a high single-digit conversion lift for one audience, and learning why content, not the model, was the constraint.