AivexaNewsSearch
AI news for builders and product teamsChecked every hour
PyTorchFirst partyProducts

How Shopify built a continual learning loop with PyTorch and vLLM

Collected Oct 1, 2026

A PyTorch case study details how Shopify built a continual learning loop with PyTorch and vLLM, compressing production failures into model weights daily, surpassing frontier-model quality on its task and cutting serving costs by 96%.

The loop begins with a quality rubric covering completeness, execution, response quality and safety, validated by blind annotation of 25 random samples and measured with Cohen's kappa; agreement around 0.2 indicates an ambiguous rubric. Ground truth includes randomly sampled traffic, not only curated examples. A judge is calibrated using DSPy with reflection-based optimizers GEPA and Agentic Context Engineering, then backtested against prior A/B tests and checked with targeted degradation tests.

Shopify then improves its frontier-powered baseline through autoresearch: an agent proposes changes to prompts, tool definitions or harness code, evaluates them against the judge and keeps only improvements. Once harness gains plateau, anonymized production traffic is mined for hard negatives. A panel of frontier reasoning models critiques failures, an arbiter merges critiques into a repair instruction injected before the user turn, and the conversation is replayed and rescored. Successful replays become reinforcement learning trajectories; failures go to Toloka expert annotators using the same rubric.

Training has two stages: supervised fine-tuning on complete trajectories including reasoning, then GRPO using the calibrated judge as reward. The self-healing pipeline runs daily with full-parameter fine-tuning, and training on new and previous trajectories limits drift and catastrophic forgetting. PyTorch distributes training across GPUs with tensor, context and data parallelism.

Gist compression replaces a roughly 6,000-token system prompt with about 1,500 learned gist tokens, with no measured quality loss on the judge. In a load test at 350 requests per minute, time-to-first-token dropped about 19% and end-to-end latency about 38%; throughput rose about 16% more requests per second and about 12% more output tokens per second on identical GPUs, roughly 14% fewer GPUs.

The GraphQL agent serves up to 2,000 requests per minute, answering merchant questions by writing and running queries against Shopify's Admin GraphQL API. The fine-tuned model's serving cost is estimated near $1M per year versus about $27M on a frontier model, a 96% reduction. PyTorch says the article was originally published on the Shopify Engineering Blog and adapted for the PyTorch community.

Read at PyTorch

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

TL;DR: This case study explores how Shopify compresses production failures into model weights every day, beats frontier-model quality, and cuts serving costs 96% by building a continual learning loop with...