From Upstream Changes to Downstream Confidence: Inside Torch Spyre’s Integration with PyTorch CRCR

Torch Spyre, the PyTorch backend for the IBM Spyre Accelerator, has reached L2 integration with PyTorch's Cross-Repository CI Relay (CRCR), according to a post describing the mechanisms behind it. CRCR lets out-of-tree (OOT) accelerator repositories receive dispatches from pytorch/pytorch, report results back to PyTorch CI CRCR HUD, and incrementally onboard through four levels of integration toward non-blocking or blocking upstream PR validation. Integration needs an allowlist entry, a workflow that listens for repository_dispatch, and a composite callback action.
The post frames the engineering problem as three evolving candidates: the OOT backend, PyTorch core, and the test suite. The primary combination tested is the newest backend code against the newest PyTorch core and the test suite at dispatch time, which makes a regression attributable because the only thing that moved since the last green run is upstream. Other combinations, such as a range of backend versions against a moving core, can be reported to HUD as separate jobs.
To choose what to run from PyTorch's tens of thousands of tests, Torch Spyre built an agentic pipeline with four stages: high-level selection, a Repository Memory Generator that indexes symbols, files, summaries and per-test embeddings, low-level selection that buckets tests into mandatory_success or skip, and refinement using real-hardware execution logs to catch runtime failures and numerical differences. Each bucketing choice is recorded as a config comment. For a PyTorch 2.13 to 2.14 upgrade, the delta approach meant evaluating about 4,000 changed tests instead of tens of thousands.
Adaptation uses a declarative YAML framework with three controls: parameter-level knobs for dtypes and shapes; outcome buckets mandatory_success, xfail, xfail_strict and skip; and capability-driven inclusion via global supported_ops and supported_dtypes. The strict variant treats an unexpected pass as a signal to promote a test rather than silence. The framework patches upstream's @ops, @modules and @dtypes decorators at collection time and emits pytest marks such as op__ and dtype__, so the same config drives selection and adaptation; -m op__add runs exactly the tests exercising that operator. The upstream test tree stays unpatched.
Workflows split tests at feature, duration and balanced-bucketing levels, build PyTorch and backend wheels once for reuse across splits, and add retry patterns. For callbacks, a matched in_progress/completed pair must come from the same job because CRCR tracks retries by check_run_id, and each matrix leg gets its own.
Why it matters: The approach is not Spyre-specific by design, with the accelerator sitting behind the config rather than in the CI logic, and the authors say it applies to any generic privateuse1 device. The intended outcome is that upstreaming reusable parts, with PyTorch OpenReg as the reference point, spares other backends from reimplementing the same test-selection and workflow plumbing.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
TL;DR PyTorch’s Cross-Repository CI Relay (CRCR) gives out-of-tree accelerators a clean, scalable way to plug into upstream CI – and it leaves each backend free to decide which of PyTorch’s... The post From Upstream Changes to Downstream Confidence: Inside Torch Spyre’s Integration with PyTorch CRCR appeared first on PyTorch .