AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Faster Scientific Image Analysis with NVIDIA cuPhoton

Collected Oct 7, 2026

NVIDIA introduced cuPhoton, an open-source CUDA-X toolkit of GPU-accelerated building blocks for scientific image analysis, spanning spectral and optical astronomy to time-domain laser and X-ray analysis.

The toolkit includes modules named xDataReader for direct FITS loading, xRep for image alignment, xPois for point-spread-function matching and subtraction, xFit for dipole fitting, xScan for candidate review, and xRay for time-domain X-ray detector analysis. cuPhoton 0.1.3 supports Python 3.12–3.14 and CUDA 13 on Linux.

NVIDIA says cuPhoton keeps image data on the GPU from sensor read through classification. On representative workloads involving hundreds of terabytes of data using multiple GPUs, the company reports image loading and reading accelerated by up to 14,900x and signal processing by up to 14,550x compared to an x86 CPU baseline. NVIDIA states these are individual operation speedups, not an end-to-end pipeline speedup.

NVIDIA cites the NSF-DOE Vera C. Rubin Observatory as a use case. Its LSSTCam records a 3.2-gigapixel exposure every 39 seconds, generating up to 20 terabytes of images and 10 million candidate objects per night; the Prompt Processing pipeline must classify roughly 10,000 detections within about 60 to 120 seconds. NVIDIA says data analytics that previously took nine months have been demonstrated in four hours on GPU-accelerated Python.

cuPhoton scales across multi-GPU, multi-node NVIDIA Grace Blackwell and NVIDIA Vera Rubin systems, distributing image pairs and candidate coordinates across GPU workers while keeping arrays on the device. The toolkit's xDataReader accelerates FITS decompression, reading, and loading to the GPU using CFITSIO, KvikIO, nvCOMP, and CuPy; it supports uncompressed 2–16-axis image HDUs and two-dimensional images with GZIP_1 or GZIP_2 tile compression, but not Rice compression or dithered floating-point quantization. Setup uses a locked CUDA 13 environment; building xDataReader from source requires a C++17 compiler, CUDA and cuFile development headers, and a reentrant CFITSIO development installation.

Read at NVIDIA Developer Blog

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Observatories and telescopes, lasers and X-ray light sources, and other high-throughput instruments generate image data faster than CPU-bound pipelines can...