AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Introducing preemptible compute: the same compute, half the price

Collected Oct 1, 2026

Together AI announced the public preview of preemptible compute for Together GPU Clusters, available on Kubernetes clusters in all regions. Preemptible nodes are billed sub-hourly at a flat 50% of the on-demand rate, with usage metered every one to two minutes, so a node running for 12 minutes is billed for approximately 12 minutes. The rate is fixed rather than moving with a spot market.

Preemptible compute adds a second compute type to Together GPU Clusters. Standard nodes are fulfilled synchronously and are never preempted. Preemptible nodes use the same NVIDIA accelerated compute, draw from unused capacity, and can be reclaimed when that capacity is needed elsewhere.

When a node is reclaimed, the cluster follows a drain sequence of up to five minutes. At T+0 the node is cordoned, a TogetherPreempted Kubernetes event fires, and pods receive SIGTERM. Workloads have up to five minutes (terminationGracePeriodSeconds) to checkpoint and exit, after which the node is removed at T+5:00. The cluster retains its preemptible target and automatically refills toward it as capacity becomes available, without requiring a request for replacement capacity.

Preemptible nodes join existing clusters labeled together.ai/compute-class=preemptible, rather than forming a separate cluster type. Together AI states that critical components should remain on standard nodes, and that work which can resume, retry, or requeue is suited to preemptible capacity, while multi-day runs without checkpointing and strict-SLO serving with no fallback are not a fit.

Setup involves setting a preemptible target when creating or updating a cluster via the Together Cloud console, CLI, or API, and scheduling eligible work onto labeled nodes. The API reports requested target as desired_preemptible_gpus and live capacity as allocated_preemptible_gpus; allocated capacity can remain below the desired target when capacity is tight. Clusters require at least one standard node, and nodes cannot be converted between compute types in place. Slurm support, additional regions, and in-place conversion are planned next.

Read at Together AI

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Together GPU Clusters now supports preemptible compute: the same GPU capacity at a flat 50% of the on-demand rate, with a five-minute drain window.