Record, train, and deploy from one place with Strands Agents, LeRobot, and Hugging Face Storage Buckets
Hugging Face published a walkthrough describing a streaming data loop built around Strands Robots, LeRobot, and Hugging Face Storage Buckets, a mutable, non-versioned, Xet-backed object-storage repository type announced in March 2026.
The post describes a four-stage loop where a Robot("so100") records a LeRobotDataset, sync_dataset_to_bucket(...) syncs it into a Storage Bucket, stream_dataset(...) reads it back over the Hub without a full download, and a trained checkpoint deploys to the same Robot with mode="real". The on-disk format stays as LeRobot wrote it. The runnable companion is at examples/notebooks/05_streaming_data_loop.ipynb.
Strands Robots is described as an open source SDK from AWS (Apache 2.0) exposing robot abstractions, simulation, and the LeRobot stack as AgentTools composed into a Strands agent. The post states LeRobot's dataset format is used by over 90,000 datasets and models on the Hub from more than 8,000 publishers (LeRobot Project Pulse).
Storage Buckets are backed by Xet, which the post says deduplicates uploads at the byte level using content-defined chunking. It cites Hugging Face measurements showing content-defined chunking reduces data transferred per upload by about four times across the Hub, and bucket benchmarks where a 500 MB upload re-sent 5.5 MB after a 1% byte change, 27.5 MB after 5%, and 55 MB after 10%.
The training section says stream_dataset() reads batches via LeRobot's StreamingLeRobotDataset, decoding camera frames from remote MP4 shards on the fly, with only the small meta/ folder on local disk. It notes buckets are streaming-only, so --dataset.repo_type=bucket requires --dataset.streaming=true. A measured single configuration is cited: 500 optimizer steps of ACT (51.6M parameters, effective batch size 8) over a 120-frame episode completed in 133 seconds on a single NVIDIA L4 (g6.4xlarge), described as one measured configuration rather than a benchmark.
The post also cites a warm CDN bucket read of about 1,086 MB/s on a 10 GB payload versus 780 MB/s cold, and roughly 1,124 MB/s warm at 100 GB, measured on an m5dn.24xlarge in us-east-1. It states Storage Regions are a Team and Enterprise plan setting, currently US and EU, with Asia-Pacific and GCC regions announced as coming; outside those plans repositories are stored in the US.
Based on reporting from the original publisher. Visit the source for full context and later updates.