Capturing token IDs during agentic interactions for better reinforcement learning
Amazon Science announced the release of Turnstile, a proxy written in the Rust programming language that sits between an agent harness and the backend system that runs a model. According to the announcement, Turnstile records the exact token-level history of every request at the moment of generation and exports a framework-neutral trajectory for reinforcement learning training stacks.
The stated problem is bookkeeping: a harness transcript faithfully records a conversation but not necessarily the tokens a model actually produced. The post describes retokenization drift, where rerunning a tokenizer over previously seen text can yield different token IDs, and chat template drift, where the surrounding format changes. It says policy-gradient RL works cleanly only when the trainer optimizes against the context the behavior policy actually saw.
Turnstile speaks the OpenAI Chat Completions API. Per the post, the harness points its client at Turnstile instead of the inference backend and otherwise runs unchanged. Requests flow through Turnstile to the inference backend, currently SGLang, with vLLM planned. Turnstile records sampled token IDs, per-token log probabilities, a loss mask, and which version of the model's weights was active during which spans. It returns TrainingSequence objects.
The post says Turnstile merges multiturn requests into a single growing token path when the previously captured token IDs appear unchanged at the start of a new request; when a prefix cannot be proven equivalent, it starts a new sequence, a process the post calls exploding the trajectory. Optional mixture-of-experts capture records routing traces and splits trajectories when shared-prefix routing does not match. For vision-language models, Turnstile decodes images, hashes and stores original bytes, runs the configured processor, and records processed pixel features alongside token IDs.
The post reports two validations: a text-only coding agent using OpenHands as the harness, training Qwen3-1.7B on the Mostly Basic Python Problems dataset, and a multimodal computer-use agent using OSWorld's PromptAgent harness, training Qwen3-VL-8B on OSWorld tasks, where mean reward per prompt rose from about 0.2 to about 0.71 over roughly 165 rollout steps. The post states both agents improved steadily and harnesses were left unchanged. Near-term work listed includes a vLLM backend, more training-framework adapters, and more multimodal-model coverage. Turnstile is available on Github. The post thanks Keagan Long, Daisy Lin, Changlong Yu, and Yifei Wang for contributions.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
A new Rust proxy called Turnstile sits between the model backend and the agent harness to capture information lost in mere text transcripts.