tokenizers v1: encode, decode and scaling, measured
Hugging Face announced a release candidate for tokenizers v1, a performance-focused version of its tokenization library. The article states v1 will produce the same token IDs as v0.23 and preserve the API, vocabulary and merge ranks while loading everything v0.23 loaded.
The post reports that across the ten model families v1's encode path covers, it encodes text 3 to 30 times faster than v0.23 with one thread on an Apple M4 Max. The low end is t5-base and the high end gpt2, according to the article. It also states v1 scales at 76% of linear across eight workers.
The changes described include splitting one crate into a workspace with tk-encode, tk-serialize, tk-convert and tk-train; a no-allocation model using a caller-owned scratch buffer; bitcannon, which replaces regex splitting with bitstream operations using SIMD; a merge-loop rewrite using an intrusive doubly-linked list in a preallocated buffer; a thread-local word cache mapping pre-token bytes to finished IDs; and native parallelism where one shared tokenizer encodes from many threads, with each thread drawing from its own sub-pool (#2365).
Hugging Face credits open source work including gigatoken, tiktoken, kitoken, tokie, fastokens, wordchipper and ai-tokenizer. It thanks IBM, NVIDIA and the ExecuTorch team for patches and testing across hardware. The article says benchmarks come from the tokbench repository with a command to rerun them.
A release candidate is on crates.io. The post says the next priority is support for more model families before 1.0.0, with remaining work including one encoding implementation, optional offsets and masks, reworked normalizers, simpler Python bindings, and inference-only C and C++ bindings for ExecuTorch and llama.cpp. It notes possible JVM, Swift and Go bindings to follow and lists tok-devices GPU encoding and batch decoding as after-1.0.0 exploration.
Based on reporting from the original publisher. Visit the source for full context and later updates.