Rust-Powered Tokenizer Gigatoken Hits 24.53 GB/s, Outpaces HuggingFace by Over 1,000x
Summary
A new Rust-based tokenizer called Gigatoken is shattering speed records, achieving 24.53 GB/s and running over 1,000x faster than HuggingFace's tokenizers by leveraging SIMD instructions and aggressive caching, with the ability to tokenize all of Common Crawl in just 6.5 hours.
Key Points
- Gigatoken is a high-performance language model tokenizer written in Rust that achieves tokenization speeds of up to 24.53 GB/s, making it up to 1,353x faster than HuggingFace's tokenizers by leveraging SIMD instructions, aggressive caching of pretoken mappings, and minimized Python overhead.
- It supports nearly all commonly used tokenizers including GPT-2, Llama 3, Qwen, DeepSeek, and more, and can be used as a drop-in replacement for HuggingFace Tokenizers or Tiktoken via a simple compatibility mode.
- At peak speeds observed on a 144-core AMD EPYC processor, Gigatoken could theoretically tokenize the entirety of Common Crawl — roughly 130 trillion tokens — in just under 6.5 hours, and is installable via pip with straightforward Python API support.