Rust-Powered Tokenizer Gigatoken Hits 24.53 GB/s, Outpaces HuggingFace by Over 1,000x

Jul 22, 2026
GitHub
Article image for Rust-Powered Tokenizer Gigatoken Hits 24.53 GB/s, Outpaces HuggingFace by Over 1,000x

Summary

A new Rust-based tokenizer called Gigatoken is shattering speed records, achieving 24.53 GB/s and running over 1,000x faster than HuggingFace's tokenizers by leveraging SIMD instructions and aggressive caching, with the ability to tokenize all of Common Crawl in just 6.5 hours.

Key Points

  • Gigatoken is a high-performance language model tokenizer written in Rust that achieves tokenization speeds of up to 24.53 GB/s, making it up to 1,353x faster than HuggingFace's tokenizers by leveraging SIMD instructions, aggressive caching of pretoken mappings, and minimized Python overhead.
  • It supports nearly all commonly used tokenizers including GPT-2, Llama 3, Qwen, DeepSeek, and more, and can be used as a drop-in replacement for HuggingFace Tokenizers or Tiktoken via a simple compatibility mode.
  • At peak speeds observed on a 144-core AMD EPYC processor, Gigatoken could theoretically tokenize the entirety of Common Crawl — roughly 130 trillion tokens — in just under 6.5 hours, and is installable via pip with straightforward Python API support.

Tags

Read Original Article