Skip to content

Extropic's Z1T Chip Delivers 100x Energy Efficiency Over GPUs With New Sparse Transformer Architecture

Sep 07, 2026
Extropic
Article image for Extropic's Z1T Chip Delivers 100x Energy Efficiency Over GPUs With New Sparse Transformer Architecture

Summary

Extropic's new Z1T chip delivers over 100x energy efficiency gains compared to Nvidia's H100 GPU by using sparse, in-memory computing and a disaggregated inference pipeline, consuming just 294.52 nJ per token versus 40.9 µJ on the H100, while also achieving faster per-token latency — with open-source weights and training recipes now publicly released.

Key Points

  • Extropic introduces Z1T, a family of sparse transformer-like models designed for its probabilistic Z1 chip, achieving over 100x energy efficiency gains compared to GPUs by leveraging sparse, in-memory computing instead of traditional dense matrix multiplications.
  • Z1T uses a disaggregated inference approach, splitting computations between Z1 probabilistic chips and FPGAs in a pipeline-parallel fashion, with energy estimates showing 294.52 nJ per token versus 40.9 µJ on an H100 GPU at 10% utilization, while also delivering faster per-token latency of approximately 58.8 µs compared to 102 µs for a compiled H100.
  • Extropic establishes empirical scaling laws for sparse transformer models on probabilistic hardware, finding that Z1T requires roughly 10x more FLOPs than a dense GPT-2 model to reach equivalent loss, but the 1000x energy advantage of Z1 operations still yields a net two-orders-of-magnitude efficiency gain, with open-source weights and training recipes now released.

Tags

Read Original Article