Moondream's Photon Engine Cuts GPU Idle Time, Boosts VLM Decode Throughput by Up to 35%
Moondream's new Photon inference engine eliminates GPU idle time through pipelined decoding, boosting vision-language model throughput by up to 35% on NVIDIA B200 hardware and achieving near-realtime inference at approximately 33ms.