JD.com Launches JoyAI-Echo: Open-Source AI Platform Generating 5-Minute Audio-Visual Content and Interactive World Models
Summary
JD.com launches JoyAI-Echo, an open-source AI platform featuring two groundbreaking tools: Echo-LongVideo, capable of generating up to 5-minute multi-shot audio-visual content, and Echo-WM, an omnimodal world model that simultaneously evolves video, sound, music, and speech in response to real-time navigation inputs.
Key Points
- JoyAI-Echo is an open-source GitHub repository housing two independent AI projects: Echo-LongVideo, which generates long-horizon multi-shot audio-visual content up to approximately 5 minutes, and Echo-WM, an omnimodal world model that responds to continuous navigation while evolving video, sound, music, and speech together.
- Echo-WM is currently built on the LTX-2.3 backbone and has a roadmap to upgrade to LTX-2.5, with planned performance improvements including sparse attention, paged KV-cache, FlashAttention, and FP8/TensorRT compilation for faster and more efficient generation.
- The project, developed by JD.com and based on Lightricks' LTX-2 codebase, is available for academic and non-commercial use only, has garnered 1.9k stars on GitHub, and provides separate Python environments and checkpoints for each of its two sub-projects.