D4RT: Teaching AI to see the world in four dimensions
Google DeepMind
Researchers introduced D4RT, an AI model that reconstructs dynamic 3D scenes from 2D video by tracking objects through space and time as a unified system rather than using separate specialized models. The model processes one-minute videos in roughly five seconds on a single TPU chip, performing 18x to 300x faster than previous methods that required up to ten minutes for the same task. This efficiency enables real-time applications in robotics and augmented reality that previously required too much computational power for practical use.
Why it matters
D4RT: Unified, efficient 4D reconstruction and tracking up to 300x faster than prior methods.