
Yann LeCun advocates for JEPA (Joint Embedding Predictive Architecture) as a superior alternative to current Vision-Language-Action (VLA) models in AI, especially for robotics. JEPA models learn world dynamics through predictive embeddings, enabling explicit planning and better generalization without relying on human demonstrations. While JEPA shows promise, it currently lags behind VLA in performance but offers a compelling path forward with hierarchical planning and efficient learning.
The field of artificial intelligence is rapidly evolving, with various architectures competing to become the foundation for future AI systems. Yann LeCun, a pioneer in deep learning, has placed a significant bet on an alternative approach called JEPA (Joint Embedding Predictive Architecture), which he believes could eventually overtake the mainstream Vision-Language-Action (VLA) models, especially in robotics.
Physical Intelligence has developed some of the most impressive robot brains, such as their PI07 model, which can perform tasks like peeling a zucchini, folding a pinwheel, and taking out the trash. These robots use Vision-Language-Action (VLA) models, which combine vision encoders, large language models (LLMs), and action modules to control robots.
JEPA models also control robots but operate differently. Instead of directly generating outputs like VLA models, JEPA learns to predict embeddings of future states, enabling it to build a world model that can be used for planning and control. However, JEPA's demonstrated capabilities currently lag behind VLA models, such as taking 60 seconds to move a cup off a platform.
VLA models are built on Vision-Language Models (VLMs), which combine vision encoders and large language models. For example, OpenAI's CLIP model uses contrastive learning to align image and text embeddings, enabling multimodal understanding.
Meta's VJEPA 2, trained on 1 million hours of video with up to 1 billion parameters, uses a self-supervised approach to predict missing patches in videos. This trains the model to understand how the world works by filling in gaps, without relying on language supervision.
While CLIP aligns images with language captions, VJEPA 2 learns purely from visual data, allowing it to develop representations unconstrained by human language. Despite this, VJEPA 2 can still be aligned with language models and achieves state-of-the-art results on video understanding benchmarks.
JEPA can be applied not only to vision encoders but also to the entire vision-language model stack. Instead of generating text outputs directly, JEPA models predict embeddings of the output text, abstracting away irrelevant semantic details and improving learning efficiency.
Meta's VLJA architecture demonstrated that JEPA-based vision-language models learn significantly faster and achieve better performance on visual question answering benchmarks compared to traditional VLMs.
LeCun criticizes VLA models for two main reasons:
JEPA learns an action-conditioned world model that predicts future states given current states and actions. This enables explicit planning by simulating sequences of actions and their outcomes.
For example, in the "push T" task, a JEPA-trained model predicts how a T-shaped object moves when pushed by a robot's end-effector. The model can simulate future states and plan optimal action sequences using methods like the cross-entropy method.
Long-term planning is challenging due to prediction drift. LeCun proposes hierarchical models where low-level predictors make detailed short-term predictions, and higher-level models make abstract long-term predictions. This hierarchy allows planning over longer horizons efficiently.
While JEPA offers a compelling framework addressing key limitations of VLA models, its current performance in robotics tasks is limited compared to VLA. However, ongoing research aims to improve JEPA's capabilities, including hierarchical planning and applying the approach to complex industrial systems.
LeCun's team at Omni Labs plans to apply JEPA-based models to industrial applications within the next few years to gain practical experience and refine the methodology.
Yann LeCun's billion-dollar bet on JEPA represents a bold vision for the future of AI, emphasizing learning world models and explicit planning over end-to-end behavioral cloning. JEPA models like VJEPA 2 and VLJA demonstrate promising results, particularly in efficient learning and video understanding.
Although JEPA-based robotics systems currently lag behind VLA models in performance, their potential for better generalization, planning, and scalability makes them a fascinating area of research. The coming years will reveal whether JEPA can follow a trajectory similar to early deep learning systems and become a foundational AI architecture.
This exploration of JEPA versus VLA models highlights the evolving landscape of AI architectures and the ongoing quest to build more intelligent, reliable, and versatile robotic systems.
Paste a YouTube link and let Magica create the key takeaways.
Summarize another video