Dynamic LeJEPA: Maximum Entropy Representations for Sequential Prediction and Latent Planning with Theoretical Guarantees
Mohsen Mostafa
Source record
Source: Crossref
Published: Sep 14, 2026
DOI: 10.20944/preprints202609.0981.v1
Open original source ↗Source abstract
Joint-Embedding Predictive Architectures (JEPAs) are emerging as the backbone for latent world models in robotics and autonomous driving, yet inject-ing domain knowledge (physics, kinematics, geometry) into these models consistently degrades performance—with no theoretical explanation. We resolve this paradox by proving, through six theorems, that prediction and representation learning are strictly decoupled: sequential prediction losses do not alter the optimal maximum-entropy embedding distribution. This yields a world-model design principle with a provable guarantee: encoder maximum-entropy, predictor dynamics, decoder physics—and explains why prior physics-informed JEPAs failed: they constrained the encoder, where constraints provably destroy the entropy guarantee, instead of the decoder, where they are provably benign. We validate this principle across two domains through a phased protocol covering measurement artifacts (Proposition 1: the N/K ≥ 5 reliability threshold), powered statistical testing (60 seeded runs across autonomous driving and bimanual robotic manipulation, paired Wilcoxon significance in every com-parison), and a budget replication demonstrating that the encoder-physics violation deepens with training. Phase 3 additionally characterizes an entropy–utility frontier on low-intrinsic-dimension robotics data, discovering that per-dimension variance matching alone produces correlated collapse—a failure mode invisible on high-dimensional data—and that optimization budget, not regularization weight, is the binding entropy constraint. To test whether the placement rule matters beyond representation quality, we further probe its consequence for downstream control: a latent cross-entropy-method (CEM) planner built on the trained Phase 3 models consistently outperforms real-action replay on model-internal cost across all three abla-tion conditions—an impossible result for genuine planning, since replaying the true actions is itself the ground-truth solution. A model-exploitation diagnostic traces this to the planner discovering action sequences ∼26% of the action range away from the true trajectory while still lowering the learned decoder’s cost, uniformly regardless of physics placement. We report this transparently as a boundary condition on the practical claim: offline latent planning against a learned decoder is not, by itself, suffi-cient evidence that a representation supports downstream control, and we identify closed-loop, simulator-verified execution as the necessary next test. The result is the first formal, cross-domain-validated blueprint for physics-informed world models: place physics where the theorem says it is safe, never where intuition suggests—together with a concrete, reproducible cautionary result on how easily offline latent planning can be mistaken for evidence of control competence.
Evidence graph
No public relationships recorded yet.
Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.