LeCun Continuously Endorses, New Work VISReg Tackles the Core Challenge of 'Representation Collapse' in JEPA World Models

marsbitPublicado a 2026-07-28Actualizado a 2026-07-28

Resumen

"VISReg: A New Self-Supervised Learning Method Tackles Representation Collapse in JEPA World Models" Yann LeCun has highlighted the new self-supervised learning (SSL) work VISReg (Variance-Invariance-Sketching Regularization), which addresses the core challenge of "representation collapse" in JEPA-based world models. SSL often collapses, mapping different inputs to similar vectors, losing discriminative power. While methods like VICReg and SIGReg attempted to regularize the representation distribution, they suffered from issues like vanishing gradients during collapse or coupled scale and shape optimization. VISReg overcomes these by decoupling the regularization into independent "scale" and "shape" objectives. It uses a variance term to prevent amplitude collapse (providing constant gradient even during collapse) and a sliced Wasserstein distance (SWD) based sketching target with stop-gradient to align the distribution shape to an isotropic Gaussian, without interfering with scale. This approach, requiring only ~15 lines of PyTorch core code, offers linear computational complexity and scales efficiently across multiple GPUs. Evaluated across 15 datasets (in-domain, out-of-distribution/OOD, dense prediction), VISReg outperforms 7 mainstream SSL methods (MoCoV3, DINO, iBOT, I-JEPA, MAE, data2vec) without relying on heuristic tricks like EMA or stop-gradient. Key results include: superior average OOD accuracy; matching DINOv2's OOD performance using only ~1/10 of the data (I...

The foundation of JEPA world models is Self-Supervised Learning (SSL), advocated by Yann LeCun since 2017.

SSL can learn general representations from massive amounts of data without manual annotation, but it commonly faces a core challenge—representation collapse: the model tends to map different inputs to the same or very few vectors, seemingly completing training without actually learning discriminative representations.

To suppress collapse, mainstream methods mostly rely on a series of heuristic techniques (EMA, teacher-student networks, stop-gradient, freezing layers, etc.). These techniques make training fragile, difficult to tune, and also weaken the method's interpretability and scalability.

Another approach is to directly constrain the representation distribution through regularization terms.

VICReg, proposed by LeCun's team, decomposes the learning objective into three terms: variance, invariance, and covariance, using covariance to constrain correlations between dimensions; but covariance only captures second-order statistics, unable to distinguish between two representations with "the same mean and variance, but drastically different distribution shapes."

Subsequently proposed SIGReg, based on the Cramér–Wold theorem, aligns the entire embedding distribution to a standard Gaussian using sketching techniques, thereby constraining the complete distribution shape.

However, SIGReg still has two key flaws:

  • Gradient Vanishing upon Collapse: When representations start to collapse, SIGReg's gradient decays accordingly—the more severe the collapse, the weaker the correction signal, making it difficult for the model to recover on its own;
  • Coupling of Scale and Shape: It does not separate the two independent attributes of "magnitude (scale)" and "distribution morphology (shape)". They interfere with each other during optimization, leading to poor adaptability on long-tailed, low-quality, low-rank data.

In other words, when the model most needs gradient signals to escape the collapsed state, SIGReg's gradient tends to vanish.

This is precisely the core problem that VISReg aims to solve.

Recently, the new SSL work VISReg (Variance-Invariance-Sketching Regularization) has been continuously endorsed and highly recognized by Turing Award winner Yann LeCun—he commented on the repost, "VICReg begat SIGReg which begat VISReg," succinctly outlining the technical lineage of this regularization route.

To gain such recognition from LeCun, what makes VISReg so strong?

The answer lies in the fact that it precisely targets the core challenge of JEPA world models that LeCun has long bet on—representation collapse.

Paper Link: https://arxiv.org/abs/2606.02572

Code / Pre-trained Weights: https://github.com/HaiyuWu/visreg

Project Page: https://haiyuwu.github.io/visreg/

VISReg decouples the collapse-preventing regularization term into two independent objectives: "scale" and "shape." Without relying on any heuristic training tricks or massive data, it outperforms 7 mainstream SSL methods on a comprehensive evaluation across 15 datasets; using only about 1/10 of the training data, it matches DINOv2 on out-of-distribution (OOD) benchmarks.

Figure 2: Simulation of gradient magnitude ‖∇L‖ for different regularization methods at various stages of representation collapse. VISReg maintains strong gradients even in collapsed states, while SIGReg's gradient almost vanishes.

Core Method

VISReg combines the strengths of VICReg and SIGReg: it retains VICReg's variance term to control scale, while replacing the covariance term with a sketching objective based on Sliced Wasserstein Distance (SWD) to control shape, and completely decouples the two via stop-gradient. The entire regularization objective consists of three parts.

1. Scale Regularization

The first part constrains the variance of each dimension to prevent amplitude collapse:

Its key property is: when the model collapses, the gradient of this term approaches a constant, ensuring the model can recover stably—this exactly compensates for SIGReg's gradient vanishing flaw.

2. Shape Regularization

The second part first normalizes to eliminate scale influence, then constrains shape separately. The crucial step is normalization with "stop-gradient" (sg):

Here, stop-gradient is applied to the standard deviation σ, ensuring that optimizing the shape loss does not conversely alter the scale—this is the mechanism enabling true decoupling and non-interference between the "scale" and "shape" objectives.

After normalization, the geometric shape of the distribution is aligned to an isotropic Gaussian using the Sliced Wasserstein Distance:

where

is the standard Gaussian quantile, and

is the random projection direction (i.e., "slice / sketching").

The theoretical basis is the Cramér–Wold theorem (Lemma 3.1 in the paper): two distributions are equal if and only if all their one-dimensional projections along directions on the unit sphere are equal. Therefore, by slicing the high-dimensional representation along sufficiently many random one-dimensional directions and aligning each to a Gaussian, we equivalently align the entire distribution in the high-dimensional space—this allows us to characterize the complete distribution shape using cheap one-dimensional sorting operations, rather than just second-order statistics.

3. Combined Objective

The third part is a centering loss that pulls the batch mean μ towards the origin:

The three regularization terms are combined with weights:

The predictive loss follows the invariance objective of JEPA / LeJEPA—aligning embeddings from each view (global + local, total V views)

to the mean of the global view

:

Finally, a single hyperparameter λ balances prediction and regularization, yielding the full objective:

Comparison with VICReg: VICReg also decouples regularization into variance + covariance, but covariance only captures second-order statistics; VISReg uses a sketching objective based on Sliced Wasserstein to fully characterize the distribution shape while retaining the variance term for scale control—preserving VICReg's flexibility while gaining distribution-level rigor.

Requires Only About 15 Lines of PyTorch Code

This regularization objective is very lightweight to implement; the core logic takes only about 15 lines:

Computational Complexity and Scalability

In terms of computation and scalability, VISReg also has advantages. The computational complexity of its regularization part is

(N is batch size, D is dimension, K is number of slices), which is linear for all scaling factors; in contrast, VICReg's covariance term is

, scaling quadratically with dimension.

Under the same batch size, VISReg's runtime speed and GPU memory usage on a single H100 GPU are superior to SIGReg.

More importantly, the K random slices can be distributed across multiple GPUs: each of M GPUs generates K/M slices, achieving an effect equivalent to a single GPU generating all K slices.

In experiments, when the number of slices per GPU was insufficient, switching to 8 GPUs, each with 128 slices (total 1024), reduced the accuracy gap with "single GPU 1024 slices" from about 2.4% to 0.22%. This means that K can remain constant when scaling up training, adding almost no per-GPU burden.

Figure: Change in linear probing accuracy when increasing GPU count with fixed K and D. When K is insufficient (K = 1⁄4D), using 8 times the number of GPUs can raise accuracy to the level of K = 2D—making it possible to keep K constant during large-scale training.

Experimental Results

Back to the question in the title—where exactly is VISReg strong? The research team compared VISReg with 7 mainstream SSL methods—MoCoV3, DINO, iBOT, I-JEPA, MAE, data2vec—on 15 datasets (8 in-domain + 6 out-of-distribution + ADE20K dense prediction), covering domains such as astronomy, medical, remote sensing, texture, flowers, etc. The answer manifests across multiple dimensions from recognition to segmentation and generation.

1. In-Domain Linear Probing

To ensure a fair comparison, experiments are divided into two groups based on whether heuristic techniques are used. In the group not using any heuristic techniques, VISReg leads: ViT-B/16 achieves in-domain linear probing accuracy of 75.7%, higher than MAE (75.1%); ViT-L/14 further improves to 77.0%, higher than LeJEPA (75.6%). Compared to iBOT and DINO which use heuristic techniques, VISReg is only slightly lower on conventional datasets, but surpasses all methods on the texture dataset DTD—indicating its cross-domain generalization ability stems from the method itself, not stacking manual tricks.

2. Out-of-Distribution (OOD) Generalization: Comprehensive Superiority

OOD generalization is a stricter test than in-domain accuracy: methods relying on heuristics are often finely tuned on the ImageNet in-domain but may not transfer well to significantly different new distributions. The team evaluated on 6 OOD datasets covering medical (ChestXRay, RetinaMNIST, OrganAMNIST), astronomy (Galaxy10), remote sensing (AID), texture (DTD), which are completely unrelated to the ImageNet training domain. Results show VISReg achieved the best average OOD accuracy across all methods and all backbone sizes, even surpassing some methods using heuristic techniques with larger backbones.

Figure 4: Average OOD linear probing accuracy. VISReg comprehensively outperforms iBOT, DINO, MoCoV3, I-JEPA, MAE, data2vec, etc.

As shown in Figure 4, ViT-B/16 VISReg average OOD accuracy is 70.19%, ViT-L/14 is 70.63%, significantly higher than MAE (67.85%), and better than MoCoV3 (69.46%), DINO (69.56%), I-JEPA (68.55%), etc.

3. Data Efficiency: Matching DINOv2 with 1/10 the Data

After pre-training VISReg (ViT-L/14) on ImageNet-22K (~14 million images), its average accuracy on 6 OOD datasets reaches 72.94%, essentially on par with DINOv2 (72.93%) trained on 10x larger scale LVD-142M (142 million images). In other words, VISReg achieves comparable performance using about 1/10 of the data. (As a control, VISReg with ViT-L/14 pre-trained only on ImageNet-1K has an average accuracy of 70.63%.) This indicates its learned representations have strong generality.

Figure 5: VISReg pre-trained on ImageNet-22K matches DINOv2 trained on 10x more data (LVD-142M) on OOD benchmarks.

4. Transfer Fine-tuning: Comprehensively Surpassing DINO

Although VISReg's linear probing accuracy on some in-domain datasets is slightly lower than DINO, after fine-tuning, it surpasses both DINO and supervised pre-training on all five tested datasets—CIFAR-10, CIFAR-100, Flowers, ImageNet-1K, Galaxy10—indicating its representation distribution is more uniform, less redundant, and more transferable.

Figure: Transfer learning comparison. After fine-tuning, VISReg outperforms DINO and supervised pre-training on all tested datasets (CIFAR-10, CIFAR-100, Flowers, ImageNet-1K, Galaxy10).

5. Dense Prediction and Generation Guidance

VISReg's advantages are not limited to classification. On ADE20K linear semantic segmentation (ViT-B/16), its mIoU is 30.16, higher than DINO (29.40) and MAE (23.60), second only to MoCoV3 (31.69); this result is competitive without using any heuristic tricks. The paper also acknowledges that there is still a gap with the best methods in dense prediction, a focus for future optimization.

Figure 7: Linear semantic segmentation on ADE20K. Without any heuristic tricks, VISReg achieves competitive mIoU, second only to MoCoV3.

In generation guidance (SiT-B/2, iREPA framework, 100k training steps), generation guided by VISReg features outperforms DINO on three out of four metrics: gFID 40.36 (DINO 41.15), Precision 51.38 (DINO 50.51), Recall 61.26 (DINO 60.70), IS essentially tied (33.48 vs 33.47). This shows VISReg's learned representations are also superior as generation guidance signals.

Figure 8: Image generation guided by VISReg vs. DINO features using SiT-B/2. VISReg provides better guidance on most metrics (lower gFID, higher Precision and Recall).

6. Robustness on Low-Quality Data

On low-quality datasets such as long-tail distribution (ImageNet-LT) and low-rank (Galaxy10), VISReg can stably prevent collapse and learn meaningful representations, while DINO fails directly without fine-tuned hyperparameters.

Table 1: Linear probing accuracy on ImageNet-LT

Table 2: In-domain linear probing accuracy on Galaxy10

Conclusion

VISReg demonstrates that by decoupling representation regularization into two independent components, "scale" and "shape," one can obtain an SSL method that is more stable, efficient, and has stronger generalization than existing methods.

Without using any training heuristic tricks, it achieves leading or near-leading results across multiple dimensions including image recognition, segmentation, and generation guidance, and matches DINOv2's OOD performance with about 1/10 of the data. This provides a new regularization-based solution to the long-standing representation collapse problem in JEPA world models.

References:

https://arxiv.org/abs/2606.02572

This article is from the WeChat public account "Xin Zhi Yuan," author: LRST

Criptos en tendencia

Preguntas relacionadas

QWhat is the core problem that VISReg aims to solve in self-supervised learning?

AVISReg aims to solve the core problem of representation collapse in self-supervised learning, where a model maps different inputs to the same or very few vectors, failing to learn discriminative representations.

QHow does VISReg fundamentally differ from its predecessor SIGReg in handling representation collapse?

AVISReg differs from SIGReg by decoupling the regularization into two independent targets: 'scale' and 'shape'. Crucially, it maintains a strong gradient signal even when the model is collapsing, whereas SIGReg suffers from vanishing gradients in such states, making recovery difficult.

QWhat are the three main components of the VISReg regularization objective?

AThe three main components are: 1) Scale Regularization (variance term to prevent magnitude collapse), 2) Shape Regularization (using Sliced Wasserstein Distance after stop-gradient normalization to align the distribution shape to an isotropic Gaussian), and 3) A centering loss that pulls the batch mean towards the origin.

QWhat key advantage in data efficiency does VISReg demonstrate compared to DINOv2 according to the experimental results?

AAccording to the experimental results, VISReg achieves comparable performance to DINOv2 on out-of-distribution benchmarks using only about 1/10th of the training data. Specifically, VISReg trained on ImageNet-22K (~14M images) matched the average OOD performance of DINOv2 trained on LVD-142M (~142M images).

QIn which evaluation scenarios did VISReg outperform DINO after fine-tuning?

AAfter fine-tuning, VISReg outperformed DINO and supervised pre-training across all five tested datasets: CIFAR-10, CIFAR-100, Flowers, ImageNet-1K, and Galaxy10.

Lecturas Relacionadas

¿Una transformación trascendental en la Fed? Informe: Warsh evalúa reducir la frecuencia de reuniones sobre tasas, rompiendo una práctica de 40 años

La presidenta de la Reserva Federal, Michelle Warsh, está considerando reducir la frecuencia de las reuniones periódicas del Comité Federal de Mercado Abierto (FOMC). Este cambio potencial, que rompería la práctica de "ocho reuniones al año, aproximadamente cada seis semanas" vigente desde 1981, representaría una de las transformaciones más significativas en las operaciones del banco central en décadas. Según el New York Times, la propuesta se discutió en una reunión esta semana, y un nuevo calendario podría acordarse antes de la próxima reunión de mediados de septiembre. La Ley Bancaria de 1935 exige al menos cuatro reuniones anuales del FOMC. El presidente Warsh ha solicitado comentarios a los funcionarios sobre la idea. La reducción del número de reuniones disminuiría las oportunidades de votar sobre las tasas de interés y podría afectar la capacidad de respuesta de la Fed a los cambios económicos, al tiempo que reduciría la transparencia al limitar las señales de política disponibles para el mercado. Esta iniciativa se enmarca en los esfuerzos de "reforma institucional" de Warsh, que ya han incluido declaraciones de política más breves y menos comentarios públicos sobre la economía. Históricamente, la frecuencia de las reuniones ha variado; en 1956 hubo 19 reuniones. Una evaluación interna de 1988 consideró que el esquema de ocho reuniones seguía siendo apropiado. La implementación de este cambio tendría un impacto profundo en los flujos de información del mercado, la flexibilidad de la política y la comunicación del banco central.

marsbitHace 58 min(s)

¿Una transformación trascendental en la Fed? Informe: Warsh evalúa reducir la frecuencia de reuniones sobre tasas, rompiendo una práctica de 40 años

marsbitHace 58 min(s)

Selecciones semanales del editor (Weekly Editor's Picks) (25-31 de julio)

**Resumen de la Selección del Editor Semanal (25-31 de julio)** La edición de esta semana destaca análisis profundos sobre incertidumbre financiera, tendencias tecnológicas y desarrollo regulatorio. La reunión de la Reserva Federal se perfila como una de las más inciertas en años, con el mercado dividido entre señales de desaceleración inflacionaria y posturas aún firmes contra la inflación persistente. En inversión y emprendimiento, se analiza la naturaleza de largo plazo del valor en criptomonedas, donde la persistencia supera al oportunismo. Se discute cómo los mercados globales de acciones, especialmente en tecnología, adoptan dinámicas similares a las de las criptomonedas, con narrativas y apalancamiento impulsando la volatilidad. Plataformas líderes como Hyperliquid y Polymarket enfrentan desafíos al expandirse más allá de sus mercados centrales, subrayando la importancia de la liquidez y los hábitos de usuario consolidados. En el ámbito de IA y almacenamiento, se examinan las crecientes preocupaciones sobre la deuda ligada a la expansión de infraestructura en la nube para IA y la reacción del mercado ante posibles excesos de oferta futuros en chips de memoria. El extraordinario rendimiento trimestral de SK Hynix no logró satisfacer las elevadas expectativas del mercado, reflejando la presión sobre los líderes del sector. Políticamente, la aprobación de la Ley CLARITY en EE.UU. se encuentra en una etapa crucial y políticamente compleja, con su destino aún incierto a pocos meses de las elecciones. Su posible fracaso podría tener un impacto limitado inmediato en el mercado pero complicaría futuros esfuerzos legislativos. Finalmente, en Ethereum, se anticipa un cambio estructural en el staking tras la actualización Pectra, que permitirá una gestión de capital más eficiente, aunque persiste una desconexión entre los fundamentos de la red y el precio de ETH. La semana también registró eventos de seguridad, como la compensación por liquidaciones anómalas en Trade.xyz, recordando los riesgos en la custodia de activos.

marsbitHace 1 hora(s)

Selecciones semanales del editor (Weekly Editor's Picks) (25-31 de julio)

marsbitHace 1 hora(s)

Trading

Spot

Artículos destacados

Cómo comprar CORE

¡Bienvenido a HTX.com! Hemos hecho que comprar CORE (CORE) sea simple y conveniente. Sigue nuestra guía paso a paso para iniciar tu viaje de criptos.Paso 1: crea tu cuenta HTXUtiliza tu correo electrónico o número de teléfono para registrarte y obtener una cuenta gratuita en HTX. Experimenta un proceso de registro sin complicaciones y desbloquea todas las funciones.Obtener mi cuentaPaso 2: ve a Comprar cripto y elige tu método de pagoTarjeta de crédito/débito: usa tu Visa o Mastercard para comprar CORE (CORE) al instante.Saldo: utiliza fondos del saldo de tu cuenta HTX para tradear sin problemas.Terceros: hemos agregado métodos de pago populares como Google Pay y Apple Pay para mejorar la comodidad.P2P: tradear directamente con otros usuarios en HTX.Over-the-Counter (OTC): ofrecemos servicios personalizados y tipos de cambio competitivos para los traders.Paso 3: guarda tu CORE (CORE)Después de comprar tu CORE (CORE), guárdalo en tu cuenta HTX. Alternativamente, puedes enviarlo a otro lugar mediante transferencia blockchain o utilizarlo para tradear otras criptomonedas.Paso 4: tradear CORE (CORE)Tradear fácilmente con CORE (CORE) en HTX's mercado spot. Simplemente accede a tu cuenta, selecciona tu par de trading, ejecuta tus trades y monitorea en tiempo real. Ofrecemos una experiencia fácil de usar tanto para principiantes como para traders experimentados.

440 Vistas totalesPublicado en 2024.12.13Actualizado en 2026.06.02

Cómo comprar CORE

Discusiones

Bienvenido a la comunidad de HTX. Aquí puedes mantenerte informado sobre los últimos desarrollos de la plataforma y acceder a análisis profesionales del mercado. A continuación se presentan las opiniones de los usuarios sobre el precio de CORE (CORE).

活动图片