LeCun Continuously Endorses, New Work VISReg Tackles the Core Challenge of 'Representation Collapse' in JEPA World Models

marsbitPublicado a 2026-07-28Actualizado a 2026-07-28

Resumen

"VISReg: A New Self-Supervised Learning Method Tackles Representation Collapse in JEPA World Models" Yann LeCun has highlighted the new self-supervised learning (SSL) work VISReg (Variance-Invariance-Sketching Regularization), which addresses the core challenge of "representation collapse" in JEPA-based world models. SSL often collapses, mapping different inputs to similar vectors, losing discriminative power. While methods like VICReg and SIGReg attempted to regularize the representation distribution, they suffered from issues like vanishing gradients during collapse or coupled scale and shape optimization. VISReg overcomes these by decoupling the regularization into independent "scale" and "shape" objectives. It uses a variance term to prevent amplitude collapse (providing constant gradient even during collapse) and a sliced Wasserstein distance (SWD) based sketching target with stop-gradient to align the distribution shape to an isotropic Gaussian, without interfering with scale. This approach, requiring only ~15 lines of PyTorch core code, offers linear computational complexity and scales efficiently across multiple GPUs. Evaluated across 15 datasets (in-domain, out-of-distribution/OOD, dense prediction), VISReg outperforms 7 mainstream SSL methods (MoCoV3, DINO, iBOT, I-JEPA, MAE, data2vec) without relying on heuristic tricks like EMA or stop-gradient. Key results include: superior average OOD accuracy; matching DINOv2's OOD performance using only ~1/10 of the data (I...

The foundation of JEPA world models is Self-Supervised Learning (SSL), advocated by Yann LeCun since 2017.

SSL can learn general representations from massive amounts of data without manual annotation, but it commonly faces a core challenge—representation collapse: the model tends to map different inputs to the same or very few vectors, seemingly completing training without actually learning discriminative representations.

To suppress collapse, mainstream methods mostly rely on a series of heuristic techniques (EMA, teacher-student networks, stop-gradient, freezing layers, etc.). These techniques make training fragile, difficult to tune, and also weaken the method's interpretability and scalability.

Another approach is to directly constrain the representation distribution through regularization terms.

VICReg, proposed by LeCun's team, decomposes the learning objective into three terms: variance, invariance, and covariance, using covariance to constrain correlations between dimensions; but covariance only captures second-order statistics, unable to distinguish between two representations with "the same mean and variance, but drastically different distribution shapes."

Subsequently proposed SIGReg, based on the Cramér–Wold theorem, aligns the entire embedding distribution to a standard Gaussian using sketching techniques, thereby constraining the complete distribution shape.

However, SIGReg still has two key flaws:

  • Gradient Vanishing upon Collapse: When representations start to collapse, SIGReg's gradient decays accordingly—the more severe the collapse, the weaker the correction signal, making it difficult for the model to recover on its own;
  • Coupling of Scale and Shape: It does not separate the two independent attributes of "magnitude (scale)" and "distribution morphology (shape)". They interfere with each other during optimization, leading to poor adaptability on long-tailed, low-quality, low-rank data.

In other words, when the model most needs gradient signals to escape the collapsed state, SIGReg's gradient tends to vanish.

This is precisely the core problem that VISReg aims to solve.

Recently, the new SSL work VISReg (Variance-Invariance-Sketching Regularization) has been continuously endorsed and highly recognized by Turing Award winner Yann LeCun—he commented on the repost, "VICReg begat SIGReg which begat VISReg," succinctly outlining the technical lineage of this regularization route.

To gain such recognition from LeCun, what makes VISReg so strong?

The answer lies in the fact that it precisely targets the core challenge of JEPA world models that LeCun has long bet on—representation collapse.

Paper Link: https://arxiv.org/abs/2606.02572

Code / Pre-trained Weights: https://github.com/HaiyuWu/visreg

Project Page: https://haiyuwu.github.io/visreg/

VISReg decouples the collapse-preventing regularization term into two independent objectives: "scale" and "shape." Without relying on any heuristic training tricks or massive data, it outperforms 7 mainstream SSL methods on a comprehensive evaluation across 15 datasets; using only about 1/10 of the training data, it matches DINOv2 on out-of-distribution (OOD) benchmarks.

Figure 2: Simulation of gradient magnitude ‖∇L‖ for different regularization methods at various stages of representation collapse. VISReg maintains strong gradients even in collapsed states, while SIGReg's gradient almost vanishes.

Core Method

VISReg combines the strengths of VICReg and SIGReg: it retains VICReg's variance term to control scale, while replacing the covariance term with a sketching objective based on Sliced Wasserstein Distance (SWD) to control shape, and completely decouples the two via stop-gradient. The entire regularization objective consists of three parts.

1. Scale Regularization

The first part constrains the variance of each dimension to prevent amplitude collapse:

Its key property is: when the model collapses, the gradient of this term approaches a constant, ensuring the model can recover stably—this exactly compensates for SIGReg's gradient vanishing flaw.

2. Shape Regularization

The second part first normalizes to eliminate scale influence, then constrains shape separately. The crucial step is normalization with "stop-gradient" (sg):

Here, stop-gradient is applied to the standard deviation σ, ensuring that optimizing the shape loss does not conversely alter the scale—this is the mechanism enabling true decoupling and non-interference between the "scale" and "shape" objectives.

After normalization, the geometric shape of the distribution is aligned to an isotropic Gaussian using the Sliced Wasserstein Distance:

where

is the standard Gaussian quantile, and

is the random projection direction (i.e., "slice / sketching").

The theoretical basis is the Cramér–Wold theorem (Lemma 3.1 in the paper): two distributions are equal if and only if all their one-dimensional projections along directions on the unit sphere are equal. Therefore, by slicing the high-dimensional representation along sufficiently many random one-dimensional directions and aligning each to a Gaussian, we equivalently align the entire distribution in the high-dimensional space—this allows us to characterize the complete distribution shape using cheap one-dimensional sorting operations, rather than just second-order statistics.

3. Combined Objective

The third part is a centering loss that pulls the batch mean μ towards the origin:

The three regularization terms are combined with weights:

The predictive loss follows the invariance objective of JEPA / LeJEPA—aligning embeddings from each view (global + local, total V views)

to the mean of the global view

:

Finally, a single hyperparameter λ balances prediction and regularization, yielding the full objective:

Comparison with VICReg: VICReg also decouples regularization into variance + covariance, but covariance only captures second-order statistics; VISReg uses a sketching objective based on Sliced Wasserstein to fully characterize the distribution shape while retaining the variance term for scale control—preserving VICReg's flexibility while gaining distribution-level rigor.

Requires Only About 15 Lines of PyTorch Code

This regularization objective is very lightweight to implement; the core logic takes only about 15 lines:

Computational Complexity and Scalability

In terms of computation and scalability, VISReg also has advantages. The computational complexity of its regularization part is

(N is batch size, D is dimension, K is number of slices), which is linear for all scaling factors; in contrast, VICReg's covariance term is

, scaling quadratically with dimension.

Under the same batch size, VISReg's runtime speed and GPU memory usage on a single H100 GPU are superior to SIGReg.

More importantly, the K random slices can be distributed across multiple GPUs: each of M GPUs generates K/M slices, achieving an effect equivalent to a single GPU generating all K slices.

In experiments, when the number of slices per GPU was insufficient, switching to 8 GPUs, each with 128 slices (total 1024), reduced the accuracy gap with "single GPU 1024 slices" from about 2.4% to 0.22%. This means that K can remain constant when scaling up training, adding almost no per-GPU burden.

Figure: Change in linear probing accuracy when increasing GPU count with fixed K and D. When K is insufficient (K = 1⁄4D), using 8 times the number of GPUs can raise accuracy to the level of K = 2D—making it possible to keep K constant during large-scale training.

Experimental Results

Back to the question in the title—where exactly is VISReg strong? The research team compared VISReg with 7 mainstream SSL methods—MoCoV3, DINO, iBOT, I-JEPA, MAE, data2vec—on 15 datasets (8 in-domain + 6 out-of-distribution + ADE20K dense prediction), covering domains such as astronomy, medical, remote sensing, texture, flowers, etc. The answer manifests across multiple dimensions from recognition to segmentation and generation.

1. In-Domain Linear Probing

To ensure a fair comparison, experiments are divided into two groups based on whether heuristic techniques are used. In the group not using any heuristic techniques, VISReg leads: ViT-B/16 achieves in-domain linear probing accuracy of 75.7%, higher than MAE (75.1%); ViT-L/14 further improves to 77.0%, higher than LeJEPA (75.6%). Compared to iBOT and DINO which use heuristic techniques, VISReg is only slightly lower on conventional datasets, but surpasses all methods on the texture dataset DTD—indicating its cross-domain generalization ability stems from the method itself, not stacking manual tricks.

2. Out-of-Distribution (OOD) Generalization: Comprehensive Superiority

OOD generalization is a stricter test than in-domain accuracy: methods relying on heuristics are often finely tuned on the ImageNet in-domain but may not transfer well to significantly different new distributions. The team evaluated on 6 OOD datasets covering medical (ChestXRay, RetinaMNIST, OrganAMNIST), astronomy (Galaxy10), remote sensing (AID), texture (DTD), which are completely unrelated to the ImageNet training domain. Results show VISReg achieved the best average OOD accuracy across all methods and all backbone sizes, even surpassing some methods using heuristic techniques with larger backbones.

Figure 4: Average OOD linear probing accuracy. VISReg comprehensively outperforms iBOT, DINO, MoCoV3, I-JEPA, MAE, data2vec, etc.

As shown in Figure 4, ViT-B/16 VISReg average OOD accuracy is 70.19%, ViT-L/14 is 70.63%, significantly higher than MAE (67.85%), and better than MoCoV3 (69.46%), DINO (69.56%), I-JEPA (68.55%), etc.

3. Data Efficiency: Matching DINOv2 with 1/10 the Data

After pre-training VISReg (ViT-L/14) on ImageNet-22K (~14 million images), its average accuracy on 6 OOD datasets reaches 72.94%, essentially on par with DINOv2 (72.93%) trained on 10x larger scale LVD-142M (142 million images). In other words, VISReg achieves comparable performance using about 1/10 of the data. (As a control, VISReg with ViT-L/14 pre-trained only on ImageNet-1K has an average accuracy of 70.63%.) This indicates its learned representations have strong generality.

Figure 5: VISReg pre-trained on ImageNet-22K matches DINOv2 trained on 10x more data (LVD-142M) on OOD benchmarks.

4. Transfer Fine-tuning: Comprehensively Surpassing DINO

Although VISReg's linear probing accuracy on some in-domain datasets is slightly lower than DINO, after fine-tuning, it surpasses both DINO and supervised pre-training on all five tested datasets—CIFAR-10, CIFAR-100, Flowers, ImageNet-1K, Galaxy10—indicating its representation distribution is more uniform, less redundant, and more transferable.

Figure: Transfer learning comparison. After fine-tuning, VISReg outperforms DINO and supervised pre-training on all tested datasets (CIFAR-10, CIFAR-100, Flowers, ImageNet-1K, Galaxy10).

5. Dense Prediction and Generation Guidance

VISReg's advantages are not limited to classification. On ADE20K linear semantic segmentation (ViT-B/16), its mIoU is 30.16, higher than DINO (29.40) and MAE (23.60), second only to MoCoV3 (31.69); this result is competitive without using any heuristic tricks. The paper also acknowledges that there is still a gap with the best methods in dense prediction, a focus for future optimization.

Figure 7: Linear semantic segmentation on ADE20K. Without any heuristic tricks, VISReg achieves competitive mIoU, second only to MoCoV3.

In generation guidance (SiT-B/2, iREPA framework, 100k training steps), generation guided by VISReg features outperforms DINO on three out of four metrics: gFID 40.36 (DINO 41.15), Precision 51.38 (DINO 50.51), Recall 61.26 (DINO 60.70), IS essentially tied (33.48 vs 33.47). This shows VISReg's learned representations are also superior as generation guidance signals.

Figure 8: Image generation guided by VISReg vs. DINO features using SiT-B/2. VISReg provides better guidance on most metrics (lower gFID, higher Precision and Recall).

6. Robustness on Low-Quality Data

On low-quality datasets such as long-tail distribution (ImageNet-LT) and low-rank (Galaxy10), VISReg can stably prevent collapse and learn meaningful representations, while DINO fails directly without fine-tuned hyperparameters.

Table 1: Linear probing accuracy on ImageNet-LT

Table 2: In-domain linear probing accuracy on Galaxy10

Conclusion

VISReg demonstrates that by decoupling representation regularization into two independent components, "scale" and "shape," one can obtain an SSL method that is more stable, efficient, and has stronger generalization than existing methods.

Without using any training heuristic tricks, it achieves leading or near-leading results across multiple dimensions including image recognition, segmentation, and generation guidance, and matches DINOv2's OOD performance with about 1/10 of the data. This provides a new regularization-based solution to the long-standing representation collapse problem in JEPA world models.

References:

https://arxiv.org/abs/2606.02572

This article is from the WeChat public account "Xin Zhi Yuan," author: LRST

Criptos en tendencia

Preguntas relacionadas

QWhat is the core problem that VISReg aims to solve in self-supervised learning?

AVISReg aims to solve the core problem of representation collapse in self-supervised learning, where a model maps different inputs to the same or very few vectors, failing to learn discriminative representations.

QHow does VISReg fundamentally differ from its predecessor SIGReg in handling representation collapse?

AVISReg differs from SIGReg by decoupling the regularization into two independent targets: 'scale' and 'shape'. Crucially, it maintains a strong gradient signal even when the model is collapsing, whereas SIGReg suffers from vanishing gradients in such states, making recovery difficult.

QWhat are the three main components of the VISReg regularization objective?

AThe three main components are: 1) Scale Regularization (variance term to prevent magnitude collapse), 2) Shape Regularization (using Sliced Wasserstein Distance after stop-gradient normalization to align the distribution shape to an isotropic Gaussian), and 3) A centering loss that pulls the batch mean towards the origin.

QWhat key advantage in data efficiency does VISReg demonstrate compared to DINOv2 according to the experimental results?

AAccording to the experimental results, VISReg achieves comparable performance to DINOv2 on out-of-distribution benchmarks using only about 1/10th of the training data. Specifically, VISReg trained on ImageNet-22K (~14M images) matched the average OOD performance of DINOv2 trained on LVD-142M (~142M images).

QIn which evaluation scenarios did VISReg outperform DINO after fine-tuning?

AAfter fine-tuning, VISReg outperformed DINO and supervised pre-training across all five tested datasets: CIFAR-10, CIFAR-100, Flowers, ImageNet-1K, and Galaxy10.

Lecturas Relacionadas

Qualcomm Chip Price Hike Deals a Blow to Android Phones

Qualcomm has officially announced a new round of price increases for its entire chip portfolio, effective September 1. The hikes, reaching up to 18% for the flagship Snapdragon 8 Elite Gen 6 Pro, follow similar moves by MediaTek, intensifying cost pressures on the already strained smartphone industry. CEO Cristiano Amon confirmed the plan, citing the need to offset rising industry-wide costs and restore declining profit margins. Qualcomm's Q3 FY2026 results showed a 25% drop in net profit, with mobile revenue plunging 20% year-over-year, hitting its lowest level since 2021. The surge in AI computing demand has led memory manufacturers to prioritize HBM production, creating a shortage in general-purpose DRAM and NAND Flash chips. Their prices soared by 93%-98% and 55%-60% respectively in Q1 2026, causing the memory cost share in smartphones to jump from 10%-15% to over 30%. Coupled with the soaring cost of advanced nodes like TSMC's 2nm and packaging, overall chip costs have reached historic highs. These upstream pressures are forcing downstream smartphone brands like Xiaomi, OPPO, and vivo to cut orders for mid-to-low-end models by up to 20% and use cost-saving measures like older chipsets. Reportedly, the Snapdragon 8E5 will be repurposed as a "long-lasting" chip for sub-brand phones in H2 2026. Amid this cost crisis, the Android market remains sluggish. Q2 2026 smartphone shipments in China fell 4.3% year-over-year, marking five consecutive quarters of decline. Major Android brands saw market share drop, while Huawei and Apple, with their in-house chip advantages, gained share. Qualcomm is diversifying into automotive and IoT sectors to reduce reliance on smartphones, but these new segments cannot yet fill the mobile revenue gap. Industry observers warn that the full impact of component cost hikes will hit in the second half of 2026, likely leading to higher-than-expected price increases for Android flagships and a further contraction in the Chinese smartphone market.

marsbitHace 4 min(s)

Qualcomm Chip Price Hike Deals a Blow to Android Phones

marsbitHace 4 min(s)

Breaking: GPT-5.6 Prices Slashed Effective Today

OpenAI has announced significant price cuts for its GPT-5.6 model API, effective immediately. The entry-level **GPT-5.6 Luna** sees the most drastic reduction, with input prices dropping 80% to $0.20 per million tokens and output prices falling to $1.20 per million tokens. The mid-tier **GPT-5.6 Terra** is reduced by 20%, now costing $2.00 (input) and $12.00 (output) per million tokens. The flagship **GPT-5.6 Sol** maintains its original price but introduces a new **Fast mode**, offering speeds up to 2.5 times faster for double the cost. The company attributes these price reductions to efficiency gains achieved through **GPT-5.6 Sol's own involvement in optimizing its production systems**. The model assisted in rewriting GPU kernels and improving speculative decoding, leading to a 20% reduction in end-to-end service costs and over 15% improvement in token generation efficiency. OpenAI emphasizes this process remained human-led. A key focus of the降价 is to lower the barrier for running **AI agent workflows**. By making the capable, tool-calling Luna model significantly cheaper, OpenAI aims to enable more frequent use in cost-sensitive, high-volume tasks like code review and monitoring. This creates a potential feedback loop: model-assisted efficiency gains lead to lower costs, which enables broader agent deployment, which in turn drives further optimization. The new pricing and features will also apply to Codex and ChatGPT Work subscriptions. The changes intensify competition in the large language model market, with OpenAI directly challenging rivals like Anthropic to respond.

marsbitHace 33 min(s)

Breaking: GPT-5.6 Prices Slashed Effective Today

marsbitHace 33 min(s)

South Koreans' 'Gambling Nature' is Actually Forced by Life

This article explores how systemic pressures in South Korea, rather than inherent "gambling" tendencies, drive widespread speculative financial behavior. It begins by noting the high frequency of flights from South Korea to Macau, symbolizing the search for outlets beyond domestic restrictions. The core argument is that ordinary life goals—stable employment, home ownership, and financial security—have become increasingly tied to asset markets due to structural economic factors. South Korea's development model, historically reliant on corporate leverage (chaebols), has evolved into a society where household debt and personal leverage are normalized as pathways to social mobility. Key mechanisms discussed include: * **Housing Policy:** Government measures to improve affordability, like extending mortgage terms to 50 years and the unique *jeonse* (key money) rental system, embed high leverage into the housing market. * **Financial Products:** The recent approval and explosive popularity of single-stock 2x leveraged ETFs (e.g., on Samsung and SK Hynix), easily accessed via mobile apps, lowered barriers to high-risk trading. * **Social Pressure:** Media narratives around soaring corporate profits (e.g., SK Hynix) and employee bonuses create a fear of missing out, pushing individuals to use leverage to "catch up." The article concludes that this "leveraged life" is a product of institutional history and policy choices. When traditional paths to success feel constrained, and policy facilitates debt-based solutions for housing and investment, speculative behavior becomes a rational, if risky, strategy for many. The rapid cycle of regulatory approval for leveraged ETFs followed by a market crash and official apology in mid-2026 exemplifies the system's inherent contradictions. Ultimately, the "bet" is not just on assets, but on using future earnings to secure a place in the present society.

marsbitHace 33 min(s)

South Koreans' 'Gambling Nature' is Actually Forced by Life

marsbitHace 33 min(s)

Trading

Spot

Artículos destacados

Cómo comprar CORE

¡Bienvenido a HTX.com! Hemos hecho que comprar CORE (CORE) sea simple y conveniente. Sigue nuestra guía paso a paso para iniciar tu viaje de criptos.Paso 1: crea tu cuenta HTXUtiliza tu correo electrónico o número de teléfono para registrarte y obtener una cuenta gratuita en HTX. Experimenta un proceso de registro sin complicaciones y desbloquea todas las funciones.Obtener mi cuentaPaso 2: ve a Comprar cripto y elige tu método de pagoTarjeta de crédito/débito: usa tu Visa o Mastercard para comprar CORE (CORE) al instante.Saldo: utiliza fondos del saldo de tu cuenta HTX para tradear sin problemas.Terceros: hemos agregado métodos de pago populares como Google Pay y Apple Pay para mejorar la comodidad.P2P: tradear directamente con otros usuarios en HTX.Over-the-Counter (OTC): ofrecemos servicios personalizados y tipos de cambio competitivos para los traders.Paso 3: guarda tu CORE (CORE)Después de comprar tu CORE (CORE), guárdalo en tu cuenta HTX. Alternativamente, puedes enviarlo a otro lugar mediante transferencia blockchain o utilizarlo para tradear otras criptomonedas.Paso 4: tradear CORE (CORE)Tradear fácilmente con CORE (CORE) en HTX's mercado spot. Simplemente accede a tu cuenta, selecciona tu par de trading, ejecuta tus trades y monitorea en tiempo real. Ofrecemos una experiencia fácil de usar tanto para principiantes como para traders experimentados.

293 Vistas totalesPublicado en 2024.12.13Actualizado en 2026.06.02

Cómo comprar CORE

Discusiones

Bienvenido a la comunidad de HTX. Aquí puedes mantenerte informado sobre los últimos desarrollos de la plataforma y acceder a análisis profesionales del mercado. A continuación se presentan las opiniones de los usuarios sobre el precio de CORE (CORE).

活动图片