LeCun Continuously Endorses, New Work VISReg Tackles the Core Challenge of 'Representation Collapse' in JEPA World Models

marsbitPublicado em 2026-07-28Última atualização em 2026-07-28

Resumo

"VISReg: A New Self-Supervised Learning Method Tackles Representation Collapse in JEPA World Models" Yann LeCun has highlighted the new self-supervised learning (SSL) work VISReg (Variance-Invariance-Sketching Regularization), which addresses the core challenge of "representation collapse" in JEPA-based world models. SSL often collapses, mapping different inputs to similar vectors, losing discriminative power. While methods like VICReg and SIGReg attempted to regularize the representation distribution, they suffered from issues like vanishing gradients during collapse or coupled scale and shape optimization. VISReg overcomes these by decoupling the regularization into independent "scale" and "shape" objectives. It uses a variance term to prevent amplitude collapse (providing constant gradient even during collapse) and a sliced Wasserstein distance (SWD) based sketching target with stop-gradient to align the distribution shape to an isotropic Gaussian, without interfering with scale. This approach, requiring only ~15 lines of PyTorch core code, offers linear computational complexity and scales efficiently across multiple GPUs. Evaluated across 15 datasets (in-domain, out-of-distribution/OOD, dense prediction), VISReg outperforms 7 mainstream SSL methods (MoCoV3, DINO, iBOT, I-JEPA, MAE, data2vec) without relying on heuristic tricks like EMA or stop-gradient. Key results include: superior average OOD accuracy; matching DINOv2's OOD performance using only ~1/10 of the data (I...

The foundation of JEPA world models is Self-Supervised Learning (SSL), advocated by Yann LeCun since 2017.

SSL can learn general representations from massive amounts of data without manual annotation, but it commonly faces a core challenge—representation collapse: the model tends to map different inputs to the same or very few vectors, seemingly completing training without actually learning discriminative representations.

To suppress collapse, mainstream methods mostly rely on a series of heuristic techniques (EMA, teacher-student networks, stop-gradient, freezing layers, etc.). These techniques make training fragile, difficult to tune, and also weaken the method's interpretability and scalability.

Another approach is to directly constrain the representation distribution through regularization terms.

VICReg, proposed by LeCun's team, decomposes the learning objective into three terms: variance, invariance, and covariance, using covariance to constrain correlations between dimensions; but covariance only captures second-order statistics, unable to distinguish between two representations with "the same mean and variance, but drastically different distribution shapes."

Subsequently proposed SIGReg, based on the Cramér–Wold theorem, aligns the entire embedding distribution to a standard Gaussian using sketching techniques, thereby constraining the complete distribution shape.

However, SIGReg still has two key flaws:

  • Gradient Vanishing upon Collapse: When representations start to collapse, SIGReg's gradient decays accordingly—the more severe the collapse, the weaker the correction signal, making it difficult for the model to recover on its own;
  • Coupling of Scale and Shape: It does not separate the two independent attributes of "magnitude (scale)" and "distribution morphology (shape)". They interfere with each other during optimization, leading to poor adaptability on long-tailed, low-quality, low-rank data.

In other words, when the model most needs gradient signals to escape the collapsed state, SIGReg's gradient tends to vanish.

This is precisely the core problem that VISReg aims to solve.

Recently, the new SSL work VISReg (Variance-Invariance-Sketching Regularization) has been continuously endorsed and highly recognized by Turing Award winner Yann LeCun—he commented on the repost, "VICReg begat SIGReg which begat VISReg," succinctly outlining the technical lineage of this regularization route.

To gain such recognition from LeCun, what makes VISReg so strong?

The answer lies in the fact that it precisely targets the core challenge of JEPA world models that LeCun has long bet on—representation collapse.

Paper Link: https://arxiv.org/abs/2606.02572

Code / Pre-trained Weights: https://github.com/HaiyuWu/visreg

Project Page: https://haiyuwu.github.io/visreg/

VISReg decouples the collapse-preventing regularization term into two independent objectives: "scale" and "shape." Without relying on any heuristic training tricks or massive data, it outperforms 7 mainstream SSL methods on a comprehensive evaluation across 15 datasets; using only about 1/10 of the training data, it matches DINOv2 on out-of-distribution (OOD) benchmarks.

Figure 2: Simulation of gradient magnitude ‖∇L‖ for different regularization methods at various stages of representation collapse. VISReg maintains strong gradients even in collapsed states, while SIGReg's gradient almost vanishes.

Core Method

VISReg combines the strengths of VICReg and SIGReg: it retains VICReg's variance term to control scale, while replacing the covariance term with a sketching objective based on Sliced Wasserstein Distance (SWD) to control shape, and completely decouples the two via stop-gradient. The entire regularization objective consists of three parts.

1. Scale Regularization

The first part constrains the variance of each dimension to prevent amplitude collapse:

Its key property is: when the model collapses, the gradient of this term approaches a constant, ensuring the model can recover stably—this exactly compensates for SIGReg's gradient vanishing flaw.

2. Shape Regularization

The second part first normalizes to eliminate scale influence, then constrains shape separately. The crucial step is normalization with "stop-gradient" (sg):

Here, stop-gradient is applied to the standard deviation σ, ensuring that optimizing the shape loss does not conversely alter the scale—this is the mechanism enabling true decoupling and non-interference between the "scale" and "shape" objectives.

After normalization, the geometric shape of the distribution is aligned to an isotropic Gaussian using the Sliced Wasserstein Distance:

where

is the standard Gaussian quantile, and

is the random projection direction (i.e., "slice / sketching").

The theoretical basis is the Cramér–Wold theorem (Lemma 3.1 in the paper): two distributions are equal if and only if all their one-dimensional projections along directions on the unit sphere are equal. Therefore, by slicing the high-dimensional representation along sufficiently many random one-dimensional directions and aligning each to a Gaussian, we equivalently align the entire distribution in the high-dimensional space—this allows us to characterize the complete distribution shape using cheap one-dimensional sorting operations, rather than just second-order statistics.

3. Combined Objective

The third part is a centering loss that pulls the batch mean μ towards the origin:

The three regularization terms are combined with weights:

The predictive loss follows the invariance objective of JEPA / LeJEPA—aligning embeddings from each view (global + local, total V views)

to the mean of the global view

:

Finally, a single hyperparameter λ balances prediction and regularization, yielding the full objective:

Comparison with VICReg: VICReg also decouples regularization into variance + covariance, but covariance only captures second-order statistics; VISReg uses a sketching objective based on Sliced Wasserstein to fully characterize the distribution shape while retaining the variance term for scale control—preserving VICReg's flexibility while gaining distribution-level rigor.

Requires Only About 15 Lines of PyTorch Code

This regularization objective is very lightweight to implement; the core logic takes only about 15 lines:

Computational Complexity and Scalability

In terms of computation and scalability, VISReg also has advantages. The computational complexity of its regularization part is

(N is batch size, D is dimension, K is number of slices), which is linear for all scaling factors; in contrast, VICReg's covariance term is

, scaling quadratically with dimension.

Under the same batch size, VISReg's runtime speed and GPU memory usage on a single H100 GPU are superior to SIGReg.

More importantly, the K random slices can be distributed across multiple GPUs: each of M GPUs generates K/M slices, achieving an effect equivalent to a single GPU generating all K slices.

In experiments, when the number of slices per GPU was insufficient, switching to 8 GPUs, each with 128 slices (total 1024), reduced the accuracy gap with "single GPU 1024 slices" from about 2.4% to 0.22%. This means that K can remain constant when scaling up training, adding almost no per-GPU burden.

Figure: Change in linear probing accuracy when increasing GPU count with fixed K and D. When K is insufficient (K = 1⁄4D), using 8 times the number of GPUs can raise accuracy to the level of K = 2D—making it possible to keep K constant during large-scale training.

Experimental Results

Back to the question in the title—where exactly is VISReg strong? The research team compared VISReg with 7 mainstream SSL methods—MoCoV3, DINO, iBOT, I-JEPA, MAE, data2vec—on 15 datasets (8 in-domain + 6 out-of-distribution + ADE20K dense prediction), covering domains such as astronomy, medical, remote sensing, texture, flowers, etc. The answer manifests across multiple dimensions from recognition to segmentation and generation.

1. In-Domain Linear Probing

To ensure a fair comparison, experiments are divided into two groups based on whether heuristic techniques are used. In the group not using any heuristic techniques, VISReg leads: ViT-B/16 achieves in-domain linear probing accuracy of 75.7%, higher than MAE (75.1%); ViT-L/14 further improves to 77.0%, higher than LeJEPA (75.6%). Compared to iBOT and DINO which use heuristic techniques, VISReg is only slightly lower on conventional datasets, but surpasses all methods on the texture dataset DTD—indicating its cross-domain generalization ability stems from the method itself, not stacking manual tricks.

2. Out-of-Distribution (OOD) Generalization: Comprehensive Superiority

OOD generalization is a stricter test than in-domain accuracy: methods relying on heuristics are often finely tuned on the ImageNet in-domain but may not transfer well to significantly different new distributions. The team evaluated on 6 OOD datasets covering medical (ChestXRay, RetinaMNIST, OrganAMNIST), astronomy (Galaxy10), remote sensing (AID), texture (DTD), which are completely unrelated to the ImageNet training domain. Results show VISReg achieved the best average OOD accuracy across all methods and all backbone sizes, even surpassing some methods using heuristic techniques with larger backbones.

Figure 4: Average OOD linear probing accuracy. VISReg comprehensively outperforms iBOT, DINO, MoCoV3, I-JEPA, MAE, data2vec, etc.

As shown in Figure 4, ViT-B/16 VISReg average OOD accuracy is 70.19%, ViT-L/14 is 70.63%, significantly higher than MAE (67.85%), and better than MoCoV3 (69.46%), DINO (69.56%), I-JEPA (68.55%), etc.

3. Data Efficiency: Matching DINOv2 with 1/10 the Data

After pre-training VISReg (ViT-L/14) on ImageNet-22K (~14 million images), its average accuracy on 6 OOD datasets reaches 72.94%, essentially on par with DINOv2 (72.93%) trained on 10x larger scale LVD-142M (142 million images). In other words, VISReg achieves comparable performance using about 1/10 of the data. (As a control, VISReg with ViT-L/14 pre-trained only on ImageNet-1K has an average accuracy of 70.63%.) This indicates its learned representations have strong generality.

Figure 5: VISReg pre-trained on ImageNet-22K matches DINOv2 trained on 10x more data (LVD-142M) on OOD benchmarks.

4. Transfer Fine-tuning: Comprehensively Surpassing DINO

Although VISReg's linear probing accuracy on some in-domain datasets is slightly lower than DINO, after fine-tuning, it surpasses both DINO and supervised pre-training on all five tested datasets—CIFAR-10, CIFAR-100, Flowers, ImageNet-1K, Galaxy10—indicating its representation distribution is more uniform, less redundant, and more transferable.

Figure: Transfer learning comparison. After fine-tuning, VISReg outperforms DINO and supervised pre-training on all tested datasets (CIFAR-10, CIFAR-100, Flowers, ImageNet-1K, Galaxy10).

5. Dense Prediction and Generation Guidance

VISReg's advantages are not limited to classification. On ADE20K linear semantic segmentation (ViT-B/16), its mIoU is 30.16, higher than DINO (29.40) and MAE (23.60), second only to MoCoV3 (31.69); this result is competitive without using any heuristic tricks. The paper also acknowledges that there is still a gap with the best methods in dense prediction, a focus for future optimization.

Figure 7: Linear semantic segmentation on ADE20K. Without any heuristic tricks, VISReg achieves competitive mIoU, second only to MoCoV3.

In generation guidance (SiT-B/2, iREPA framework, 100k training steps), generation guided by VISReg features outperforms DINO on three out of four metrics: gFID 40.36 (DINO 41.15), Precision 51.38 (DINO 50.51), Recall 61.26 (DINO 60.70), IS essentially tied (33.48 vs 33.47). This shows VISReg's learned representations are also superior as generation guidance signals.

Figure 8: Image generation guided by VISReg vs. DINO features using SiT-B/2. VISReg provides better guidance on most metrics (lower gFID, higher Precision and Recall).

6. Robustness on Low-Quality Data

On low-quality datasets such as long-tail distribution (ImageNet-LT) and low-rank (Galaxy10), VISReg can stably prevent collapse and learn meaningful representations, while DINO fails directly without fine-tuned hyperparameters.

Table 1: Linear probing accuracy on ImageNet-LT

Table 2: In-domain linear probing accuracy on Galaxy10

Conclusion

VISReg demonstrates that by decoupling representation regularization into two independent components, "scale" and "shape," one can obtain an SSL method that is more stable, efficient, and has stronger generalization than existing methods.

Without using any training heuristic tricks, it achieves leading or near-leading results across multiple dimensions including image recognition, segmentation, and generation guidance, and matches DINOv2's OOD performance with about 1/10 of the data. This provides a new regularization-based solution to the long-standing representation collapse problem in JEPA world models.

References:

https://arxiv.org/abs/2606.02572

This article is from the WeChat public account "Xin Zhi Yuan," author: LRST

Criptomoedas em alta

Perguntas relacionadas

QWhat is the core problem that VISReg aims to solve in self-supervised learning?

AVISReg aims to solve the core problem of representation collapse in self-supervised learning, where a model maps different inputs to the same or very few vectors, failing to learn discriminative representations.

QHow does VISReg fundamentally differ from its predecessor SIGReg in handling representation collapse?

AVISReg differs from SIGReg by decoupling the regularization into two independent targets: 'scale' and 'shape'. Crucially, it maintains a strong gradient signal even when the model is collapsing, whereas SIGReg suffers from vanishing gradients in such states, making recovery difficult.

QWhat are the three main components of the VISReg regularization objective?

AThe three main components are: 1) Scale Regularization (variance term to prevent magnitude collapse), 2) Shape Regularization (using Sliced Wasserstein Distance after stop-gradient normalization to align the distribution shape to an isotropic Gaussian), and 3) A centering loss that pulls the batch mean towards the origin.

QWhat key advantage in data efficiency does VISReg demonstrate compared to DINOv2 according to the experimental results?

AAccording to the experimental results, VISReg achieves comparable performance to DINOv2 on out-of-distribution benchmarks using only about 1/10th of the training data. Specifically, VISReg trained on ImageNet-22K (~14M images) matched the average OOD performance of DINOv2 trained on LVD-142M (~142M images).

QIn which evaluation scenarios did VISReg outperform DINO after fine-tuning?

AAfter fine-tuning, VISReg outperformed DINO and supervised pre-training across all five tested datasets: CIFAR-10, CIFAR-100, Flowers, ImageNet-1K, and Galaxy10.

Leituras Relacionadas

SemiAnalysis on the Epic Plunge: It's Not Over Yet

SemiAnalysis Weekly discusses the recent sharp correction in the semiconductor market after a historic first half. Analysts Doug O'Loughlin and Dylan note that despite healthy fundamentals, markets like South Korea's KOSPI have plunged 40%, wiping out leveraged retail investors. A core debate focuses on AI demand versus supply constraints. Dylan cites SemiAnalysis's internal use of AI coding agents, leading to a 100x increase in AI spending, as evidence of powerful demand. Doug agrees demand is strong but questions its exact magnitude, calling it a "trillion-dollar question." His primary concern is physical and financial bottlenecks: a shortage of 100,000 electricians in the US, massive $450B in corporate debt issuance by hyperscalers (funded by a shrinking pension pool), and labor/scale limits in regions like Taiwan, where TSMC constitutes 20% of GDP. The conversation covers market dynamics, including the typical semiconductor cycle where over-ordering leads to crashes, China's growing memory capacity, and the potential for older chips like the H100 to lose value as models scale. Politically, AI is seen as a likely scapegoat in upcoming elections, though not a top-tier voter priority. The analysts conclude that while the long-term potential is significant, the scaling path is narrowing. The challenge is matching exponential compute demands with real-world constraints on capital, labor, and permits, risking scenarios where massive investment outpaces near-term revenue generation.

marsbitHá 32m

SemiAnalysis on the Epic Plunge: It's Not Over Yet

marsbitHá 32m

After Robinhood's Aggressive Entry, Why is PUMP Defying the Trend and Rising?

Amid a prolonged crypto bear market that has strained many projects, Pump.fun stands out by maintaining strong profitability. Despite a major token unlock in mid-July that was expected to cause selling pressure, its native token PUMP has risen nearly 48% over the past month. Analysis shows only a small portion of unlocked tokens were sold, with 92% of investor allocations still held. A buyback-and-burn mechanism has also supported the price, though daily回购 funds have declined by 77.5% from their peak, partly due to a change in tokenomics and lower platform revenue. Pump.fun remains a top revenue generator in crypto, with annualized revenue of around $450 million, ranking just behind giants like Tether and Circle. However, its daily revenue is down about 69% from its all-time high, reflecting dependence on meme coin market cycles. The platform also faces a class-action lawsuit alleging market manipulation and securities violations. New competition has emerged with Robinhood Chain's launchpad, which recently saw higher weekly volume and active addresses than Pump.fun. While Pump.fun's metrics haven't declined yet, it has responded with a new token launch mechanism called BOOST to improve liquidity. This has increased platform activity but raised community concerns about potential exploitation for short-term pumps. Ultimately, Pump.fun's success remains tightly linked to meme coin speculation and broader market sentiment.

marsbitHá 32m

After Robinhood's Aggressive Entry, Why is PUMP Defying the Trend and Rising?

marsbitHá 32m

China Business Journal Unveils Cryptocurrency Extortion Scheme

China Business Journal, a state-run Chinese newspaper, has issued a warning about scammers using its name to extort cryptocurrency from companies. According to the publication, fraudsters are sending emails impersonating its journalists. These emails threaten to publish damaging information about a company unless a Bitcoin payment is made. The scammers, using Proton Mail addresses, claim to possess compromising material from a secretive investigation. The newspaper's representatives have clarified that it has no involvement with these messages. They emphasized that the publication does not engage in such practices, never demands payment to suppress stories, and does not negotiate in this manner. The editorial team believes the criminals are exploiting the newspaper's authoritative reputation to pressure companies into making hasty decisions under threat of reputational harm. China Business Journal stated it is collecting evidence of this brand misuse, analyzing the fraudulent emails, and preparing materials for law enforcement. It reserves the right to pursue both civil and criminal liability against the scheme's organizers. Companies are urged to carefully verify any communication purportedly from the newspaper, avoid transferring cryptocurrency to unknown recipients, and report such incidents to the authorities. The newspaper confirmed it will continue its own investigation into the matter.

cryptonews.ruHá 1h

China Business Journal Unveils Cryptocurrency Extortion Scheme

cryptonews.ruHá 1h

Trading

Spot

Artigos em Destaque

Como comprar CORE

Bem-vindo à HTX.com!Tornámos a compra de CORE (CORE) simples e conveniente.Segue o nosso guia passo a passo para iniciar a tua jornada no mundo das criptos.Passo 1: cria a tua conta HTXUtiliza o teu e-mail ou número de telefone para te inscreveres numa conta gratuita na HTX.Desfruta de um processo de inscrição sem complicações e desbloqueia todas as funcionalidades.Obter a minha contaPasso 2: vai para Comprar Cripto e escolhe o teu método de pagamentoCartão de crédito/débito: usa o teu visa ou mastercard para comprar CORE (CORE) instantaneamente.Saldo: usa os fundos da tua conta HTX para transacionar sem problemas.Terceiros: adicionamos métodos de pagamento populares, como Google Pay e Apple Pay, para aumentar a conveniência.P2P: transaciona diretamente com outros utilizadores na HTX.Mercado de balcão (OTC): oferecemos serviços personalizados e taxas de câmbio competitivas para os traders.Passo 3: armazena teu CORE (CORE)Depois de comprar o teu CORE (CORE), armazena-o na tua conta HTX.Alternativamente, podes enviá-lo para outro lugar através de transferência blockchain ou usá-lo para transacionar outras criptomoedas.Passo 4: transaciona CORE (CORE)Transaciona facilmente CORE (CORE) no mercado à vista da HTX.Acede simplesmente à tua conta, seleciona o teu par de trading, executa as tuas transações e monitoriza em tempo real.Oferecemos uma experiência de fácil utilização tanto para principiantes como para traders experientes.

368 Visualizações TotaisPublicado em {updateTime}Atualizado em 2026.06.02

Como comprar CORE

Discussões

Bem-vindo à Comunidade HTX. Aqui, pode manter-se informado sobre os mais recentes desenvolvimentos da plataforma e obter acesso a análises profissionais de mercado. As opiniões dos utilizadores sobre o preço de CORE (CORE) são apresentadas abaixo.

活动图片