Mysterious Model HappyHorse Tops the Chart Overnight: Is the Video Generation Arena Welcoming a "Game Changer"?

marsbitXuất bản vào 2026-04-08Cập nhật gần nhất vào 2026-04-08

Tóm tắt

A mysterious AI video generation model named "HappyHorse-1.0" has quietly topped the AI Video Arena leaderboard on Artificial Analysis, surpassing established models like Seedance 2.0 and others in Elo score—a user-blind-test-based ranking reflecting real perceived quality. The model’s origin was initially unknown, but technical analysis later linked it to the open-source model "daVinci-MagiHuman," jointly developed by Shanghai SII GAIR Lab and Beijing-based Sand.ai. HappyHorse-1.0, likely an optimized iteration by Sand.ai, uses a 15-billion-parameter transformer architecture for joint audio-video-text modeling. Its strong performance in human-centric scenes (e.g., portraits, narrations) helped it excel in blind tests, though it still lags in multi-character or complex motion scenarios. The achievement signals a potential shift: an open-source model rivaling closed-source alternatives in perceived quality, which could lower costs and increase flexibility for developers in vertical applications like virtual avatars. However, limitations remain, including high computational requirements (H100 GPU needed) and shorter generation lengths. While not yet threatening market leaders, HappyHorse represents progress toward open models reaching "production-ready" quality, potentially accelerating community-driven improvements in the video AI space.

No launch event, no technical blog, no corporate backing—a text-to-video model named HappyHorse-1.0 quietly topped the AI Video Arena rankings on the authoritative AI evaluation platform Artificial Analysis, surpassing Seedance 2.0 with a higher Elo score and leaving mainstream players like Keling and Tiangang far behind, sparking a "decryption race" in the tech community.

Artificial Analysis' ranking is not based on technical parameter evaluations but on aggregated blind test results from real users, reflected through Elo scores. This makes the ranking harder to question than typical benchmark scores and turns "Who made this?" into an unavoidable question.

"Happy Horse" Quietly Tops the Chart, Sparking a Guessing Game in Tech Circles

Speculations on X emerged quickly. The first clue noticed was the language order on the official website: Mandarin and Cantonese were listed before English. For a product targeting global users, this order is unusual—if the team were U.S.-based, English would almost certainly be first. This strongly suggests the team behind it is from China.

The name itself is also a clue. 2026 is the Year of the Horse in the lunar calendar, and the name "HappyHorse" subtly references this, similar to the earlier "Pony Alpha." Suspects quickly piled up: Tencent and Alibaba's founders both have the surname Ma" (horse), putting them naturally on the list; some bet on Xiaomi, noting Lei Jun's low-key style and penchant for surprise reveals; others felt it aligned more with DeepSeek, which had quietly released a visual model before taking it down. Speculations ran wild, but no one had solid evidence.

The real breakthrough came from technical comparisons. X user Vigo Zhao cross-referenced HappyHorse-1.0's public benchmark data with known models and found a highly matching candidate: daVinci-MagiHuman, an open-source model called "DaVinci Magic Human" launched on GitHub in March.

Visual quality 4.80, text alignment 4.18, physical consistency 4.52, word error rate in speech 14.60%—each metric matched. The official website structure was nearly identical too: architecture descriptions, performance tables, and demo video styles all seemed to follow the same template. Both use a single-stream Transformer architecture, both support joint audio-video generation, and both support the same list of languages. This level of overlap is hard to dismiss as coincidence.

The most widely accepted conclusion in tech circles is that HappyHorse is an optimized iteration of the open-source model daVinci-MagiHuman, developed by Sand.ai, one of the joint developers. The core goal is to validate the model's performance上限 under real user preferences, paving the way for future commercialization.

daVinci-MagiHuman was officially open-sourced on March 23, 2026, a collaboration between two young teams. One is from the Generative Artificial Intelligence Research Laboratory (GAIR) at Shanghai Institute of Intelligence (SII), led by scholar Liu Pengfei; the other is Beijing-based Sand.ai (San Dai Tech), founded by Cao Yue, who also has an academic background, with a focus on autoregressive world models.

The model uses a 15-billion-parameter pure self-attention single-stream Transformer, packing text, video, and audio tokens into the same sequence for joint modeling—no one in the open-source community had previously attempted true joint pre-training of audio and video from scratch, as most efforts involved stitching together single-modal bases.

How Did an Open-Source Video Model Achieve a Two-Week Comeback?

Once the identity was clarified, another question became even harder to answer: daVinci-MagiHuman was only open-sourced in late March, so how did HappyHorse-1.0 manage to secure a higher Elo score than Seedance 2.0 in just two weeks?

Based on information disclosed on the official website, it's reasonable to speculate that HappyHorse made targeted adjustments to the default generation strategy for the evaluation scenario.

The Elo system essentially accumulates user preferences. Slight improvements in perceptually sensitive areas—like stable facial expressions, audio-visual alignment, and visual appeal—can make a big difference in blind tests. The model's capability上限 remains unchanged, but its "evaluation performance" can be polished.

In fact, over 60% of the blind test samples on Artificial Analysis involve portrait generation and voice-over content. daVinci-MagiHuman was trained with a focus on portrait performance, giving it a natural advantage in such scenarios, which is the main reason for its领先 blind test win rate. If blind test samples are dominated by portrait close-ups, models skilled in portraits will systematically benefit, unrelated to their actual performance in multi-character, complex camera work, or long-term narrative scenarios.

The result is a noticeable gap between the ranking numbers and actual test experiences, splitting X discussants into two camps. Skeptics, after testing, believe that HappyHorse-1.0 still lags behind Seedance 2.0 in character details and motion coherence, questioning the representativeness of the Elo score itself.

Supporters, however, hold high hopes for HappyHorse's potential, hoping it can address the industry pain point of "visual consistency across multi-shot sequences," something current mainstream video models haven't solved well. If daVinci-MagiHuman truly makes a breakthrough here, it could be far more significant than a ranking.

The model's limitations shouldn't be overshadowed by the numbers. Xiaohongshu blogger @JACK's AI World was among the first to deploy and test daVinci-MagiHuman. He found that it requires an H100 to run, making it nearly impossible for consumer-grade GPUs. Although the community is researching quantization solutions, local deployment for individual users remains challenging in the short term.

In terms of scenarios, it currently excels mainly with single characters; once multiple people appear or the scene becomes high, the quality drops—this isn't something tuning parameters can fix, as it's directly related to its design focus on portraits. Generation length is typically around 10 seconds; going longer risks instability, and high-definition output requires super-resolution plugins.

@JACK's AI World concluded: daVinci-MagiHuman's overall usability is not as good as LTX 2.3; it will only be suitable for daily use after the community successfully implements quantization.

Has the Video Generation Arena Finally Welcomed a True "Game Changer"?

Of course, leading the rankings once doesn't say much. Next, HappyHorse will need to undergo more thorough testing in areas like stability, high-concurrency access speed, cross-scene consistency, character control precision, and generalization beyond the test set. These are the core metrics that determine whether a model can truly enter creators' workflows.

But if we zoom out to the broader industry landscape, the signal this event sends is already clear enough.

Open-source video models themselves aren't new. But a visible gap in effectiveness has long existed between open-source and closed-source models—in scenarios requiring delivery to clients, the generation quality of open-source models has consistently failed to cross the threshold from "usable" to "deliverable." The pricing power of closed-source products like Keling and Seedance is, to a considerable extent, built upon this gap.

The significance this time lies in the fact that a product based on an open-source model has, for the first time, matched mainstream closed-source competitors in a blind test ranking based on real user perception. Regardless of how much tuning was done for the evaluation scenario, for closed-source vendors relying on this gap to maintain pricing power, this is at least a signal worth taking seriously.

For developers, the implications of this turning point are more concrete. In vertical scenarios like portraits, digital humans, and virtual anchors, once the generation quality of an open-source base reaches the "deliverable" threshold, the cost structure of self-deployment will undergo substantial changes—not just compressing API call costs, but more importantly, bringing data, models, and the entire inference pipeline under one's own control, offering customization depth and privacy compliance flexibility that closed-source solutions can hardly match.

HappyHorse-1.0 won't shake the market positions of Seedance 2.0 or Keling in the short term. But once the perception that open-source models can rival closed-source ones is established, subsequent quantization optimizations, vertical fine-tuning, and inference acceleration will be pushed forward by the community at a pace far exceeding that of closed-source products.

In this Year of the Horse, what's truly worth watching might not be which horse runs the fastest, but the fact that the track itself is widening.

This article is from the WeChat public account "AI Value Official," author: Xingye, editor: Meiqi

Câu hỏi Liên quan

QWhat is the name of the text-to-video model that recently topped the AI Video Arena leaderboard on Artificial Analysis?

AHappyHorse-1.0

QWhich open-source model is HappyHorse-1.0 highly suspected to be based on, according to technical comparisons?

AdaVinci-MagiHuman

QWhat is the core architectural approach used by the daVinci-MagiHuman model for joint audio-video modeling?

AA single-stream Transformer architecture that models text, video, and audio tokens in a unified sequence.

QWhat is the primary reason HappyHorse-1.0 performed so well in the user-blind-test-based Elo ranking system?

AIt was likely optimized for the evaluation scenarios, particularly excelling in human portrait generation and narration content, which made up over 60% of the test samples.

QWhat broader industry signal does HappyHorse-1.0's performance send, according to the article?

AIt signals that open-source models can achieve user-perceived quality comparable to closed-source commercial products, potentially changing cost structures and offering greater flexibility for developers in vertical scenarios.

Nội dung Liên quan

Lại chơi trò 'thoát ràng buộc'? Các module quang nội địa đối mặt với một bài kiểm tra sức ép

**Tóm tắt: Mỹ dự kiến cấm nhập khẩu module quang Trung Quốc, ngành công nghiệp nội địa đối mặt thử thách** Ngày 4/8, thông tin từ FCC (Ủy ban Truyền thông Liên bang Mỹ) cho biết cơ quan này đang soạn thảo lệnh cấm nhập khẩu các mẫu module thu phát quang mới từ Trung Quốc, dự kiến công bố và thực thi vào năm 2026, mặc dù quy định cuối cùng vẫn có thể thay đổi. Thị trường module quang toàn cầu hiện do các nhà sản xuất Trung Quốc thống trị, chiếm hơn 60% thị phần, với những tên tuổi lớn như Zhongji Innolight, Eoptolink. Tỷ trọng này thậm chí còn cao hơn ở phân khúc sản phẩm cao tốc 800G và 1.6T, đang có nhu cầu bùng nổ nhờ sức mạnh tính toán AI. Tuy nhiên, việc Mỹ "thoái vốn" khỏi module quang Trung Quốc gặp nhiều thách thức. Các hãng công nghệ lớn như Meta, Google, Microsoft, Amazon và NVIDIA có nhu cầu cực lớn về module tốc độ cao, ước tính tổng cộng khoảng 40 triệu đơn vị vào năm 2026. Trong khi đó, năng lực sản xuất trong nước của Mỹ rất hạn chế, tổng công suất hàng tháng của các nhà sản xuất Mỹ chưa bằng 1/5 một nhà sản xuất hàng đầu Trung Quốc. Các công ty Trung Quốc như Zhongji Innolight và Eoptolink có doanh thu từ thị trường nước ngoài (chủ yếu là Mỹ) chiếm trên 90%. Để đối phó với rủi ro chính sý trong týõng lai, các công ty Trung Quốc hàng đầu đã và đang xây dựng nhà máý lắp ráp ở nước ngoài, chủ yếu tại Đông Nam Á và Mexico. Tuy nhiên, hiệu quả của chiến lược này phụ thuộc vào phạm vi và cách thức hạn chế cuối cùng của Mỹ (nhắm vào sản phẩm hay toàn bộ doanh nghiệp). Phản ứng thị trường chứng khoán ban đầu khá biến động, nhưng cổ phiếu các công ty module quang chính ở Trung Quốc sau đó đã thu hẹp phần lớn mức giảm. Các chuyên gia cho rằng, áp lực từ bên ngoài có thể thúc đẩy quyết tâm tự chủ nghiên cứu và phát triển của Trung Quốc, mở rộng thị trường nội địa và toàn cầu, từ đó củng cố hõn nữa vị thế ngành công nghiệp.

marsbit27 phút trước

Lại chơi trò 'thoát ràng buộc'? Các module quang nội địa đối mặt với một bài kiểm tra sức ép

marsbit27 phút trước

Khi cuộc cạnh tranh thiết bị chip bắt đầu không chỉ xem ai tiên tiến hơn

Cuộc cạnh tranh trong lĩnh vực thiết bị chip hiện nay không chỉ đơn thuần xoay quanh việc ai tiên tiến hơn, mà còn phụ thuộc vào tính ổn định và sự đáng tin cậy lâu dài của chuỗi cung ứng. Các quy định xuất khẩu của Mỹ đã thay đổi cách các công ty chip toàn cầu đánh giá thiết bị. Bên cạnh các tiêu chí truyền thống như hiệu suất, năng suất và chi phí, khả năng tiếp tục nhận được hỗ trợ, nâng cấp và linh kiện thay thế trong dài hạn mà không bị gián đoạn bởi các yếu tố chính trị hiện là mối quan tâm lớn. Điều này đang làm suy giảm lợi thế cốt lõi về độ tin cậy của các nhà cung cấp thiết bị Mỹ. Trong bối cảnh đó, các công ty Hàn Quốc như Samsung và SK Hynix được cho là đang đánh giá thiết bị khắc của công ty Trung Quốc AMEC tại các nhà máy ở Trung Quốc. Hành động này không nhất thiết dẫn đến các đơn hàng lớn ngay lập tức, nhưng nó phản ánh nhu cầu chiến lược về một "nhà cung cấp dự phòng" để giảm thiểu rủi ro chuỗi cung ứng. Các công ty cần xác minh sớm liệu thiết bị nội địa có thể duy trì hoạt động sản xuất nếu nguồn cung từ Mỹ gặp trở ngại. Sự phát triển của các thiết bị chip Trung Quốc, đặc biệt trong các lĩnh vực như khắc (ví dụ: AMEC) và các công đoạn khác, là điều kiện tiên quyết để nắm bắt cơ hội này. Chúng đã vượt qua ngưỡng cơ bản và được tích lũy kinh nghiệm trong môi trường sản xuất thực tế ở thị trường nội địa. Tuy nhiên, đây không phải là sự thay thế toàn cầu. Việc đánh giá hiện chỉ giới hạn ở các nhà máy tại Trung Quốc, và con đường từ thử nghiệm đến mua hàng hàng loạt vẫn còn dài. Ngành công nghiệp thiết bị Trung Quốc vẫn thiếu hụt ở nhiều khâu then chốt như máy quang khắc cao cấp và một số linh kiện lõi. Tóm lại, các biện pháp kiểm soát của Mỹ đã vô tình tạo ra động lực để các công ty đa quốc gia xem xét kỹ hơn các nhà cung cấp thiết bị Trung Quốc như một biện pháp phòng ngừa rủi ro. Cuộc cạnh tranh giờ đây không chỉ là về công nghệ tiên tiến, mà còn về việc ai có thể cung cấp một chuỗi cung ứng ổn định và đáng tin cậy trong dài hạn. Các thiết bị Trung Quốc có được cơ hội tham gia, nhưng thành công cuối cùng vẫn phụ thuộc vào năng lực công nghệ và dịch vụ thực sự.

marsbit27 phút trước

Khi cuộc cạnh tranh thiết bị chip bắt đầu không chỉ xem ai tiên tiến hơn

marsbit27 phút trước

Khối lượng giao dịch tăng gấp 2.5 lần, tại sao doanh thu của Circle chỉ tăng 7%?

Circle đã công bố kết quả kinh doanh quý 2 với một nghịch lý: khối lượng giao dịch trên chuỗi của USDC tăng 151% lên 14,8 nghìn tỷ USD, nhưng "Tổng doanh thu và doanh thu từ dự trữ" chỉ tăng 7% lên 701 triệu USD. Nguyên nhân chính là doanh thu chủ yếu đến từ lãi suất thu được trên lượng USDC lưu thông trung bình, chứ không phải từ số lần giao dịch. Mặc dù lượng USDC lưu thông trung bình tăng 25%, nhưng tỷ suất sinh lời từ dự trữ lại giảm 66 điểm cơ bản, khiến mức tăng doanh thu thực tế bị hạn chế. Sau khi trừ đi "Tổng chi phí phân phối, giao dịch và các chi phí khác", chỉ số RLDC (Doanh thu còn lại sau chi phí phân phối) của Circle tăng lên 289 triệu USD, với tỷ suất RLDC cải thiện từ 38,2% lên 41,2%. Tuy nhiên, chi phí hoạt động điều chỉnh tăng 23%, còn EBITDA điều chỉnh chỉ tăng 8%, cho thấy việc đầu tư vào các sản phẩm mới và cơ sở hạ tầng chưa mang lại hiệu quả lợi nhuận ngay lập tức. Circle cũng báo cáo một số tiến triển về mạng lưới, như Circle Payments Network đạt khối lượng giao dịch 14,7 tỷ USD (tính theo năm trong 30 ngày qua) và có 175 tổ chức tài chính kết nối. Dù vậy, các chỉ số này phản ánh mức độ phát triển mạng lưới chứ chưa phải là nguồn doanh thu đáng kể trong quý. Tóm lại, báo cáo nhấn mạnh rằng tăng trưởng doanh thu của Circle hiện chủ yếu phụ thuộc vào quy mô dự trữ USDC và môi trường lãi suất, hơn là trực tiếp từ cường độ giao dịch trên chuỗi.

marsbit45 phút trước

Khối lượng giao dịch tăng gấp 2.5 lần, tại sao doanh thu của Circle chỉ tăng 7%?

marsbit45 phút trước

Samsung tại Trung Quốc, lại rút thêm một bước

Samsung Electronics tiếp tục thu hẹp sự hiện diện tại thị trường Trung Quốc. Sau khi rút khỏi thị trường thiết bị gia dụng vào tháng 5, giờ đây các cửa hàng bán lẻ điện thoại Samsung có doanh thu hàng tháng dưới 300.000 NDT (khoảng 30.000 USD) cũng đang bị đóng cửa tại nhiều thành phố như Thâm Quyến, Phúc Châu, Trịnh Châu và Tây An. Thị phần điện thoại Samsung tại Trung Quốc đã giảm mạnh, chỉ còn 0,1% vào quý 2/2026, trong khi các thương hiệu nội địa như Huawei, Apple, OPPO, vivo thống trị thị trường. Nguyên nhân được cho là do giá thành cao, tính năng không còn cạnh tranh, hệ sinh thái "không hợp thổ nhưỡng" và ảnh hưởng từ các yếu tố văn hóa. Tuy nhiên, trên phạm vi toàn cầu, Samsung vẫn dẫn đầu với 22% thị phần điện thoại thông minh. Công ty hiện đang tập trung mạnh vào thị trường cao cấp, như điện thoại gập cao cấp Galaxy Z Fold8, để tối ưu hóa lợi nhuận. Chiến lược này diễn ra song song với việc bộ phận bán dẫn (chủ yếu là bộ nhớ) của tập đoàn đang tạo ra lợi nhuận khổng lồ, chiếm 99,7% lợi nhuận hoạt động trong quý 2/2026. Bài toán đặt ra cho Samsung là sự mất cân bằng nội bộ: bán dẫn quá mạnh khiến chi phí cho điện thoại tăng cao do cơ chế chuyển giá nội bộ, trong khi các bộ phận khác như màn hình cũng không muốn nhường sân cho đối thủ. Việc rút lui khỏi các thị trường đại chúng tại Trung Quốc, tập trung vào phân khúc siêu cao cấp có thể là lựa chọn thực tế, nhưng cũng đồng nghĩa với việc từ bỏ ảnh hưởng và hiệu ứng quy mô. Tương lai của Samsung đang ngày càng phụ thuộc vào chu kỳ thăng trầm của ngành bán dẫn.

marsbit48 phút trước

Samsung tại Trung Quốc, lại rút thêm một bước

marsbit48 phút trước

Giao dịch

Giao ngay
活动图片