# Multimodal Related Articles

HTX News Center provides the latest articles and in-depth analysis on "Multimodal", covering market trends, project updates, tech developments, and regulatory policies in the crypto industry.

Google Officially Declares War

Google Declares War with AI-First I/O 2026 At its 2026 I/O developer conference, Google launched an aggressive, multi-pronged offensive, embedding AI across its ecosystem and challenging rivals on performance and price. The event showcased three major releases: Gemini 3.5 Flash, the video-centric Gemini Omni Flash, and the system-level AI assistant Spark. Gemini 3.5 Flash, despite being a smaller "Flash" model, outperforms its Pro counterpart in key benchmarks like mathematical reasoning (GSM8K) and coding (SWE-bench). Google attributes this to "extreme knowledge distillation" from a larger teacher model and a novel, highly granular MoE (Mixture of Experts) architecture with 256 experts, achieving sub-65ms response times. The native multi-modal model, Gemini Omni Flash, offers real-time video understanding with 120ms latency, enabling applications like preventing a cup from overfilling. The new Spark assistant gains deep Android system integration, allowing it to automate complex multi-app workflows based on voice commands. Complementing these, Google unveiled lightweight AI glasses featuring Micro-OLED displays and on-device Gemini chips for instant, offline translation and scene analysis. CEO Sundar Pichai announced Gemini has reached 900 million monthly active users, leveraged through integration into Chrome, Android, and Workspace. Google also slashed prices dramatically: the Gemini 3.5 Flash API is priced at a fraction of competitor rates. This price war is enabled by Google's vertically integrated TPU infrastructure. The strategy signals a shift: standalone AI models are becoming commoditized. Google's advantage lies in its "device + cloud + ecosystem + hardware" integration, aiming to reshape internet traffic from user-initiated searches to AI-driven service distribution. This move pressures pure-play AI firms like OpenAI and Anthropic on business models, and challenges Apple to respond in the next-generation, screen-less device race.

链捕手05/21 13:40

Google Officially Declares War

链捕手05/21 13:40

The Creator of Kling Returns to Alibaba and Builds Another Dark Horse

The article discusses the rise of HappyHorse-1.0, an AI video generation model developed by Alibaba, which topped the Artificial Analysis leaderboard in both text-to-video and image-to-video categories in April 2026. The model was created under the leadership of Zhang Di, who returned to Alibaba in November 2025 after working at Kuaishou, where he led the development of the Kling model. HappyHorse is open-source and commercially available, similar to Alibaba's Qwen model. Zhang Di's background includes extensive experience in large-scale data systems and machine learning at Alibaba and Kuaishou, which contributed to the rapid development of HappyHorse within just five months. The model uses a 15-billion-parameter transformer architecture with native multimodal training, supporting multiple languages and lip-sync capabilities. It also focuses on reducing inference time and cost, making it practical for commercial use. The primary application of HappyHorse is in e-commerce, where it can generate product videos to enhance user engagement and conversion rates by creating contextual and personalized content. This aligns with Alibaba's strengths in commerce, advertising, and data feedback loops. The model's success with open-source approach contrasts with challenges faced by closed-source models like OpenAI's Sora (shut down due to high costs) and ByteDance's Seedance 2.0 (paused over copyright issues). HappyHorse represents a strategic move for Alibaba to integrate AI video generation into its core business ecosystems.

marsbit04/13 05:10

The Creator of Kling Returns to Alibaba and Builds Another Dark Horse

marsbit04/13 05:10

Claiming the "Happy Horse": Alibaba's AI Lays Out the "Eight Trigrams Formation"

Alibaba has officially claimed the "HappyHorse" (HappyHorse-1.0) AI video generation model, which recently topped the global benchmark on Artificial Analysis with an Elo score of 1357. Developed by Alibaba’s ATH (Alibaba Token Hub) innovation unit, the model is notable for its ability to generate high-definition video with synchronized audio and sound effects from text input, significantly improving motion coherence and reducing production time and cost. This launch is part of a broader acceleration in Alibaba’s AI strategy. In late March and early April, the company released three flagship models in quick succession: Qwen3.5-Omni, Wan2.7-Image, and Qwen3.6-Plus. The latter broke global daily call volume records with 1.4 trillion tokens processed shortly after release. Alibaba has also undergone significant organizational restructuring to support its AI ambitions. In March, it established the ATH business group, led by CEO Wu Yongming, to integrate AI development, cloud services, and application deployment. Further changes in April included forming a group-level technology committee and consolidating the Tongyi Lab into a dedicated AI model division. The company is investing heavily in AI, with plans to spend over 380 billion RMB on cloud and AI infrastructure over three years. Its self-developed GPUs have already seen mass production. While the market has responded positively to these moves, challenges remain in balancing centralized control with operational flexibility and maintaining team stability amid rapid changes.

marsbit04/11 04:07

Claiming the "Happy Horse": Alibaba's AI Lays Out the "Eight Trigrams Formation"

marsbit04/11 04:07

From 'Word Unit' to 'Symbol Unit': The Debate Over the Chinese Translation of 'Token' and Its Underlying AI Cognitive Implications

Recent discussions have emerged regarding the official Chinese translation of the AI term "Token," which has been recommended as “词元” (Cíyuán, meaning "word unit") by the National Committee for Terminology in Science and Technology. While this translation is argued to align with historical usage in natural language processing (NLP) and is considered concise and communicable, this article presents a critical counterview advocating for “符元” (Fúyuán, meaning "symbol unit") as a more structurally accurate and future-proof alternative. The author argues that defining Token based on its origin in NLP—as a linguistic semantic unit—overlooks its evolution into a general-purpose, discrete symbolic unit used across multimodal systems (text, image, audio, etc.). Using “词元” ties the concept too narrowly to language, causing cognitive misalignment and semantic drift when applied in non-linguistic contexts. By contrast, “符元” reflects Token’s fundamental role as a symbol in information theory and computation, independent of modality. The article further critiques the reliance on metaphorical extensions (e.g., comparing image patches to “words”) as insufficient for rigorous terminology. It highlights risks including confusion with existing linguistic terms like Lemma (also translated as “词元”), poor cross-lingual reversibility (e.g., difficult back-translation to English), and systemic misunderstanding among non-expert audiences. In conclusion, the author emphasizes that terminology should align with computational essence—not historical usage or explanatory convenience—to ensure conceptual clarity and scalability in AI’s multidisciplinary future. “符元” is proposed as a more neutral, stable, and structurally coherent translation for Token.

marsbit04/10 10:43

From 'Word Unit' to 'Symbol Unit': The Debate Over the Chinese Translation of 'Token' and Its Underlying AI Cognitive Implications

marsbit04/10 10:43

Mysterious Model HappyHorse Tops the Chart Overnight: Is the Video Generation Arena Welcoming a "Game Changer"?

A mysterious AI video generation model named "HappyHorse-1.0" has quietly topped the AI Video Arena leaderboard on Artificial Analysis, surpassing established models like Seedance 2.0 and others in Elo score—a user-blind-test-based ranking reflecting real perceived quality. The model’s origin was initially unknown, but technical analysis later linked it to the open-source model "daVinci-MagiHuman," jointly developed by Shanghai SII GAIR Lab and Beijing-based Sand.ai. HappyHorse-1.0, likely an optimized iteration by Sand.ai, uses a 15-billion-parameter transformer architecture for joint audio-video-text modeling. Its strong performance in human-centric scenes (e.g., portraits, narrations) helped it excel in blind tests, though it still lags in multi-character or complex motion scenarios. The achievement signals a potential shift: an open-source model rivaling closed-source alternatives in perceived quality, which could lower costs and increase flexibility for developers in vertical applications like virtual avatars. However, limitations remain, including high computational requirements (H100 GPU needed) and shorter generation lengths. While not yet threatening market leaders, HappyHorse represents progress toward open models reaching "production-ready" quality, potentially accelerating community-driven improvements in the video AI space.

marsbit04/08 07:57

Mysterious Model HappyHorse Tops the Chart Overnight: Is the Video Generation Arena Welcoming a "Game Changer"?

marsbit04/08 07:57

活动图片