# Generation Related Articles

HTX News Center provides the latest articles and in-depth analysis on "Generation", covering market trends, project updates, tech developments, and regulatory policies in the crypto industry.

1600 Lines of Code Create an Underwater Manhattan, Fable 5 Leaves Karpathy Stunned

Title: Fable 5 Stuns Karpathy with 3D Worlds Built from 1600 Lines of Code An AI model, Fable 5, has demonstrated a remarkable leap in generating complex, interactive 3D worlds with minimal code. Showcased by Peter Gostev of Arena.ai, the model created 63 diverse 3D environments across themes like immersive cityscapes, explorable famous paintings, natural wonders, and cosmic phenomena. Many were generated in a single attempt. A standout creation is a detailed, submerged Manhattan built with only 1600 lines of Three.js code. Other highlights include traversable versions of Van Gogh's "Starry Night," a bear catching a salmon with realistic physics, and a split Red Sea. The model's ability to cohesively manage vast numbers of elements within a scene represents a significant technical advancement. Andrej Karpathy, who recently joined Anthropic's pre-training team, expressed amazement, particularly at how the model intuitively understood and rendered complex real-world interactions like a fish struggling when caught. He coined the term "fablemaxxing" to describe this qualitative leap. While Fable 5 excels at world-building, Gostev notes current limitations in creating engaging, long-form gameplay. The model also sometimes requires prompting to fully utilize its capabilities. Having topped the Agent Arena benchmark for real-world task completion, Fable 5 signals that the boundaries of AI-generated content are rapidly expanding, with its full potential yet to be discovered.

marsbit07/06 09:48

1600 Lines of Code Create an Underwater Manhattan, Fable 5 Leaves Karpathy Stunned

marsbit07/06 09:48

He Kaiming's Team's New Work: After Deleting VAE and Private Data, Text-to-Image Generation Becomes Even Stronger

KaiMing He's team introduces **MiniT2I**, a minimalist text-to-image (T2I) model that challenges the complexity of mainstream approaches. It eliminates components commonly considered essential: the VAE encoder-decoder, AdaLN conditioning mechanisms, auxiliary losses, private training data, and post-training alignment stages like RL/DPO. Instead, it uses a pure flow-matching objective trained directly on RGB pixels. The model employs a simplified **MM-JiT** Transformer architecture. It removes AdaLN blocks for conditioning and instead prepends two lightweight text adapter blocks to a standard pre-norm Transformer, allowing frozen T5 text features to adapt to the denoiser. Training follows a two-stage, LLM-like paradigm using only public datasets: pre-training on LLaVA-recaptioned CC12M for coverage, followed by fine-tuning on ~120k high-quality image-text pairs. With just 258M parameters (B/16), MiniT2I achieves competitive scores (0.87 on GenEval, 84.2 on DPG-Bench), outperforming larger pixel-space models. Scaling to 912M parameters (L/16) yields results comparable to SD3-Medium (~2B parameters) in style, composition, and imagination, though it lags in text rendering and named entities due to public data limitations. Key advantages include lower computational cost (~570 GFLOPs vs. ~1379 for latent models) and architectural simplicity. Acknowledged limitations include patch boundary artifacts in pixel space, side effects of high CFG scales, resolution ceilings for sequences longer than 1024 tokens, and the aforementioned data bottlenecks. The work demonstrates that high-performance T2I generation is possible with a radically simplified, publicly reproducible baseline.

marsbit06/22 10:17

He Kaiming's Team's New Work: After Deleting VAE and Private Data, Text-to-Image Generation Becomes Even Stronger

marsbit06/22 10:17

The Creator of Kling Returns to Alibaba and Builds Another Dark Horse

The article discusses the rise of HappyHorse-1.0, an AI video generation model developed by Alibaba, which topped the Artificial Analysis leaderboard in both text-to-video and image-to-video categories in April 2026. The model was created under the leadership of Zhang Di, who returned to Alibaba in November 2025 after working at Kuaishou, where he led the development of the Kling model. HappyHorse is open-source and commercially available, similar to Alibaba's Qwen model. Zhang Di's background includes extensive experience in large-scale data systems and machine learning at Alibaba and Kuaishou, which contributed to the rapid development of HappyHorse within just five months. The model uses a 15-billion-parameter transformer architecture with native multimodal training, supporting multiple languages and lip-sync capabilities. It also focuses on reducing inference time and cost, making it practical for commercial use. The primary application of HappyHorse is in e-commerce, where it can generate product videos to enhance user engagement and conversion rates by creating contextual and personalized content. This aligns with Alibaba's strengths in commerce, advertising, and data feedback loops. The model's success with open-source approach contrasts with challenges faced by closed-source models like OpenAI's Sora (shut down due to high costs) and ByteDance's Seedance 2.0 (paused over copyright issues). HappyHorse represents a strategic move for Alibaba to integrate AI video generation into its core business ecosystems.

marsbit04/13 05:10

The Creator of Kling Returns to Alibaba and Builds Another Dark Horse

marsbit04/13 05:10

Claiming the "Happy Horse": Alibaba's AI Lays Out the "Eight Trigrams Formation"

Alibaba has officially claimed the "HappyHorse" (HappyHorse-1.0) AI video generation model, which recently topped the global benchmark on Artificial Analysis with an Elo score of 1357. Developed by Alibaba’s ATH (Alibaba Token Hub) innovation unit, the model is notable for its ability to generate high-definition video with synchronized audio and sound effects from text input, significantly improving motion coherence and reducing production time and cost. This launch is part of a broader acceleration in Alibaba’s AI strategy. In late March and early April, the company released three flagship models in quick succession: Qwen3.5-Omni, Wan2.7-Image, and Qwen3.6-Plus. The latter broke global daily call volume records with 1.4 trillion tokens processed shortly after release. Alibaba has also undergone significant organizational restructuring to support its AI ambitions. In March, it established the ATH business group, led by CEO Wu Yongming, to integrate AI development, cloud services, and application deployment. Further changes in April included forming a group-level technology committee and consolidating the Tongyi Lab into a dedicated AI model division. The company is investing heavily in AI, with plans to spend over 380 billion RMB on cloud and AI infrastructure over three years. Its self-developed GPUs have already seen mass production. While the market has responded positively to these moves, challenges remain in balancing centralized control with operational flexibility and maintaining team stability amid rapid changes.

marsbit04/11 04:07

Claiming the "Happy Horse": Alibaba's AI Lays Out the "Eight Trigrams Formation"

marsbit04/11 04:07

活动图片