# Architecture Related Articles

HTX News Center provides the latest articles and in-depth analysis on "Architecture", covering market trends, project updates, tech developments, and regulatory policies in the crypto industry.

Cross-Chain Bridges Actively Adapt, LI.FI Leverages Intent Architecture to Become the Liquidity Hub for TradFi Institutions

Cross-Chain Bridge LI.FI Transforms with Intents Architecture to Serve as Liquidity Hub for TradFi Institutions Facing declining cross-chain transaction volumes and overall crypto market liquidity, cross-chain bridge protocol LI.FI is proactively shifting its strategy. Moving beyond its role as a "liquidity transfer protocol," LI.FI is targeting new assets, clients, and operational systems. Key to this transformation is the launch of LI.FI Intents, an intent-based execution architecture. This product positions itself as a foundational layer for stablecoin payments, Real World Assets (RWA), and compliant on-chain liquidity, catering specifically to fintech companies, neobanks, wallets, and regulated financial institutions. LI.FI Intents simplifies user experience by offering a turnkey solution. It leverages a solver network for market-maker level execution, enabling precise cross-chain swaps (e.g., between USDC and USDT) without users managing gas tokens or complex blockchain steps. It lowers barriers to entry by integrating with applications like Jumper and Rabby, allowing enterprise users to bypass direct wallet interactions for transactions like payments and asset transfers. The architecture emphasizes compliance. Its network consists of verified legal entities, and enterprises can review and approve orders within their compliance frameworks before processing. All interacting wallets undergo OFAC screening. For ecosystem coverage, LI.FI Intents supports major networks including EVM chains, Solana, and Tron, mitigating risks associated with single-chain dependency. In essence, as tokenized assets like RWAs gain traction, LI.FI Intents focuses on efficiently integrating stablecoin payments and compliant liquidity into enterprise ecosystems. By automating complex execution steps—allowing users to simply declare their intent (the "destination")—it aims to enhance operational efficiency and capital utilization for institutional clients.

Odaily星球日报06/03 06:07

Cross-Chain Bridges Actively Adapt, LI.FI Leverages Intent Architecture to Become the Liquidity Hub for TradFi Institutions

Odaily星球日报06/03 06:07

Xiaomi MiMo's 99% Price Cut is Not Marketing! Luo Fuli Posts on X to Refute Critics

The price of Xiaomi's MiMo-V2.5 series API has been permanently reduced by up to 99%, specifically for the "Input (Cache Hit)" cost, which covers users re-reading historical context in long conversations. MiMo's head, Luo Fuli, published a detailed technical blog to clarify that this drastic price cut stems from genuine engineering breakthroughs, not a marketing stunt or a simple price war. The core of the achievement lies in six key engineering optimizations. First, the model architecture adopts a Hybrid Sliding Window Attention (SWA), reducing the memory footprint (KVCache) to 1/7th of a traditional model. Second, a dual-pool memory management system actually utilizes these savings, allowing a single GPU to handle over 5 times more concurrent users. Third, an upgraded prefix caching mechanism achieves a cache hit rate of 93-95% for repeated reads, meaning most such requests bypass GPU computation entirely. Fourth, a self-developed distributed cache (GCache) utilizes idle SSD space on existing GPU servers, eliminating additional storage costs. Fifth, an intelligent scheduling system (LLM-Router) efficiently routes requests to maximize cache reuse and performance. Sixth, Multi-Token Prediction (MTP) accelerates the model's text generation ("output") side. Together, these systemic optimizations dramatically lower the real computational cost per request, enabling the 99% price reduction for cached inputs while reportedly maintaining positive gross margins. Luo Fuli's disclosure aims to shift the narrative from "price war" to a demonstration of substantive AI engineering progress.

marsbit05/31 10:37

Xiaomi MiMo's 99% Price Cut is Not Marketing! Luo Fuli Posts on X to Refute Critics

marsbit05/31 10:37

Why Did Zhipu Surge Nearly 30% in a Single Day?

"Global AI Model Unicorn" Zhipu's stock surged nearly 30% in a single day, reaching a new market cap high. The catalyst was the launch of its GLM-5.1-highspeed API, boasting a generation speed of **400 tokens per second**, setting a new global benchmark. This speed, roughly 3-5 times faster than industry leaders like OpenAI's GPT-4o and Anthropic's Claude, is achieved **without compromising the full-scale model's capabilities**. In the era of AI Agents requiring dozens of self-calls, such latency reduction is critical, transforming speed from a system metric into a determinant of intelligence limits. The breakthrough stems from a three-layer technical overhaul: 1. **TileRT Inference Engine**: Compiles the entire model into a continuous, always-on computation pipeline using "Warp Specialization," minimizing GPU idle time by having different processor groups handle data loading, computation, and communication in parallel. 2. **Heterogeneous Parallelism for MLA**: To efficiently run the GLM-5.1 model using the MLA attention mechanism, TileRT employs a heterogeneous strategy. One GPU handles sparse indexing/routing, while the others perform dense computation, optimizing for MLA's unique workflow. 3. **ZCube Network Architecture**: Replaces the standard Spine-Leaf (ROFT) network topology with a flat, dual-group interconnect. This design creates a single optimal path between any two GPUs, eliminating network congestion at scale and reducing latency. The business impact is significant: a 15% increase in cluster throughput (free extra capacity), a 40.6% reduction in tail latency (improved stability), and a one-third cut in networking hardware costs. Long-term, this innovation challenges the dominance of NVIDIA's integrated hardware-software stack (GPU+NVLink+InfiniBand), potentially benefiting manufacturers of high-density Leaf switches and optical modules while lowering the software barrier for domestic AI chips like Huawei's Ascend. The innovation proves that more can be achieved with the same compute, reshaping the infrastructure beyond just GPUs.

marsbit05/23 01:23

Why Did Zhipu Surge Nearly 30% in a Single Day?

marsbit05/23 01:23

A Decade's Bet on Cerebras: How the 'Wafer-Scale AI Chip' Reached NASDAQ

"Cerebras, a pioneering AI chip company, successfully debuted on NASDAQ (CBRS) on May 14, 2026, with its stock price surging approximately 68% on the first day. This marks a significant milestone following a decade-long journey, as recounted by early investor Steve Vassallo. The story begins not in 2016, but with the deep, 19-year relationship between Vassallo and founder Andrew Feldman, which started with Feldman’s previous company, SeaMicro (acquired by AMD in 2012). In 2016, Feldman and a core team of chip and system experts sought to challenge the emerging consensus. At a time when AI’s practical utility was still debated and GPUs were becoming the default hardware, they envisioned a fundamentally new computer architecture purpose-built for AI workloads. They identified memory bandwidth, not raw compute power, as the critical bottleneck for neural networks. Defying industry inertia, Cerebras pursued a radical, wafer-scale chip design—58 times larger than the biggest existing chips. This meant confronting and solving a cascade of unprecedented engineering challenges: power delivery, thermal management, and maintaining electrical continuity across tens of thousands of connections. It required reinventing nearly every aspect of modern computing—semiconductors, systems, data structures, software, and algorithms. The path was fraught with setbacks, including a prototype that caught fire on its first power-up. Progress was marked by intense, iterative problem-solving, with the board meeting every 6-8 weeks to tackle the latest technical frontier. Through disciplined perseverance and deep trust within the team, they achieved a breakthrough in August 2019 when their first wafer-scale computer successfully operated. Feldman’s drive for a 1000x leap, his formative upbringing among intellectual giants who modeled both brilliance and kindness, and his belief in building a loyal, mission-driven team were central to Cerebras’s culture. His competitive strategy was that of David vs. Goliath—finding innovative, human-centric approaches that larger incumbents would overlook. From the symbolic delivery of the first term sheet over a backyard fence in 2016 to the NASDAQ bell ringing in 2026, Cerebras’s journey is a testament to long-term vision, technical audacity, and the power of foundational founder-investor relationships. It stands as a reminder that the computing revolution can come not just from more GPUs, but from a complete reimagining of the architecture itself."

marsbit05/15 03:55

A Decade's Bet on Cerebras: How the 'Wafer-Scale AI Chip' Reached NASDAQ

marsbit05/15 03:55

Ant Digital Tech Proposes New Architecture for Agent Economy, Covering Four Layers: Identity, Payment, Risk Control, and Compliance

Ant Digital Technologies (Ant Digital) has introduced a new architectural framework for the agentic economy, named the "4R Full-Stack Architecture," at the Hong Kong Web3 Festival. The framework is designed to address four core challenges in AI agent operations: identity, payment, risk control, and compliance. The four layers include: - **Agentic Runtime**, featuring DTClaw with the CARLI security model to enforce behavioral constraints and ensure controllability and auditability; - **Payment Rails**, which provide on-chain payment channels supporting smart decision-making, verifiable credentials, instant settlement, and cross-chain asset transfers; - **Agent Registry**, leveraging DIDs and the ERC-8004 standard to assign verifiable on-chain identities to agents; - **Root Infrastructure**, built on Jovay Layer2 and ZKVM technology to enable high-speed micro-payments and trusted off-chain computation with on-chain verification. According to CTO Yan Ying, the architecture aims to resolve fundamental gaps in the current agent economy—such as execution vulnerabilities, identity issues, payment barriers, and trust deficits—by redesigning underlying infrastructure rather than applying superficial fixes. The initiative builds on Ant Digital’s extensive experience in financial-grade security, privacy computing, and blockchain.

marsbit04/20 09:24

Ant Digital Tech Proposes New Architecture for Agent Economy, Covering Four Layers: Identity, Payment, Risk Control, and Compliance

marsbit04/20 09:24

a16z Founder: In the Agent Era, What Truly Matters Has Changed

Marc Andreessen, co-founder of a16z, argues that the current AI boom is not an overnight success but the culmination of 80 years of research, now delivering practical results. He emphasizes that this era is defined by the convergence of four key capabilities: large language models (LLMs), reasoning, coding, and agents capable of recursive self-improvement. Andreessen describes the agent architecture—combining an LLM with a shell, file system, markdown, and cron/loop—as a fundamental shift beyond chatbots. This structure leverages existing software components, allowing agents to maintain state, introspect, and extend their own functionality. He predicts a move away from traditional GUI and browser-based interactions toward an "agent-first" world where software is primarily operated by bots, not humans, with people simply stating their goals. He draws parallels to the 2000 internet bubble but notes key differences: current AI infrastructure investments are led by cash-rich giants and quickly monetized. He highlights that scaling constraints involve not just GPUs but the entire chip ecosystem. Open source and edge inference are crucial for democratizing knowledge and enabling low-latency, cost-effective applications on local hardware. Finally, Andreessen identifies significant non-technical challenges: potential short-term cybersecurity crises, the need for "proof of human" identity solutions, financial infrastructure for agents, and institutional resistance from sectors like education and healthcare. He cautions that societal adoption will be slower than technological change.

marsbit04/20 00:02

a16z Founder: In the Agent Era, What Truly Matters Has Changed

marsbit04/20 00:02

Thin Harness, Fat Skills: The True Source of 100x AI Productivity

The article "Thin Harness, Fat Skills: The True Source of 100x AI Productivity" argues that the key to massive productivity gains in AI is not more advanced models, but a superior system architecture. This framework, "fat skills + thin harness," decouples intelligence from execution. Core components are defined: 1. **Skill Files:** Reusable markdown documents that teach a model *how* to perform a process, acting like parameterized function calls. 2. **Harness:** A thin runtime layer that manages the model's execution loop, context, and security, staying minimal and fast. 3. **Resolver:** A context router that loads the correct documentation or skill at the right time, preventing context window pollution. 4. **Latent vs. Deterministic:** A strict separation between tasks requiring AI judgment (latent space) and those needing predictable, repeatable results (deterministic). 5. **Diarization:** The critical process where the model reads all materials on a topic and synthesizes a structured, one-page summary, capturing nuanced intelligence. The architecture prioritizes pushing intelligence into reusable skills and execution into deterministic tools, with a thin harness in between. This allows the system to learn and improve over time, as demonstrated by a YC system that matches startup founders. Skills like `/enrich-founder` and `/match` perform complex analysis and matching that pure embedding searches cannot. A learning loop allows skills to rewrite themselves based on feedback, creating a compound improvement effect without code changes. The conclusion is that 10x to 1000x efficiency gains come from this disciplined system design, not just smarter models. Skills represent permanent upgrades that automatically improve with each new model release.

marsbit04/13 04:19

Thin Harness, Fat Skills: The True Source of 100x AI Productivity

marsbit04/13 04:19

Chaos Labs Exits, Who Will Take Over Aave's Risk?

Chaos Labs, the core risk management provider for Aave V2 and V3 markets, has announced its decision to terminate its partnership with Aave. Despite Aave Labs increasing the budget to $5 million to retain them, Chaos Labs chose to leave due to fundamental disagreements on how risk should be managed. Key reasons for the departure include: the loss of core Aave contributors increasing operational risk, the expanded scope and complexity introduced by Aave V4 (which requires rebuilding risk infrastructure from scratch), and the fact that Chaos Labs operated at a financial loss even with increased budgets. They estimate that proper risk management for both V3 and V4 should cost at least $8 million annually (≈5.6% of protocol revenue), closer to traditional banking standards, rather than the previous 2%. Chaos Labs emphasized that Aave’s reputation and institutional adoption rely heavily on its risk management track record. They also highlighted unquantified costs like legal liability and operational security risks. The exit occurs as Aave plans its V4 upgrade and expands into institutional markets. Chaos Labs warns that migrating to V4 while maintaining V3 will double, not halve, the workload, and that accumulated operational experience cannot be easily transferred. The decision reflects a principled stance: Chaos Labs only attaches its name to work that meets its high-risk standards, even at significant financial sacrifice.

marsbit04/07 03:36

Chaos Labs Exits, Who Will Take Over Aave's Risk?

marsbit04/07 03:36

活动图片