Mysterious "Ox Alpha" Large Model Goes Viral with Limited-Time Free Access

marsbitXuất bản vào 2026-08-23Cập nhật gần nhất vào 2026-08-23

Tóm tắt

A mysterious anonymous AI model named "Ox Alpha," nicknamed "Cow is Coming" by Chinese netizens, has appeared on OpenRouter, sparking widespread speculation. The model offers a 1 million token context, supports text, image, and video inputs, can call tools, and is currently free. Its standout feature is strong coding ability. Initial tests on the DeepSWE benchmark, which evaluates real-world software engineering tasks, showed an 80% pass rate on a subset of tasks, reportedly nearing top-tier code models. However, follow-up tests yielded a 63% score, with variations attributed to different task sets and configurations. The model's true developer is a major topic of debate. The prevailing theory points to Zhipu AI's unreleased GLM-5.3 Flash or its multimodal variant. Evidence cited includes identical visual token consumption patterns with GLM-5V-Turbo for videos, a consistent offset in text token counts compared to GLM-5.3, and similar behavioral traits like refusing audio processing. Zhipu has a precedent of anonymous testing. Simultaneously, another anonymous model, "korrine," appeared on Code Arena, with guesses ranging from Moonshot's Kimi K3.1 to models from Qwen or MiMo, adding to the industry's guessing game. This trend of anonymous "undercover" testing allows for unbiased performance evaluation in platforms like Arena and provides real-world, high-pressure testing through tools like OpenRouter before official release. It also serves as an effective marketing tactic,...

It's trendy to test large models with 'anonymous accounts'.

Recently, an anonymous model named Ox Alpha suddenly appeared on OpenRouter. Ox means 'ox' or 'cow', and domestic netizens quickly gave it a more relatable nickname:

"Ox Comes" large model.

According to information disclosed by OpenRouter, Ox Alpha has a 1 million token context window, supports text, image, and video input, can call tools, and is currently completely free.

Shortly after launch, developers integrated it into a coding agent and threw it into a real code repository for testing.

Preliminary test results show that this mysterious "Ox Comes" large model's capabilities are already approaching current top-tier code models.

Meanwhile, netizens revealed that an anonymous model named korrine is being tested on Code Arena. Some speculate it's Kimi K3.1, while others point to Qwen and MiMo, with varied guesses.

In August, the large model circle has suddenly turned into a large-scale guessing game.

"Ox Comes" Model Performs Remarkably

What truly drew attention to Ox Alpha was its coding capability.

Developer Ben Davis selected 10 tasks from DeepSWE for testing, and Ox Alpha completed 8 of them, achieving an 80% pass rate. In his published comparison results, Fable 5 Max scored 65%, GLM-5.3 Max and Grok 4.6 xhigh both scored 62%, and GPT-5.6 Sol Max scored 52%.

DeepSWE examines real-world software engineering ability. The model needs to read code repositories, locate problems, modify code, run tests, and continue fixing based on error reports. Compared to single-round coding problems, it's closer to the actual work of a coding Agent.

However, 10 tasks is a very small sample. Subsequently, other developers tested on another subset of DeepSWE, reporting a result of about 63%. The task scopes and execution configurations of the two tests were not identical, making it impossible to definitively rank Ox Alpha based on this alone.

Nevertheless, these results preliminarily show that this anonymous model has demonstrated strong potential for long-context coding, with capabilities approaching current leading models.

Who Created the "Ox Comes" Large Model?

The most widespread speculation about the identity of the "Ox Comes" model is that it is an unreleased GLM-5.3 Flash from Zhipu AI, or a multimodal version of GLM-5.3.

Someone even wrote a blog to analyze this:

1. The strongest evidence comes from the video encoder. For four videos with different frame rates, durations, and resolutions, the visual tokens consumed by Ox Alpha matched exactly with GLM-5V-Turbo. MiMo, Qwen, and GLM-4.6V all showed significantly different results.

2. The text tokenizer also shows a high degree of alignment. The researcher tested 25 sets of prompts; the token count between Ox Alpha and GLM-5.3 consistently maintained a fixed difference of 75 tokens.

3. Other features also point to Zhipu AI. Ox Alpha refuses to process audio, which matches the routing method of GLM-5V; its answer style, the number of Agent execution steps, and the inference interface are also very similar to GLM. Zhipu AI previously used Pony Alpha for anonymous testing of GLM-5, establishing a precedent for this practice.

https://ox-alpha-evidence-production.up.railway.app/

Other netizens have also found clues in conversations.

Ben Davis believes he is 99% certain this is GLM-5.x.

These clues increase the credibility of the GLM theory, but are still insufficient for definitive identity confirmation.

As of now, neither OpenRouter nor Zhipu AI have publicly responded.

korrine's Identity is Even More Mysterious

While the identity of the "Ox Comes" model remains unclear, another anonymous model named korrine has appeared on Code Arena.

Initially, many speculated it was Kimi K3.1, because before the release of Kimi K3, it was believed to have been tested under the codename kivine. The similar structure of kivine and korrine sparked this association.

However, the original source of the rumor later added that the previously learned about new Moonshot model might correspond to another codename, adamant-ananke. korrine could also come from other Chinese teams like Qwen.

In the comments, some also pointed to MiMo V3.

Why Do Large Model Companies Like 'Testing in Disguise'?

Anonymous testing is becoming an important step before the official release of large models.

Hiding the manufacturer and model name in the Arena can minimize preconceptions brought by branding. Users cannot see the model's identity and can only choose based on actual outputs. The accumulated battle results are also closer to the real user experience.

OpenRouter provides a different kind of testing environment.

Developers integrate the model into various coding Agents, letting it enter real repositories, continuously call tools, and handle software engineering tasks lasting several hours. Issues like context stability, tool calling reliability, and whether the model gets stuck in loops or goes off track during long tasks can be quickly exposed under intense use.

For model developers, this is akin to a public stress test. Teams can observe failure cases in advance, verify the capacity of their inference services, and also accumulate real-world reputation before official launch.

Moreover, "guessing the model" is increasingly becoming a marketing tactic; the suspense over identity can indeed prolong the discussion cycle.

Finally, back to the model itself. If Ox Alpha is truly a Flash model and its coding ability is already approaching top-tier levels, where will the ceiling be pushed by the more resource-intensive, more capable full version?

Reference Links:

https://x.com/Adidotdev/status/2090833298713096241

https://x.com/davis7/status/2090669483740279155?s=20

https://x.com/davis7/status/2090655207831298095?s=20

https://x.com/MaxForAI/status/2090783750217162788

This article is from the WeChat public account "Almost Human" (ID: almosthuman2014), author: Following AI

Câu hỏi Liên quan

QWhat is the name of the anonymous model discussed in the article that appeared on OpenRouter, and why did it capture significant attention?

AThe anonymous model discussed is named 'Ox Alpha' (nicknamed 'Niu Lai' by Chinese netizens). It captured significant attention primarily due to its impressive performance on real-world software engineering tasks, specifically in coding tests like DeepSWE where it achieved high pass rates, showing capabilities approaching top-tier code models.

QAccording to the article's investigation, which company is the leading candidate for being behind the 'Ox Alpha' model, and what evidence supports this claim?

AThe leading candidate suggested in the article is Zhipu AI, possibly an unreleased GLM-5.3 Flash or a multimodal version of GLM-5.3. Supporting evidence includes: 1) Its video encoder's visual token consumption perfectly matches GLM-5V-Turbo, 2) Its text tokenizer shows a fixed token count offset compared to GLM-5.3, and 3) Other behavioral traits, such as refusing to process audio, align with Zhipu's GLM models. The company also has a precedent for anonymous testing.

QBesides 'Ox Alpha', what is the name of the other anonymous model mentioned in the article that is being tested on Code Arena, and what are some speculations about its origin?

AThe other anonymous model mentioned is named 'korrine', tested on Code Arena. Speculations about its origin include that it could be Moonshot AI's upcoming Kimi K3.1 (due to a similar earlier test codename), a model from Qwen, or MiMo V3. The article states its identity is even more mysterious than that of Ox Alpha.

QWhat are the two main benefits for AI model companies to conduct anonymous testing, as explained in the article?

AThe article explains two main benefits: 1) It allows for unbiased evaluation by removing brand bias, enabling users to judge models based solely on performance, leading to results that better reflect real user experience. 2) It serves as a public stress test, helping companies identify failure cases, verify the capacity of their inference services, and build genuine user reputation before an official launch.

QWhat specific capability of the 'Ox Alpha' model was tested using the DeepSWE benchmark, and what was one of its notable performance results mentioned?

AThe 'Ox Alpha' model was tested on its real-world software engineering capability using the DeepSWE benchmark, which requires tasks like reading code repositories, locating issues, modifying code, and running tests. One notable result mentioned was that in an initial test of 10 tasks by a developer, Ox Alpha completed 8, achieving an 80% pass rate, which was higher than several other top models in the comparison.

Nội dung Liên quan

Phố Wall Nhìn Nhận Ra Sao Về Buổi Trình Diễn Đầu Tay Của Walsh Tại Jackson Hall? Phái Diều Hâu 'Chỉnh Sửa' Truyền Thông Tháng 7, Không Tăng Lãi Suất Tháng 9 Có Thể Lại Làm Tổn Thương Uy Tín Của Fed

Chủ tịch Fed Christopher Waller trong bài phát biểu đầu tiên tại hội nghị Jackson Hole đã được giới Phố Wall đánh giá là một sự "điều chỉnh" theo hướng diều hậu đối với thông điệp sau cuộc họp FOMC tháng 7. Bài phát biểu nhấn mạnh mục tiêu lạm phát 2% là bất di bất dịch, cho rằng điều kiện tài chính hiện tại không thực sự hạn chế, và số liệu lạm phát gần đây chưa đủ chứng minh xu hướng cơ bản được cải thiện rõ rệt. Waller nói rõ nếu không tin chắc lạm phát đang giảm đủ nhanh, Fed "vẫn còn việc phải làm". Nhiều chuyên gia như Priya Misra (JP Morgan) nhận định đây là một bài phát biểu diều hâu, nhằm khôi phục uy tín chống lạm phát và sửa chữa "sai sót truyền thông" từ tháng 7. Aberdeen cảnh báo nếu Fed không tăng lãi suất vào tháng 9, uy tín có thể bị tổn hại thêm. Barclays và Société Générale dự báo Fed sẽ tăng lãi suất 25 điểm cơ bản vào cả tháng 9 và tháng 12. Thị trường nhanh chóng định giá lại khả năng tăng lãi suất gần hơn, với xác suất tăng vào tháng 9 tăng từ khoảng 35% lên khoảng 50-60%. Lợi tức trái phiếu kho bạc kỳ hạn ngắn tăng mạnh, phản ánh kỳ vọng này. Tuy nhiên, một số ý kiến cho rằng Waller vẫn chưa đưa ra lộ trình chính sách cụ thể hay cam kết hành động vào tháng 9, mà chỉ đưa ra các nguyên tắc định hướng, từ chối cung cấp "hàm phản ứng" rõ ràng. Có nhận xét ví ông "cho la bàn chứ không cho GPS". Đồng thuận chung là Waller đã thiết lập lại một logic chính sách diều hâu hơn: nếu kinh tế và thị trường lao động vẫn mạnh mẽ trong khi lạm phát cơ bản không giảm đủ nhanh, Fed vẫn có thể tiếp tục thắt chặt. Việc có tăng lãi suất vào tháng 9 hay không sẽ phụ thuộc vào dữ liệu lạm phát và việc làm sắp tới. Thách thức thực sự với Waller là liệu ông có sẵn sàng chuyển chẩn đoán diều hâu thành hành động thực tế nếu dữ liệu không cải thiện.

marsbit29 phút trước

Phố Wall Nhìn Nhận Ra Sao Về Buổi Trình Diễn Đầu Tay Của Walsh Tại Jackson Hall? Phái Diều Hâu 'Chỉnh Sửa' Truyền Thông Tháng 7, Không Tăng Lãi Suất Tháng 9 Có Thể Lại Làm Tổn Thương Uy Tín Của Fed

marsbit29 phút trước

Bài phát biểu mới nhất của Warsh: Thời đại của chúng ta

Bài phát biểu của Chủ tịch Fed Kevin Walsh tại Hội nghị Jackson Hole mang tín hiệu thận trọng thiên hướng diều hâu. Walsh nhấn mạnh nền kinh tế và thị trường lao động Mỹ vẫn mạnh mẽ, môi trường tài chính chưa thực sự hạn chế, và lạm phát vẫn cao hơn đáng kể so với mục tiêu 2%. Do đó, kiểm soát giá cả phải tiếp tục là ưu tiên hàng đầu của chính sách tiền tệ. Ông nêu rõ tiêu chuẩn của mình: Fed cần có sự tin tưởng rằng lạm phát cơ bản đang tiến về mục tiêu một cách rõ ràng và đủ nhanh; nếu không, vẫn còn công việc phải làm. Dù dữ liệu CPI và PCE mùa hè tích cực, chúng không khiến ông tin rằng xu hướng lạm phát cơ bản đã được cải thiện một cách có ý nghĩa. Walsh cũng trình bày quan điểm về "hướng dẫn triển vọng". Ông cho rằng trong thời kỳ bình thường, việc này nên được hạn chế, vì cam kết sớm về lộ trình lãi suất có thể gây hiểu lầm, hạn chế không gian linh hoạt của Fed và tạo ra "vấn đề sảnh gương", nơi thị trường và Fed có thể bỏ qua những thay đổi kinh tế mới. Bài phát biểu đã tác động đến thị trường: kỳ vọng về khả năng tăng lãi suất của Fed vào tháng 9 tăng lên gần 60%, và giá vàng giảm mạnh. Trong phần kết luận, Walsh cam kết một cách tiếp cận có kỷ luật, tập trung vào dữ liệu và các nguyên tắc vững chắc để thực hiện nhiệm vụ kép của Fed là ổn định giá cả và tối đa hóa việc làm.

marsbit2 giờ trước

Bài phát biểu mới nhất của Warsh: Thời đại của chúng ta

marsbit2 giờ trước

Truy cập vào các phiên hoạt động thay vì cơ sở dữ liệu: Thị trường chợ đen ở Nga đã thay đổi như thế nào

Thị trường ngầm ở Nga đang chuyển hướng từ bán cơ sở dữ liệu khổng lồ sang buôn bán quyền truy cập ngắn hạn, tích cực vào tài khoản người dùng. Giá thông tin bị đánh cắp bởi phần mềm độc hãi từ thiết bị nhiễm bệnh đã tăng gần 13% trong năm qua. Các gói dữ liệu mới chứa token phiên hoạt động, mật khẩu hiện hành, thông tin đăng nhập VPN, kho lưu trữ đám mây và quyền truy cập mạng công ty, cho phép xâm nhập tức thì mà có thể bỏ qua xác thực hai yếu tố. Trong khi đó, số lượng các vụ rò rỉ cơ sở dữ liệu lớn mới đã giảm sau các biện pháp trừng phạt nghiêm khắc của chính phủ. Tuy nhiên, các cơ chế quản lý hiện tại tỏ ra không hiệu quả trước mô hình đe dọa mới này, vì phần mềm độc hại đánh cắp dữ liệu từ thiết bị cá nhân sau khi chúng rời khỏi hệ thống công ty, mà không vi phạm trực tiếp đến cơ sở dữ liệu tập trung. Các chuyên gia chỉ ra rằng mô hình này lặp lại kịch bản tội phạm mạng toàn cầu, chẳng hạn như thị trường Genesis Market đã bị đóng cửa. Tình hình nhấn mạnh lỗ hổng nằm giữa thiết bị cá nhân và hệ thống công ty, nơi giá trị của token ngắn hạn ngày càng tăng nhưng lại nằm ngoài phạm vi kiểm soát của các quy định bảo vệ hiện hành.

cryptonews.ru5 giờ trước

Truy cập vào các phiên hoạt động thay vì cơ sở dữ liệu: Thị trường chợ đen ở Nga đã thay đổi như thế nào

cryptonews.ru5 giờ trước

Bitfinex Báo Hiệu Khởi Đầu Xu Hướng Tăng Giá Bitcoin Trong Bối Cảnh Mức Tương Quan Đạt Đỉnh Với Vàng

Các nhà phân tích từ sàn Bitfinex chỉ ra rằng Bitcoin đang bắt đầu hành xử như "vàng kỹ thuật số" khi nhà đầu tư tìm kiếm công cụ phòng ngừa rủi ro trước nợ và thao túng tiền tệ. Mối tương quan giữa giá Bitcoin và vàng đã đạt đến mức đỉnh hiếm khi duy trì lâu dài, báo hiệu một sự thay đổi xu hướng sắp tới. Cả hai tài sản đều được giao dịch như một phương tiện phòng ngừa sự mất giá tiền tệ, với Bitcoin là phiên bản có hệ số beta cao hơn. Bitfinex nhấn mạnh rằng bối cảnh giảm rủi ro tiềm năng sẽ cho thấy liệu Bitcoin sẽ giữ vững cùng vàng hay giảm theo cổ phiếu. Báo cáo cũng ghi nhận sự thay đổi môi trường vĩ mô, từ cơn sốt AI sang giao dịch phòng ngừa mất giá, do những diễn biến gần đây liên quan đến nợ công Mỹ. Đồng thời, Bitfinex xác định Bitcoin đã thoát khỏi pha tích lũy và bước vào pha tăng trưởng, với chỉ số Delta-Thermo Market Multiple là 2.03, tiến gần ngưỡng 2.5x đánh dấu bắt đầu pha tăng. Tình huống hiện tại được so sánh với cấu hình tương tự năm 2024, liên quan đến lo ngại về "xói mòn nợ" và sự đa dạng hóa rộng rãi gây bất lợi cho đồng USD. Tuy nhiên, Chủ tịch Fed Kevin Warsh gần đây đã bác bỏ lập luận này, giữ lập trường "diều hâu" và ám chỉ khả năng tăng lãi suất.

cryptonews.ru5 giờ trước

Bitfinex Báo Hiệu Khởi Đầu Xu Hướng Tăng Giá Bitcoin Trong Bối Cảnh Mức Tương Quan Đạt Đỉnh Với Vàng

cryptonews.ru5 giờ trước

Giao dịch

Giao ngay
活动图片