Exploring Physical World AGI with "Visual Reasoning", ElorianAI Raises $55 Million

marsbitDipublikasikan tanggal 2026-04-23Terakhir diperbarui pada 2026-04-23

Abstrak

ElorianAI, co-founded by ex-Google AI expert Andrew Dai and former AI specialist Yinfei Yang, has raised $55 million in early funding to develop next-generation AI systems with advanced visual reasoning capabilities. While current large models excel in text-based tasks like programming and math, they perform poorly in visual reasoning—even top models like Gemini only match a 3-year-old’s ability in basic visual benchmarks. The key limitation lies in the architecture of current vision-language models (VLMs), which first convert visual inputs into text before reasoning, losing critical spatial and structural information. ElorianAI aims to build a native multimodal model that processes and reasons directly in visual space, enabling deeper understanding of physical relationships, constraints, and environments. The company plans to release a state-of-the-art visual reasoning model by 2026, with potential applications in robotics, disaster management, engineering, healthcare, and AI hardware. By using high-quality, diverse, and synthetically generated data, ElorianAI intends to create models that don’t just perceive but truly understand and reason about the physical world—bringing us closer to visual AGI.

By Alpha Community

AI large models have surpassed average humans in certain areas, such as programming and mathematics. Reports indicate that Anthropic has almost achieved 100% AI programming internally, and Google's Gemini Deep Think solved 5 out of 6 problems in IMO 2025, reaching gold medal level.

However, in visual reasoning, even the leading Gemini 3 Pro only reached the level of a 3-year-old child on BabyVision, a benchmark testing basic visual reasoning abilities.

Why are large models strong in programming and mathematics but weak in visual reasoning? This is due to limitations in their "thinking process." Visual Language Models (VLMs) need to first convert visual input into language and then perform text-based reasoning. However, many visual tasks cannot be accurately described in words, resulting in poor visual reasoning capabilities of the models.

Andrew Dai, who worked at Google DeepMind for 14 years, teamed up with Apple's seasoned AI expert Yinfei Yang to establish a company called Elorian AI. Their goal is to elevate the model's visual reasoning ability from "child level" to "adult level," enabling the model to natively "think" within the "visual space" and thereby advance toward AGI in the physical world.

Elorian AI raised $55 million in early-stage funding co-led by Striker Venture Partners, Menlo Ventures, and Altimeter, with participation from 49 Palms and top AI scientists including Jeff Dean.

Pioneers in Multimodal Models Aim to Equip Visual Models with Reasoning Abilities

Andrew Dai, who is of Chinese descent, holds a bachelor's degree in computer science from Cambridge and a PhD in machine learning from Edinburgh. He interned at Google during his PhD and joined the company in 2012, staying for 14 years until starting his own business.


Image Source: Andrew Dai's LinkedIn

Shortly after joining Google, he co-authored the first paper on language model pre-training and supervised fine-tuning, "Semi-supervised Sequence Learning," with Quoc V. Le. This paper laid the foundation for the birth of GPT. Another foundational paper of his is "Glam: Efficient scaling of language models with mixture-of-experts," which paved the way for the now mainstream MoE architecture.

Image Source: Google

During his time at Google, he was deeply involved in almost all large model trainings, from Palm to Gemini 1.5 and Gemini 2.5. Under Jeff Dean's arrangement, he began leading the data division of Gemini (including synthetic data) in 2023, and the team later expanded to hundreds of people.

Image Source: Yinfei Yang's LinkedIn

Co-founding Elorian AI with Andrew Dai is Yinfei Yang, who worked at Google Research for four years, focusing on multimodal representation learning, before joining Apple to lead multimodal model R&D.

Image Source: arxiv

His representative research, "Scaling up visual and vision-language representation learning with noisy text supervision," advanced the development of multimodal representation learning.

Elorian AI's co-founders also include Seth Neel, who was an Assistant Professor at Harvard University and is an expert in data and AI.

Why discuss the groundbreaking papers written by Elorian AI's co-founders? Because their goal is not just engineering optimization but a paradigm shift at the foundational architecture level, upgrading AI from text-based intelligent understanding to vision-based intelligent understanding.

The current state of AI models is that, despite excelling in text-based tasks, even the most advanced frontier multimodal large models still stumble on the most basic visual grounding tasks.

For example, how to fit a part precisely into a mechanical device to make it run more accurately and efficiently? Such spatial physical tasks are simple for elementary school students but challenging for existing multimodal large models.

This brings us back to biology for clues. In the human brain, vision is the underlying substrate supporting many thinking processes. Humans' ability to use visual and spatial reasoning is far more ancient than language-based logical reasoning.

For instance, teaching someone to navigate a maze using language can be confusing, but drawing a sketch makes it instantly understandable.

Even a bird, without language, can recognize and reason about geographical features through vision to achieve global long-distance migration. This is a strong signal that vision is likely the correct direction for truly advancing machine reasoning.

So, imagine if, from the very beginning of model construction, this biological visual instinct is encoded into AI's genes, building a native multimodal model that "simultaneously understands and processes text, images, video, and audio," enabling the model to possess visual understanding capabilities. Andrew Dai and his team aim to build an innate "synesthete," teaching machines not only to "see" the world but also to "understand" it.

To Andrew Dai and his team, a deep understanding of the real "physical world" is the key to achieving the next leap in machine intelligence and ultimately reaching "Visual AGI."

VLMs with Post-Reasoning Are Not the Right Path to Visual Reasoning

There have been teams attempting this before. In fact, Andrew Dai's previous Gemini team was already among the global leaders in the multimodal field. However, traditional multimodal models are still primarily VLMs (Visual Language Models), built on a "two-step" logic: first converting visual input into language, then performing text-based reasoning (sometimes assisted by external tools).

However, post-reasoning inherently has limitations. On one hand, it is prone to model hallucinations; on the other, many visual tasks cannot be precisely described in words.

Additionally, visual generation models like NanoBanana excel in multimodal generation, but generation ability does not equal reasoning ability. The "thinking" before generation still relies on language models, not native reasoning capability.

To develop models that truly understand the spatial, structural, and relational complexities of the visual world, disruptive innovation at the underlying technology level is necessary.

So, how to innovate? Elorian AI's founders, with years of experience in the multimodal field, approach this by deeply integrating multimodal training with a new architecture specifically designed for multimodal reasoning. They abandon the traditional approach of treating images as static input, instead training models to directly interact with and manipulate visual representations to autonomously parse their structure, relationships, and physical constraints.

Of course, another core element is data, which is crucial to the performance and success of these models.

Andrew Dai stated that they place great importance on data quality, data mix ratios, data sources, and data diversity. They have innovated at the data layer, reconstructing the reasoning chain in visual space, and are extensively and deeply using synthetic data.

Combined, these efforts will give rise to new AI systems that move beyond simple visual "perception" to high-level visual "reasoning."

This AI system could be a visual reasoning foundation model: building a highly general but exceptionally proficient model in a specific capability set—visual reasoning.

As a general foundation model, its application areas should be broad.

First, in the robotics field, it could become the underlying neural center of powerful systems,赋予ing them the ability to operate autonomously in various unfamiliar environments.

For example, sending a robot to handle a sudden safety fault in a hazardous environment requires the robot to make quick and accurate instant decisions. If the robot lacks a foundation model with deep reasoning capabilities, people wouldn't dare let it randomly press buttons or operate levers. But if it has strong reasoning能力, it might think: "Before operating this panel, maybe I should pull this lever first to activate the safety mechanism."

Furthermore, in disaster management, models with visual reasoning could analyze satellite images to monitor and prevent forest fires. In engineering, they could accurately understand complex visual blueprints and system diagrams. The significance of this ability lies in the fact that the operating principles of the physical world are fundamentally different from the pure code world. You can't design an airplane wing just by typing a few lines of pure code.

However, Elorian AI's models and capabilities are currently still on paper. They plan to release a model in 2026 that achieves SOTA level in visual reasoning. At that time, we can verify if their results match their claims.

When AI Truly Possesses "Visual Reasoning" Ability, How Will It Change the Physical World?

To enable AI to understand and influence the real physical world, technology has iterated several times.

From image recognition in the traditional CV era, to image generation models/multimodal models in generative AI, to world models, the understanding of the physical world has been continuously enhanced.

Visual reasoning foundation models could take it a step further. Because achieving visual reasoning allows AI to understand the physical world more deeply, thereby achieving a higher level of machine intelligence.

Imagine, when models with deep understanding and fine operation empower the embodied intelligence industry and the AI hardware industry, it will greatly expand their application scope. For example, robots could perform more reliable industrial production or work in medical care; AI hardware, especially wearable devices, could become smarter personal assistants.

However, underlying these technologies is still data. As Andrew Dai mentioned earlier, data quality, data mix ratios, data sources, and data diversity all determine model performance.

In the physical AI field, Chinese companies, whether at the model level or the data level, are closer to world leadership compared to text large models. If they can leverage their advantages of richer data and application scenarios to accelerate iteration speed, then whether in embodied intelligence or AI hardware, whether applied in industry, healthcare, or homes, there is a greater opportunity to reach leading levels and potentially produce world-class enterprises.

Pertanyaan Terkait

QWhat is the main goal of current Vision Language Models (VLMs) according to the article, and what are their limitations?

AThe main goal of VLMs is to process visual input by first converting it into language and then performing text-based reasoning. Their limitation is that many visual tasks cannot be accurately described with text, leading to poor visual reasoning capabilities.

QWho are the founders of Elorian AI and what are their backgrounds?

AThe founders are Andrew Dai, a former Google DeepMind researcher with 14 years of experience, and Yinfei Yang, an AI expert who worked at Google Research and Apple. Andrew Dai contributed to foundational papers in language model pre-training and MoE architecture, while Yinfei Yang focused on multimodal representation learning.

QHow does Elorian AI plan to improve AI's visual reasoning capabilities?

AElorian AI aims to develop a native multimodal model that processes text, images, video, and audio simultaneously. They focus on integrating multimodal training with new architectures designed for visual reasoning, directly interacting with visual representations to parse structures and physical constraints, and using high-quality, diverse synthetic data.

QWhat potential applications are mentioned for AI with advanced visual reasoning skills?

AApplications include robotics for autonomous operations in unfamiliar environments, disaster management through satellite image analysis, engineering by interpreting complex visual diagrams, and enhancing AI hardware like wearable devices for personal assistance.

QWhen does Elorian AI plan to release their model, and what is the expected achievement?

AElorian AI plans to release a model in 2026 that achieves state-of-the-art (SOTA) performance in visual reasoning, aiming to elevate capabilities from 'child-level' to 'adult-level'.

Bacaan Terkait

Sinyal Pasar Bawah Sejarah Terulang? Messari yang Bernilai 3 Miliar Terjual Murah 10 Juta

Penanda Sinyal Sejarah Terulang? Messari yang Pernah Bernilai $300 Juta Terjual Murah $10 Juta Messari, yang pernah menjadi platform data kripto terdekat dengan Bloomberg dan bernilai $300 juta pada 2022, baru-baru ini dijual kepada pesaingnya Blockworks dengan harga hanya sekitar $10 juta. Penurunan dramatis ini bukan hanya kisah satu perusahaan. AI menjadi ancaman struktural bagi bisnis inti Messari, yaitu laporan penelitian dan data. Di saat yang sama, industri secara keseluruhan mengalami kontraksi mendalam: platform data seperti DappRadar dan Parsec tutup, media seperti CoinDesk dijual murah, dan perusahaan seperti Dune melakukan PHK besar-besaran. Venture Capital (VC) juga mundur. Investasi VC kripto anjlok lebih dari 80% dalam 6 bulan, dana dan perhatian dialihkan ke AI. Banyak firma VC kripto legendaris berhenti beroperasi atau memperluas fokus ke luar kripto, dengan salah satu mitra Dragonfly Capital menyebut situasi saat ini sebagai "kepunahan massal". Namun, kondisi suram ini mungkin justru merupakan sinyal bottom historis. Indeks Fear & Greed Kripto telah lama berada di zona "ketakutan ekstrem", mirip dengan momen sebelum siklus bull sebelumnya. Pemegang jangka panjang Bitcoin mengontrol hampir 80% pasokan yang beredar, pola yang biasanya terlihat di area bawah pasar. Aktivitas VC yang rendah juga mirip dengan periode sebelum ledakan DeFi pada 2020. Beberapa pelaku pasar justru melihat peluang dalam keputusasaan ini. Dragonfly Capital baru saja mengumpulkan dana baru senilai $650 juta, sementara Blockworks, yang membeli Messari, dengan sengaja mengkonsolidasi industri data kripto. Sejarah menunjukkan bahwa titik balik besar sering kali dimulai ketika semuanya terlihat paling suram.

marsbit34m yang lalu

Sinyal Pasar Bawah Sejarah Terulang? Messari yang Bernilai 3 Miliar Terjual Murah 10 Juta

marsbit34m yang lalu

Volume Pengiriman TPU Google Dinaikkan 50%

Kebutuhan AI akan daya komputasi terus meningkat, dan pasar sedang mengalami koreksi ekspektasi penting. Beberapa lembaga luar negeri secara diam-diam telah merevisi perkiraan pengiriman Google TPU ke atas. Dari perkiraan awal 10 juta unit untuk tahun 2027, penelitian industri terbaru menunjukkan angka ini dapat dinaikkan menjadi 15 juta unit, peningkatan 50%. Peningkatan signifikan dalam volume TPU ini akan langsung mengalir ke seluruh rantai pasokan. Skema interkoneksi optik penuh yang standar pada kluster TPU Google berarti kebutuhan perangkat keras pendukungnya kaku dan hampir tetap, termasuk NPO Optical Engine (yang dipasangkan 1:1 dengan TPU), modul optik 1.6T, OCS Optical Switch, catu daya server, serta kabel serat optik dan konektor MPO. Setiap kenaikan ekspektasi pengiriman TPU akan mendorong ekspektasi kinerja seluruh rantai ini. Di antara segmen-segmen ini, pendinginan cair (*liquid cooling*) muncul sebagai arah inti dengan perubahan terbesar. Dengan peningkatan daya chip TPU generasi baru, solusi pendingin tradisional sudah tidak memadai. Tahun 2026 diprediksi menjadi tahun peluncuran besar-besaran pendinginan cair untuk Google. Selain akselerasi kinerja, pola persaingan global sedang berubah. Produsen luar negeri menghadapi kendala kapasitas dan teknologi, membuka peluang bagi produsen domestik China untuk masuk ke rantai pasokan inti Google dengan keunggulan iterasi cepat dan kapasitas memadai. Pasar pendinginan cair khusus Google diproyeksikan melonjak dari level ratusan miliar menjadi 300 miliar dalam dua tahun, didorong oleh peningkatan volume dan nilai per unit. Logika di pasar serat optik juga diperbarui. Kebutuhan akan interkoneksi padat di pusat data AI telah mengubahnya dari komoditas siklus menjadi sumber daya strategis. Siklus ekspansi yang panjang untuk preform serat optik (18-24 bulan) menciptakan ketidakseimbangan pasokan dan permintaan global. Vendor cloud besar seperti Google mengamankan pasokan jangka panjang. Produsen serat optik China, dengan keunggulan kapasitas dan biaya, diperkirakan akan mendominasi hampir setengah dari permintaan global untuk serat AIDC pada 2026. Selain itu, peningkatan pasokan TPU juga mendorong pemulihan di segmen pendukung lain seperti modul optik 1.6T dan catu daya server bertegangan tinggi, di mana produsen China juga mendapat peluang substitusi. Intinya, fokus investasi beralih dari spekulasi chip ke pertumbuhan pasti infrastruktur pendukung komputasi. Revisi besar pasokan TPU Google ini mengunci visibilitas kinerja untuk dua tahun ke depan di seluruh rantai industri.

marsbit1j yang lalu

Volume Pengiriman TPU Google Dinaikkan 50%

marsbit1j yang lalu

Setelah Cerita Dunia Kripto Meredup, Apa yang Sebenarnya Diinginkan Wall Street?

Penulis: Blockchain in Plain Language Pada musim gugur 2008, saat Lehman Brothers runtuh, Satoshi Nakamoto meluncurkan Bitcoin dengan pesan sindiran terhadap sistem keuangan tradisional. Setelah 17 tahun, situasinya berbalik. Alih-alih mengadopsi narasi spekulatif "decentralized", Wall Street kini membangun infrastruktur keuangan tradisional yang dikendalikan dan sesuai regulasi di atas teknologi blockchain. Inti transformasi ini adalah tokenisasi aset dunia nyata. Dana BUIDL milik BlackRock, yang di-backing oleh surat utang pemerintah AS jangka pendek, telah menjadi aset dasar yang stabil di blockchain. Securitize, mitra BlackRock, akan melantai di bursa dengan valuasi $1,25 miliar dan bekerja sama dengan NYSE untuk membangun sistem penyelesaian perdagangan saham 24/7 berbasis blockchain. Wall Street juga mengemas volatilitas Bitcoin menjadi produk berpenghasilan tetap. ETF Bitcoin Premium Income (BITA) milik BlackRock akan menghasilkan pendapatan dengan menjual opsi call, mengubah Bitcoin menjadi aset yang membayar dividen bulanan. Dalam pembayaran, stablecoin kini difokuskan sebagai alat transaksi yang efisien. Stripe dan Mastercard mengintegrasikan stablecoin untuk penyelesaian pembayaran lintas batas yang instan, sementara regulasi GENIUS Act 2025 memastikan stablecoin tetap sebagai alat bayar yang tidak membayar bunga dan diawasi ketat. Kesimpulannya, Wall Street tidak lagi mengejar cerita spekulatif cryptocurrency. Mereka membangun pipa keuangan baru di blockchain yang menghasilkan pendapatan, terkendali, dan compliant—mengadopsi teknologi untuk memperkuat, bukan mengganti, sistem keuangan tradisional.

marsbit1j yang lalu

Setelah Cerita Dunia Kripto Meredup, Apa yang Sebenarnya Diinginkan Wall Street?

marsbit1j yang lalu

Terikat pada SpaceX, Jalan Cursor Menuju Kebangkitan dengan Nilai US$600 Miliar

**Ringkasan Bahasa Indonesia: Kisah Cursor, Startup AI Pemrograman yang Meroket dan Masa Depannya di Bawah Elon Musk** Pada 2019, Michael Truell, mahasiswa MIT berusia 18 tahun, menunjukkan bakat pemrograman luar biasa. Beberapa tahun kemudian, ia dan beberapa temannya mendirikan Anysphere dan meluncurkan Cursor, sebuah editor kode berbasis AI. Pada akhir 2025, Cursor digunakan jutaan pengembang dengan pendapatan melampaui $10 miliar, tumbuh 10 kali lipat dalam kurang dari setahun. Pertumbuhan pesat Cursor diiringi kontroversi, seperti proses rekrutmen yang ketat dengan "percobaan kerja" tidak berbayar selama berhari-hari hingga berminggu-minggu. Tantangan struktural utama adalah ketergantungannya pada model AI dari Anthropic (Claude). Ketika Anthropic meluncurkan alat pemrograman saingannya, Claude Code, Cursor merasa terancam dan memulai proyek darurat untuk mengembangkan model AI sendiri bernama Composer. Untuk mengatasi kebutuhan komputasi mahal dalam pengembangan model, Cursor menjalin kerja sama strategis dengan SpaceX milik Elon Musk. Cursor mendapatkan akses ke sumber daya komputasi masif SpaceX, sementara Grok (AI milik Musk/xAI) mendapatkan data pelatihan dari Cursor untuk meningkatkan kemampuan pemrogramannya. Kerja sama ini juga mencakup perjanjian akuisisi potensial senilai $600 miliar oleh SpaceX di kemudian hari. Inti cerita ini adalah pertanyaan tentang masa depan Cursor: akankah menjadi perusahaan perangkat lunak generasi berikutnya yang mandiri, atau hanya menjadi bagian dalam persaingan raksasa AI? Saat ini, dengan 700 karyawan dan melayani 60% perusahaan Fortune 500, Cursor terus tumbuh sambil menavigasi hubungan kompleks dengan mitra model AI dan komitmen besar dengan Elon Musk.

marsbit1j yang lalu

Terikat pada SpaceX, Jalan Cursor Menuju Kebangkitan dengan Nilai US$600 Miliar

marsbit1j yang lalu

Trading

Spot
Futures

Artikel Populer

Cara Membeli AR

Selamat datang di HTX.com! Kami telah membuat pembelian Arweave (AR) menjadi mudah dan nyaman. Ikuti panduan langkah demi langkah kami untuk memulai perjalanan kripto Anda.Langkah 1: Buat Akun HTX AndaGunakan alamat email atau nomor ponsel Anda untuk mendaftar akun gratis di HTX. Rasakan perjalanan pendaftaran yang mudah dan buka semua fitur.Dapatkan Akun SayaLangkah 2: Buka Beli Kripto, lalu Pilih Metode Pembayaran AndaKartu Kredit/Debit: Gunakan Visa atau Mastercard Anda untuk membeli Arweave (AR) secara instan.Saldo: Gunakan dana dari saldo akun HTX Anda untuk melakukan trading dengan lancar.Pihak Ketiga: Kami telah menambahkan metode pembayaran populer seperti Google Pay dan Apple Pay untuk meningkatkan kenyamanan.P2P: Lakukan trading langsung dengan pengguna lain di HTX.Over-the-Counter (OTC): Kami menawarkan layanan yang dibuat khusus dan kurs yang kompetitif bagi para trader.Langkah 3: Simpan Arweave (AR) AndaSetelah melakukan pembelian, simpan Arweave (AR) di akun HTX Anda. Selain itu, Anda dapat mengirimkannya ke tempat lain melalui transfer blockchain atau menggunakannya untuk memperdagangkan mata uang kripto lainnya.Langkah 4: Lakukan trading Arweave (AR)Lakukan trading Arweave (AR) dengan mudah di pasar spot HTX. Cukup akses akun Anda, pilih pasangan perdagangan, jalankan trading, lalu pantau secara real-time. Kami menawarkan pengalaman yang ramah pengguna baik untuk pemula maupun trader berpengalaman.

836 Total TayanganDipublikasikan pada 2024.12.11Diperbarui pada 2026.06.02

Cara Membeli AR

Diskusi

Selamat datang di Komunitas HTX. Di sini, Anda bisa terus mendapatkan informasi terbaru tentang perkembangan platform terkini dan mendapatkan akses ke wawasan pasar profesional. Pendapat pengguna mengenai harga AR (AR) disajikan di bawah ini.

活动图片