OpenAI Loses 'The God of CUDA Kernels'

marsbitPublished on 2026-08-17Last updated on 2026-08-17

Abstract

OpenAI has lost Scott Gray, a foundational engineer renowned as the "CUDA Kernel God" and one of the world's top GPU programmers. His departure, indicated by a subtle update to his social media bio, marks the exit of another key figure from the company's early days. Gray joined OpenAI as a full-time member in August 2016 and spent a decade there, contributing critically to performance optimization. His methodology was defined by bypassing software abstractions to push hardware to its absolute limits, exemplified by his early work on the maxas assembler and block-sparse GPU kernels. At OpenAI, his optimizations were integral to major projects including sparse transformers, GPT-3, DALL·E, and the core attention kernels running on vast GPU clusters. Gray's last original post in November 2023 stated, "OpenAI is nothing without its people," during the internal crisis following Sam Altman's brief ouster. His new direction, as noted in his bio, is to independently explore "neuroscience-inspired AI methods," a return to a long-standing personal interest mentioned in his original 2016 OpenAI introduction. His exit is part of a broader trend in 2026, which has seen at least 12 senior leaders depart OpenAI across operations, commercial, product, research, safety, and hardware divisions. While OpenAI's engineering systems will continue, losing an engineer of Gray's caliber—who embodied the deep technical prowess that shaped the company's infrastructure—signals a shift. As OpenAI prepar...

Today, a tweet from two days ago started trending widely. It contained only five words: Scott Gray has left OpenAI.

This renowned engineer, known as the "God of CUDA Kernels" and the "world's strongest GPU programmer," has become another heavyweight talent lost by OpenAI this year. He has not confirmed the news on X himself, and his LinkedIn page has not been updated yet. However, his personal bio on both X and Bluesky has already been changed from "GPU Geek at @OpenAI" to "Former GPU geek at @OpenAI." Moreover, he added a line in front: "Currently independent, exploring some neuroscience-inspired AI approaches."

An engineer who spent ten years at OpenAI announced his departure with a single line in his bio.

Quite a coincidence. Exactly ten years ago, on August 16, OpenAI published a blog post titled "Team update," introducing five full-time members who joined that month. The list included Dario Amodei, Filip Wolski, Jack Clark, Scott Gray, and Zain Shah.

https://openai.com/index/team-update-august/

Ten years later today, Dario is the CEO of Anthropic, Jack Clark handles policy at Anthropic, and Scott Gray has marked OpenAI as "former."

As of now, OpenAI has not released any statement regarding Scott Gray's departure, and Gray himself has not posted any detailed explanation. The common source for all current reports is the change in his social media bios, followed by a viral爆料post on X. Specific details such as the exact departure date, next destination, or whether he has already started a new company are not publicly available.

Another notable detail on his account: Gray's last original X post was on November 20, 2023. The content was a single sentence: "OpenAI is nothing without its people." That was three days after Sam Altman was fired by the board. Over seven hundred employees signed a petition demanding the board's resignation, and this sentence was repeatedly copied and pasted within OpenAI at the time, considered a collective show of loyalty. Since then, he has replied to some tweets (the last replies stopped after February 2025) but has not posted any new original content.

Introduction to Scott Gray

We previously had an article dedicated to introducing Scott Gray. Refer to "God of CUDA Kernels, World's Strongest GPU Programmer? Who is this Behind-the-Scenes Master at OpenAI?" Here is a brief overview.

His fame began during his time at Nervana Systems. At that time, the vast majority of developers relied on NVIDIA's CUDA C/C++ and official libraries like cuBLAS and cuDNN. Multiple layers of software abstraction shielded hardware details but also created performance ceilings. Gray's judgment was that it was necessary to bypass these abstraction layers. Thus, he wrote maxas, an assembler for the Maxwell architecture, allowing developers to write SASS machine code directly, manually allocate registers, manage memory latency, and control instruction pipelines.

To prove this path was viable, he hand-wrote an SGEMM kernel using maxas. On the GM204 GPU, this kernel achieved 98% of the hardware's theoretical peak performance and was 4.8% faster than NVIDIA's official closed-source, expert-hand-written cuBLAS. The subsequent maxDNN extended the same methodology to convolution: achieving stable computational efficiency between 93% and 95% on all convolutional layers of AlexNet, while cuDNN's efficiency at the time fluctuated wildly between 32% and 57%.

This became the foundational methodology for all his subsequent work: not accepting the performance ceilings imposed by abstraction layers.

In that August 2016 OpenAI team update, the official introduction of him was only two technical sentences, essentially saying he previously focused on optimizing the performance of deep networks on GPUs at Nervana, and his assembly-level optimizations for dense linear algebra and convolution "remain the fastest to this day." The same paragraph also mentioned that when not writing code, he usually reads the latest research in neuroscience and related fields.

It seems that ten years later, his new direction of "neuroscience-inspired AI approaches" is actually a return to that enduring interest.

After joining OpenAI, his role shifted from "optimizer" to "enabler." In 2017, together with Alec Radford and Durk Kingma, he released block-sparse GPU kernels. Unlike unstructured sparsity that removes individual weights, block sparsity cuts the weight matrix into fixed-size blocks and zeroes out entire blocks. Dedicated kernels completely skip zero-value blocks during computation, achieving speeds several orders of magnitude faster than cuBLAS handling dense matrices or cuSPARSE handling general sparse matrices. This set of kernels was open-sourced at the time, directly catalyzing a series of subsequent sparse attention works.

OpenAI Block Sparse Kernel Diagram, https://cdn.openai.com/blocksparse/blocksparsepaper.pdf

Following this thread, the 2019 paper "Generating Long Sequences with Sparse Transformers" was authored by Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever—reducing the time and memory overhead of attention from O(n²) to O(n√n). The paper explicitly listed "fast attention training kernels" as one of three key contributions.

Moving forward, his name appears in the author list for GPT-3, "Scaling Laws for Neural Language Models," DALL·E, and the technical report for OpenAI Five.

During this time, he also earned the nickname "God of CUDA Kernels." That was in September 2025. Former OpenAI employee Rohan Pandey posted on X, saying that approximately only one person at the company was responsible for the inference-side CUDA kernels. Colleagues called the attention kernel he wrote "the Bob kernel," which executed trillions of times daily on hundreds of thousands of GPUs. After the post went viral, comments widely speculated that "Bob" was Scott Gray.

Related posts also mentioned: there are probably fewer than a hundred people worldwide who can write high-performance CUDA kernels for the training process (especially backpropagation). This skill requires simultaneous expertise in parallel computing theory, GPU hardware architecture, and deep learning algorithms, while most practitioners remain at the application or inference optimization layer.

And now, Scott Gray has left.

2026: OpenAI Has Lost Quite a Few People

OpenAI has lost a significant number of talents this year. According to a compilation by X blogger Chubby, OpenAI has seen 12 senior leaders depart this year, covering operations, business, product, research, safety, ethics, and hardware.

A few days ago, Brad Lightcap, who worked at OpenAI for eight years and served as CFO and COO, announced his departure on X, stating he wanted to "start something new." Two days later, Chief Revenue Officer Denise Dresser, who just took the position in December last year, announced her departure, to be succeeded by former Wiz President & COO Dali Rajic.

On the business and product front, CEO of Applications Fidji Simo took a leave of absence in April for health reasons and stepped down from full-time to become a part-time advisor in July; former CMO Kate Rouch stepped down in April due to treatment needs; CTO of Enterprise Business Srinivas Narayanan also announced his departure in April. On the research and product side, Kevin Weil, who led OpenAI for Science, and Sora lead Bill Peebles both announced their departures on the same day in April; Weil is currently working on an AI science startup. Barret Zoph returned to OpenAI in January from Thinking Machines Lab to lead enterprise AI sales but left again in June; the company did not specify a reason.

Three people from the safety and ethics lines departed concentrated in July: Safety Systems Lead Johannes Heidecke left during a safety team reorganization, with his responsibilities merged into the research department led by Mia Glaese; Chief Futurist Joshua Achiam left after nearly nine years—the mission alignment team he previously led was disbanded in February this year; Ethics Lead Chloé Bakalar left less than a year after joining. Earlier, in March, Robotics and Consumer Hardware Lead Caitlin Kalinowski resigned due to disagreement with the company's cooperation with the Pentagon. She stated on LinkedIn that such issues should undergo more thorough review and later added that her objection was primarily about the governance process.

This list also missed one: in January this year, Research VP Jerry Tworek, responsible for reasoning models, left after nearly seven years to start his own venture, citing a desire to pursue research directions "difficult to do within OpenAI."

Conclusion

Scott Gray's departure may not immediately alter OpenAI's model release cadence. A decade of accumulated engineering systems will not come to a halt because one person leaves. However, for a company increasingly reliant on scale, efficiency, and cost control, losing an engineer who could push hardware performance to its limits and was deeply involved in multiple generations of core models is still not a change that can be taken lightly.

The timing of this departure is also noteworthy: OpenAI is preparing to go public, transforming into a giant corporation pursuing revenue and participating in a global infrastructure race. Now, some of the people who originally shaped it are also re-evaluating what they want to do next.

Scott Gray's choice wasn't to join another frontier lab or announce a startup. He simply stated in his personal bio that he is independently exploring "neuroscience-inspired AI approaches." From the engineer who read neuroscience papers in his spare time in the 2016 official introduction, to writing neuroscience back into his personal bio ten years later, this seems more like a return after a long detour: after squeezing the existing computing paradigm to its absolute limit, he began searching for another possibility.

In 2023, Gray wrote, "OpenAI is nothing without its people." Looking back today, this statement was both a stance during that company crisis and seems like a belated footnote. OpenAI will not lose everything because of any one person's departure. But ultimately, a company's direction is shaped by those who stay, those who leave, and the respective problems they choose to explore.

This article is from the WeChat public account "Almost Human" (ID: almosthuman2014), author: Focus on AI

Related Questions

QWho is Scott Gray and why is his departure from OpenAI significant?

AScott Gray was a renowned engineer at OpenAI, widely regarded as a 'CUDA kernel god' and one of the world's best GPU programmers. His departure is significant because he was a founding-era employee who made major contributions to core models like GPT-3 and DALL·E, and his expertise in writing high-performance CUDA kernels was critical for pushing hardware efficiency to its limits.

QWhat was the main methodology behind Scott Gray's technical work?

AThe core methodology behind Scott Gray's work was rejecting the performance ceilings imposed by software abstraction layers. He famously bypassed standard libraries like cuBLAS and cuDNN by writing low-level assembly code (SASS machine code) to manually control GPU registers, memory latency, and instruction pipelines, achieving near-peak hardware performance.

QAccording to the article, what new direction is Scott Gray pursuing after leaving OpenAI?

AAfter leaving OpenAI, Scott Gray's social media profiles indicate he is currently independent and exploring 'neuroscience-inspired AI methods.' This aligns with a long-standing personal interest noted in his 2016 OpenAI introduction, which mentioned he spent his non-coding time reading the latest neuroscience research.

QBesides Scott Gray, which other key OpenAI leaders have left the company in 2026, as mentioned in the article?

AIn 2026, OpenAI has seen several key leaders depart, including former CFO/COO Brad Lightcap, Chief Revenue Officer Denise Dresser, Applications CEO Fidji Simo, former CMO Kate Rouch, Enterprise CTO Srinivas Narayanan, Sora lead Bill Peebles, safety leader Johannes Heidecke, chief futurist Joshua Achiam, ethics lead Chloé Bakalar, hardware head Caitlin Kalinowski, and research VP Jerry Tworek.

QWhat does the article suggest about the broader context and potential impact of Scott Gray's departure?

AThe article suggests Gray's departure is part of a broader trend of early, influential employees leaving as OpenAI transitions into a large, revenue-driven public company. While its engineering systems won't collapse overnight, losing such a unique talent who maximized hardware efficiency represents a meaningful shift. His choice to explore neuroscience-inspired AI also symbolizes a search for new paradigms after pushing the current computational model to its limits.

Related Reads

Report on the State of the Crypto Industry by 2026, Featuring Sumsub Vice President for North America, Danielle LaBarbera

The cryptocurrency industry is entering an era of regulated maturity, driven by frameworks like the CLARITY Act and the $GENIUS Act in the US. These provide clearer rules but also raise operational standards, requiring platforms to effectively demonstrate compliance. Simultaneously, fraud is evolving into sophisticated, AI-powered, lifecycle-based attacks, moving beyond simple onboarding scams. According to Daniel LaBarbera, VP of Sumsub for North America, companies must shift from one-time KYC checks to continuous, risk-based verification. This involves integrating identity, behavioral, device, and transaction data into a unified risk view. The 2026 State of Crypto Industry report highlights that while 55% of crypto firms faced fraud last year, effective strategies now combine AI-driven detection, continuous monitoring, and behavioral analytics. Key compliance challenges persist, particularly with the Travel Rule. Only 23% of companies are fully compliant, with high implementation costs, data security concerns, and regulatory fragmentation being major hurdles. Meanwhile, stablecoins are gaining traction, accounting for 36% of all crypto transactions in 2025, and evolving from trading tools into financial infrastructure for payments and settlements. The path forward lies in risk-based approaches that balance security, speed, and user experience. This includes adopting documentless verification, reusable KYC, and breaking down silos between compliance functions to create a holistic, lifecycle view of customer risk and trust.

cryptonews.ru11m ago

Report on the State of the Crypto Industry by 2026, Featuring Sumsub Vice President for North America, Danielle LaBarbera

cryptonews.ru11m ago

U.S. National Debt Approaches $40 Trillion in 5 Months, Bitcoin Debate Gains Momentum

The US national debt has surged to nearly $40 trillion in just five months, the fastest trillion-dollar increase on record. The debt's growth has accelerated dramatically over time, from taking 192 years to reach the first $1 trillion in 1981 to adding the latest $1 trillion in only five months. The federal budget deficit for the fiscal year is already over $1.8 trillion, exceeding last year's total, and net interest payments on the debt have surpassed $1 trillion, now exceeding defense or Medicare spending. This rapid debt accumulation is fueling arguments within the cryptocurrency industry that Bitcoin serves as a hedge against government-backed currencies and fiscal irresponsibility. Proponents, including some lawmakers and industry leaders, argue Bitcoin could act as a "hard currency" fiscal control mechanism. Legislation has even been proposed for the Treasury to acquire Bitcoin to help reduce the national debt. The International Monetary Fund has warned that global public debt could hit 100% of world GDP by 2029 if current trends continue, with the US and China as primary drivers. Some analyses suggest a sovereign debt crisis could drive capital into alternative assets like Bitcoin, similar to past regional banking crises. However, skeptics point out that both Bitcoin and gold have fallen in price at times during 2026 despite record debt, indicating the depreciation hedge theory may operate over a much longer timeframe than short-term price action.

cryptonews.ru21m ago

U.S. National Debt Approaches $40 Trillion in 5 Months, Bitcoin Debate Gains Momentum

cryptonews.ru21m ago

Technological Self-Reliance in China: A War from Lithography Machines to ABF Films

"China's Tech Independence: A Battle from Lithography Machines to ABF Film" In August 2026, Japan's Ajinomoto announced a 30% supply cut of ABF film to Chinese mainland clients, causing industry-wide shock and fears of price hikes and shortages in the semiconductor supply chain. That same year, Chinese company Lotus Holdings, known for its MSG business, acquired a small domestic ABF film startup for 103 million yuan, aiming to change this passive situation. ABF film is a core insulating material for advanced CPU, GPU, and AI chip packages. Ajinomoto, originally a food flavoring company, has monopolized over 95% of the global ABF film market for nearly 30 years, deriving its technology from byproducts of monosodium glutamate production. With AI chips consuming 5-10 times more ABF film than traditional chips, the supply-demand gap is widening. Ajinomoto's supply cut to China, where domestic ABF film production accounts for less than 5%, directly threatens the production of domestic high-end AI chips and their substrates. This incident highlights a crucial but often overlooked truth: the vulnerabilities in China's quest for technological self-reliance extend beyond headline areas like lithography machines to critical but seemingly minor components—insulating films, photoresists, electronic specialty gases, etc. The article frames China's tech independence as a multi-front war. While major breakthroughs have been achieved in chip design (e.g., Huawei's HiSilicon), foundry (e.g., SMIC), and memory chips (e.g., YMTC), countless smaller "Ajinomoto-style" chokepoints remain. Lotus Holdings' acquisition represents a significant shift: the battle is no longer fought only by tech giants but has become a collective, industry-wide effort involving companies from diverse backgrounds. Historically, external blockades have often spurred China's technological breakthroughs, as seen with Huawei's Kirin chips and YMTC's 3D NAND flash memory. Ajinomoto's supply cut, while a short-term challenge, may similarly catalyze domestic innovation in ABF film and other critical materials. The path to technological sovereignty is long and arduous, requiring sustained patience and investment to fill every gap in the complex supply chain. Lotus's move is not an immediate solution but a step towards that future, symbolizing a broader, relentless march toward independence.

marsbit40m ago

Technological Self-Reliance in China: A War from Lithography Machines to ABF Films

marsbit40m ago

Trading

Spot
活动图片