OpenAI Loses 'The God of CUDA Kernels'

marsbit2026-08-17 tarihinde yayınlandı2026-08-17 tarihinde güncellendi

Özet

OpenAI has lost Scott Gray, a foundational engineer renowned as the "CUDA Kernel God" and one of the world's top GPU programmers. His departure, indicated by a subtle update to his social media bio, marks the exit of another key figure from the company's early days. Gray joined OpenAI as a full-time member in August 2016 and spent a decade there, contributing critically to performance optimization. His methodology was defined by bypassing software abstractions to push hardware to its absolute limits, exemplified by his early work on the maxas assembler and block-sparse GPU kernels. At OpenAI, his optimizations were integral to major projects including sparse transformers, GPT-3, DALL·E, and the core attention kernels running on vast GPU clusters. Gray's last original post in November 2023 stated, "OpenAI is nothing without its people," during the internal crisis following Sam Altman's brief ouster. His new direction, as noted in his bio, is to independently explore "neuroscience-inspired AI methods," a return to a long-standing personal interest mentioned in his original 2016 OpenAI introduction. His exit is part of a broader trend in 2026, which has seen at least 12 senior leaders depart OpenAI across operations, commercial, product, research, safety, and hardware divisions. While OpenAI's engineering systems will continue, losing an engineer of Gray's caliber—who embodied the deep technical prowess that shaped the company's infrastructure—signals a shift. As OpenAI prepar...

Today, a tweet from two days ago started trending widely. It contained only five words: Scott Gray has left OpenAI.

This renowned engineer, known as the "God of CUDA Kernels" and the "world's strongest GPU programmer," has become another heavyweight talent lost by OpenAI this year. He has not confirmed the news on X himself, and his LinkedIn page has not been updated yet. However, his personal bio on both X and Bluesky has already been changed from "GPU Geek at @OpenAI" to "Former GPU geek at @OpenAI." Moreover, he added a line in front: "Currently independent, exploring some neuroscience-inspired AI approaches."

An engineer who spent ten years at OpenAI announced his departure with a single line in his bio.

Quite a coincidence. Exactly ten years ago, on August 16, OpenAI published a blog post titled "Team update," introducing five full-time members who joined that month. The list included Dario Amodei, Filip Wolski, Jack Clark, Scott Gray, and Zain Shah.

https://openai.com/index/team-update-august/

Ten years later today, Dario is the CEO of Anthropic, Jack Clark handles policy at Anthropic, and Scott Gray has marked OpenAI as "former."

As of now, OpenAI has not released any statement regarding Scott Gray's departure, and Gray himself has not posted any detailed explanation. The common source for all current reports is the change in his social media bios, followed by a viral爆料post on X. Specific details such as the exact departure date, next destination, or whether he has already started a new company are not publicly available.

Another notable detail on his account: Gray's last original X post was on November 20, 2023. The content was a single sentence: "OpenAI is nothing without its people." That was three days after Sam Altman was fired by the board. Over seven hundred employees signed a petition demanding the board's resignation, and this sentence was repeatedly copied and pasted within OpenAI at the time, considered a collective show of loyalty. Since then, he has replied to some tweets (the last replies stopped after February 2025) but has not posted any new original content.

Introduction to Scott Gray

We previously had an article dedicated to introducing Scott Gray. Refer to "God of CUDA Kernels, World's Strongest GPU Programmer? Who is this Behind-the-Scenes Master at OpenAI?" Here is a brief overview.

His fame began during his time at Nervana Systems. At that time, the vast majority of developers relied on NVIDIA's CUDA C/C++ and official libraries like cuBLAS and cuDNN. Multiple layers of software abstraction shielded hardware details but also created performance ceilings. Gray's judgment was that it was necessary to bypass these abstraction layers. Thus, he wrote maxas, an assembler for the Maxwell architecture, allowing developers to write SASS machine code directly, manually allocate registers, manage memory latency, and control instruction pipelines.

To prove this path was viable, he hand-wrote an SGEMM kernel using maxas. On the GM204 GPU, this kernel achieved 98% of the hardware's theoretical peak performance and was 4.8% faster than NVIDIA's official closed-source, expert-hand-written cuBLAS. The subsequent maxDNN extended the same methodology to convolution: achieving stable computational efficiency between 93% and 95% on all convolutional layers of AlexNet, while cuDNN's efficiency at the time fluctuated wildly between 32% and 57%.

This became the foundational methodology for all his subsequent work: not accepting the performance ceilings imposed by abstraction layers.

In that August 2016 OpenAI team update, the official introduction of him was only two technical sentences, essentially saying he previously focused on optimizing the performance of deep networks on GPUs at Nervana, and his assembly-level optimizations for dense linear algebra and convolution "remain the fastest to this day." The same paragraph also mentioned that when not writing code, he usually reads the latest research in neuroscience and related fields.

It seems that ten years later, his new direction of "neuroscience-inspired AI approaches" is actually a return to that enduring interest.

After joining OpenAI, his role shifted from "optimizer" to "enabler." In 2017, together with Alec Radford and Durk Kingma, he released block-sparse GPU kernels. Unlike unstructured sparsity that removes individual weights, block sparsity cuts the weight matrix into fixed-size blocks and zeroes out entire blocks. Dedicated kernels completely skip zero-value blocks during computation, achieving speeds several orders of magnitude faster than cuBLAS handling dense matrices or cuSPARSE handling general sparse matrices. This set of kernels was open-sourced at the time, directly catalyzing a series of subsequent sparse attention works.

OpenAI Block Sparse Kernel Diagram, https://cdn.openai.com/blocksparse/blocksparsepaper.pdf

Following this thread, the 2019 paper "Generating Long Sequences with Sparse Transformers" was authored by Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever—reducing the time and memory overhead of attention from O(n²) to O(n√n). The paper explicitly listed "fast attention training kernels" as one of three key contributions.

Moving forward, his name appears in the author list for GPT-3, "Scaling Laws for Neural Language Models," DALL·E, and the technical report for OpenAI Five.

During this time, he also earned the nickname "God of CUDA Kernels." That was in September 2025. Former OpenAI employee Rohan Pandey posted on X, saying that approximately only one person at the company was responsible for the inference-side CUDA kernels. Colleagues called the attention kernel he wrote "the Bob kernel," which executed trillions of times daily on hundreds of thousands of GPUs. After the post went viral, comments widely speculated that "Bob" was Scott Gray.

Related posts also mentioned: there are probably fewer than a hundred people worldwide who can write high-performance CUDA kernels for the training process (especially backpropagation). This skill requires simultaneous expertise in parallel computing theory, GPU hardware architecture, and deep learning algorithms, while most practitioners remain at the application or inference optimization layer.

And now, Scott Gray has left.

2026: OpenAI Has Lost Quite a Few People

OpenAI has lost a significant number of talents this year. According to a compilation by X blogger Chubby, OpenAI has seen 12 senior leaders depart this year, covering operations, business, product, research, safety, ethics, and hardware.

A few days ago, Brad Lightcap, who worked at OpenAI for eight years and served as CFO and COO, announced his departure on X, stating he wanted to "start something new." Two days later, Chief Revenue Officer Denise Dresser, who just took the position in December last year, announced her departure, to be succeeded by former Wiz President & COO Dali Rajic.

On the business and product front, CEO of Applications Fidji Simo took a leave of absence in April for health reasons and stepped down from full-time to become a part-time advisor in July; former CMO Kate Rouch stepped down in April due to treatment needs; CTO of Enterprise Business Srinivas Narayanan also announced his departure in April. On the research and product side, Kevin Weil, who led OpenAI for Science, and Sora lead Bill Peebles both announced their departures on the same day in April; Weil is currently working on an AI science startup. Barret Zoph returned to OpenAI in January from Thinking Machines Lab to lead enterprise AI sales but left again in June; the company did not specify a reason.

Three people from the safety and ethics lines departed concentrated in July: Safety Systems Lead Johannes Heidecke left during a safety team reorganization, with his responsibilities merged into the research department led by Mia Glaese; Chief Futurist Joshua Achiam left after nearly nine years—the mission alignment team he previously led was disbanded in February this year; Ethics Lead Chloé Bakalar left less than a year after joining. Earlier, in March, Robotics and Consumer Hardware Lead Caitlin Kalinowski resigned due to disagreement with the company's cooperation with the Pentagon. She stated on LinkedIn that such issues should undergo more thorough review and later added that her objection was primarily about the governance process.

This list also missed one: in January this year, Research VP Jerry Tworek, responsible for reasoning models, left after nearly seven years to start his own venture, citing a desire to pursue research directions "difficult to do within OpenAI."

Conclusion

Scott Gray's departure may not immediately alter OpenAI's model release cadence. A decade of accumulated engineering systems will not come to a halt because one person leaves. However, for a company increasingly reliant on scale, efficiency, and cost control, losing an engineer who could push hardware performance to its limits and was deeply involved in multiple generations of core models is still not a change that can be taken lightly.

The timing of this departure is also noteworthy: OpenAI is preparing to go public, transforming into a giant corporation pursuing revenue and participating in a global infrastructure race. Now, some of the people who originally shaped it are also re-evaluating what they want to do next.

Scott Gray's choice wasn't to join another frontier lab or announce a startup. He simply stated in his personal bio that he is independently exploring "neuroscience-inspired AI approaches." From the engineer who read neuroscience papers in his spare time in the 2016 official introduction, to writing neuroscience back into his personal bio ten years later, this seems more like a return after a long detour: after squeezing the existing computing paradigm to its absolute limit, he began searching for another possibility.

In 2023, Gray wrote, "OpenAI is nothing without its people." Looking back today, this statement was both a stance during that company crisis and seems like a belated footnote. OpenAI will not lose everything because of any one person's departure. But ultimately, a company's direction is shaped by those who stay, those who leave, and the respective problems they choose to explore.

This article is from the WeChat public account "Almost Human" (ID: almosthuman2014), author: Focus on AI

İlgili Sorular

QWho is Scott Gray and why is his departure from OpenAI significant?

AScott Gray was a renowned engineer at OpenAI, widely regarded as a 'CUDA kernel god' and one of the world's best GPU programmers. His departure is significant because he was a founding-era employee who made major contributions to core models like GPT-3 and DALL·E, and his expertise in writing high-performance CUDA kernels was critical for pushing hardware efficiency to its limits.

QWhat was the main methodology behind Scott Gray's technical work?

AThe core methodology behind Scott Gray's work was rejecting the performance ceilings imposed by software abstraction layers. He famously bypassed standard libraries like cuBLAS and cuDNN by writing low-level assembly code (SASS machine code) to manually control GPU registers, memory latency, and instruction pipelines, achieving near-peak hardware performance.

QAccording to the article, what new direction is Scott Gray pursuing after leaving OpenAI?

AAfter leaving OpenAI, Scott Gray's social media profiles indicate he is currently independent and exploring 'neuroscience-inspired AI methods.' This aligns with a long-standing personal interest noted in his 2016 OpenAI introduction, which mentioned he spent his non-coding time reading the latest neuroscience research.

QBesides Scott Gray, which other key OpenAI leaders have left the company in 2026, as mentioned in the article?

AIn 2026, OpenAI has seen several key leaders depart, including former CFO/COO Brad Lightcap, Chief Revenue Officer Denise Dresser, Applications CEO Fidji Simo, former CMO Kate Rouch, Enterprise CTO Srinivas Narayanan, Sora lead Bill Peebles, safety leader Johannes Heidecke, chief futurist Joshua Achiam, ethics lead Chloé Bakalar, hardware head Caitlin Kalinowski, and research VP Jerry Tworek.

QWhat does the article suggest about the broader context and potential impact of Scott Gray's departure?

AThe article suggests Gray's departure is part of a broader trend of early, influential employees leaving as OpenAI transitions into a large, revenue-driven public company. While its engineering systems won't collapse overnight, losing such a unique talent who maximized hardware efficiency represents a meaningful shift. His choice to explore neuroscience-inspired AI also symbolizes a search for new paradigms after pushing the current computational model to its limits.

İlgili Okumalar

Glassnode: Consumer Confidence Falls as AI-Related Stocks Rise, Bitcoin Lags Behind

According to Glassnode, consumer confidence remains at one of its lowest levels in a decade, despite two consecutive months of improvement. This has not stopped households from moving money out of cash, as they expect further cost-of-living increases and a broader economic slowdown. The key question is where this capital is flowing. US stocks hit a new all-time high in early August, primarily driven by trading in AI-related stocks rather than a broad market rally. Bitcoin, historically seen as a hedge against declining trust in traditional finance, has not participated in this movement. Spot Bitcoin ETFs saw outflows of $389.7 million in one week, coinciding with rising equity markets—a divergence that aligns with Glassnode's data on capital flows. Bitcoin is currently trading at roughly half its October 2025 peak, stuck in a narrow range. Meanwhile, AI-related trading continues to attract fresh capital from retail traders, hedge funds, and even crypto-native institutional investors, who are redirecting funds into AI stocks and tokens. The macroeconomic backdrop has not been hostile to Bitcoin, with core inflation at a moderate 2.5% in July. However, Bitcoin's muted response to favorable inflation data is seen as a concerning signal, given its supposed role as a hedge against currency debasement. Spot Bitcoin exchange trading volume has fallen to its lowest since 2019, and recent ETF inflows are only a "fraction of any prior accumulation wave," suggesting institutional buying may have paused. This trend extends beyond trading: some Bitcoin miners are repurposing their power contracts and data center capacity for AI workloads. This appears to be a structural shift that could pressure Bitcoin's status as the default destination for capital leaving cash. The fundamental arguments for Bitcoin as a hedge against inflation or scarcity are not invalidated, but their expected impact has not materialized within the timeline anticipated by crypto optimists this summer.

cryptonews.ru3 dk önce

Glassnode: Consumer Confidence Falls as AI-Related Stocks Rise, Bitcoin Lags Behind

cryptonews.ru3 dk önce

The Optimal 'AI Bubble Trade': Simultaneously Going Long on 'Arrogance' and 'Bias'

The optimal investment strategy in the current AI bubble environment is a dual "leg" approach: going long on both "hubris" (AI tech leaders) and "humiliation" (neglected, underperforming cyclical assets). This aims to capture gains from both sides during the final surge of a nominal GDP-driven bubble, according to a Bank of America report by strategist Michael Hartnett. The bank's Bull & Bear Indicator remains in extreme bullish territory, signaling "sell", yet history shows such signals have limited immediate impact. Current fund flows show structural shifts: gold saw its largest weekly inflow since January, commodities are up 58.9% YTD, while tech stocks experienced their largest weekly outflow in seven weeks. The core thesis is that the final stage of a bubble benefits both the leading theme ("hubris" - AI) and oversold sectors ("humiliation" - like consumer stocks), similar to patterns seen in the 1999 tech bubble and 2007-2008 credit crisis. The report advises shorting "AI bonds," anticipating pressure from massive capital expenditures. Key risks include high concentration, surging bond yields, and cautious voter sentiment. The US debt burden is highlighted, with servicing costs reaching $1.4 trillion. The 10-year Treasury yield breaching 5% is seen as a red line for policymakers. For the "avoid the dollar" theme, BofA recommends gold and Hong Kong property stocks, the latter seen as deeply undervalued. The November US midterm elections, particularly the Texas governor race concerning AI data center expansion, are flagged as a critical political variable that could determine the AI bull market's trajectory. Private client data shows record-high equity allocations (66.4%) and record-low cash levels (9.4%), indicating bullish positioning. The report concludes that while overbought conditions can pause the bull market, ending it requires a combination of excessive positioning, overly optimistic earnings, and policy tightening—a scenario not yet in place.

marsbit43 dk önce

The Optimal 'AI Bubble Trade': Simultaneously Going Long on 'Arrogance' and 'Bias'

marsbit43 dk önce

İşlemler

Spot
活动图片