Tsinghua University's Special Award Winner, Gu Yuxian, Joins DeepSeek

marsbitPubblicato 2026-07-06Pubblicato ultima volta 2026-07-06

Introduzione

Tsinghua University's prestigious Graduate Special Scholarship recipient and 2021 Ph.D. candidate, Yuxian Gu, has officially joined DeepSeek. This news coincides with DeepSeek's major recruitment drive and the imminent launch of DeepSeek V4, on whose research paper Gu is listed as an author. A doctoral student in the Conversational AI group under Professor Minlie Huang at Tsinghua, Gu's research focuses on enhancing efficiency throughout the entire lifecycle of large language models. His key contributions span three areas: innovative methods for pre-training data selection (e.g., PDS), advanced knowledge distillation techniques for model compression (notably MiniLLM), and the development of efficient model architectures like Jet-Nemotron. His work has gained significant recognition, with nearly 5,000 citations on Google Scholar. Key publications include the highly cited surveys and papers on pre-trained models and the MiniLLM distillation method. As first author, he has presented at top-tier AI conferences including NeurIPS, ICLR, and ACL. One of his notable achievements is the Jet-Nemotron architecture, which combines Post-Neural Architecture Search (PostNAS) and a novel linear attention module called JetBlock. This model series demonstrates state-of-the-art performance rivaling larger models while achieving substantial efficiency gains in inference. Gu's expertise in creating powerful yet efficient AI systems aligns with industry needs, as evidenced by the adoption of h...

Recently, DeepSeek has embarked on a hiring spree, with openings across multiple departments including algorithm, R&D, product, operations, data engineering, and administrative functions.

At the same time, the official version of DeepSeek V4 is scheduled for release in mid-month. In the author list of the earlier DeepSeek V4 paper, we discovered the name of Gu Yuxian (Yuxian Gu), a Tsinghua University class of 2021 Ph.D. student and recipient of the 2025 Special Scholarship for Graduate Students.

To our knowledge, Gu Yuxian has formally joined DeepSeek.

Gu Yuxian has also received the 2025 Apple Ph.D. Fellowship and the Ant In-Tech Scholarship.

"Algorithmic innovation becomes the key to breaking computational bottlenecks when hardware resources are constrained," said Tsinghua alumnus Gu Yuxian. He is a graduating Ph.D. student in the Department of Computer Science at Tsinghua University, having also completed his undergraduate studies there.

His personal homepage shows that Gu Yuxian studied in the Conversational AI (CoAI) group at Tsinghua University under the supervision of Professor Huang Minlie.

Personal homepage address: https://t1101675.github.io/

His research primarily focuses on improving efficiency throughout the entire lifecycle of large language models, covering key stages such as pre-training, downstream adaptation, and inference. His recent work has mainly progressed in three directions:

Pre-training Data Selection: Dedicated to constructing theory and algorithms to optimize the data selection process for training large language models, thereby training more powerful and efficient models. Representative work includes PDS, Instruction Pre-training, and Learning Law.

Knowledge Distillation in Model Compression: Designing new methods to effectively transfer knowledge from large models to smaller, more deployable models. Representative achievements in this direction include MiniLLM and MiniPLM.

Efficient Model Architecture: Exploring and designing new model architectures that improve performance while reducing computational costs. Related work includes Jet-Nemotron.

On his Google Scholar homepage, Gu Yuxian's paper citation count has reached nearly 5000, with two papers exceeding 1000 citations: "Pre-trained models: Past, present and future" and "MiniLLM: Knowledge distillation of large language models".

As the first author, Gu Yuxian has published papers multiple times at top international AI academic conferences such as NeurIPS, ICLR, and ACL.

Last year, Jiqizhixin reported on "Jet-Nemotron", a new series of hybrid-architecture language models that achieved SOTA full-attention model accuracy while demonstrating outstanding efficiency.

The core innovations of Jet-Nemotron are primarily reflected in the following two points:

Post Neural Architecture Search (PostNAS): An efficient pipeline for post-training architecture exploration and adaptation, applicable to any pre-trained Transformer model.

JetBlock: A novel linear attention module whose performance significantly outperforms previous designs like Mamba2.

Paper address: https://arxiv.org/pdf/2508.15884

At that time, the 2B version of Jet-Nemotron's performance could rival the latest SOTA open-source full-attention language models like Qwen3, Qwen2.5, Gemma3, and Llama3.2, while achieving significant efficiency gains. On H100 GPUs, its generation throughput achieved a speedup of up to 53.6x (with a context length of 256K and maximum batch size).

On the MMLU and MMLU-Pro benchmarks, Jet-Nemotron's accuracy also surpassed some MoE full-attention models, such as DeepSeek-V3-Small and Moonlight, despite those models having larger parameter scales.

Earlier in 2024, Gu Yuxian and his collaborators proposed a knowledge distillation method for distilling large language models into smaller language models. They first utilized reverse Kullback-Leibler divergence (KLD) to replace the forward KLD objective in standard knowledge distillation methods, then derived an effective optimization method to learn this objective.

They named the resulting student model "MiniLLM". Extensive experiments in instruction-following scenarios showed that compared to baseline methods, MiniLLM could generate more precise answers with higher overall quality, while also having lower exposure bias, better calibration capability, and stronger long-text generation performance.

Leading open-source communities and industrial platforms like Google, Alibaba, and NVIDIA have adopted this method.

Paper address: https://arxiv.org/pdf/2306.08543

We also look forward to Gu Yuxian bringing more new achievements in the next chapter of his career at "DeepSeek".

This article is from the WeChat public account "Jiqizhixin" (ID: almosthuman2014), author: Jiqizhixin focusing on AI talent.

Domande pertinenti

QWho is Yuxian Gu and what recent career move did he make?

AYuxian Gu is a 2021-level doctoral student from Tsinghua University and a recipient of the 2025 Graduate Special Scholarship. He has recently formally joined the AI company DeepSeek.

QWhat are the main research areas of Yuxian Gu, as mentioned in the article?

AHis research focuses on improving efficiency throughout the lifecycle of large language models. His recent work covers three main areas: pre-training data selection, knowledge distillation for model compression, and efficient model architecture design.

QAccording to the article, what is a key innovation of the Jet-Nemotron model architecture?

AA key innovation of Jet-Nemotron is the JetBlock, a novel linear attention module, which significantly outperforms previous designs like Mamba2.

QWhat significant achievement is mentioned regarding Yuxian Gu's publication metrics?

AAccording to his Google Scholar profile, his publications have received nearly 5,000 citations. Two of his papers each have over 1,000 citations.

QWhat model did Yuxian Gu and collaborators introduce in 2024 for knowledge distillation, and how is it described?

AIn 2024, they introduced 'MiniLLM,' a method for distilling large language models into smaller ones. It is described as generating more accurate responses with higher overall quality, lower exposure bias, better calibration, and stronger long-text generation performance compared to baseline methods.

Letture associate

Treasury Secretary's Move to Suppress Treasury Yields Ignites 'Currency Debasement Trade'! Gold Hits Three-Month High, Bitcoin Surges Over 25% in a Single Week

US Treasury Secretary Besant's efforts to lower long-term Treasury yields by announcing expanded buybacks had only a brief market impact. However, this move fueled a "currency devaluation trade," weakening the US dollar while boosting both gold (to a three-month high) and Bitcoin (up over 25% for the week). Analysts attribute this reaction to deepening market concerns over the massive US fiscal deficit and structural pressures keeping long-term rates elevated, including fierce competition for capital from global government borrowing and massive AI sector financing. Despite the Treasury's actions, fundamental forces like growth, inflation, and capital demand are seen as limiting its ability to sustainably suppress yields. Bitcoin's strong positive correlation with gold has reinforced its narrative as a hedge against devaluation. While equity markets have shown resilience, some strategists warn that Treasury yields nearing 5% increase pressure on the dollar and high-leverage assets. Figures like Ray Dalio have advised reducing bond exposure in favor of gold and some Bitcoin, citing US debt risks. Market opinions are divided on the sustainability of the devaluation trade, with some noting the lack of a near-term catalyst for its next leg higher. The underlying tension between the Treasury's desire for lower borrowing costs and the Federal Reserve's focus on inflation and reducing market intervention remains a key theme. Upcoming events like Nvidia's earnings and the Jackson Hole symposium will test whether AI profits can continue supporting stocks and if the Fed aligns more with Washington's preference for easier financial conditions.

华尔街日报2 h fa

Treasury Secretary's Move to Suppress Treasury Yields Ignites 'Currency Debasement Trade'! Gold Hits Three-Month High, Bitcoin Surges Over 25% in a Single Week

华尔街日报2 h fa

Alexander Shokhin: Business Needs an Interest Rate Below 10% and the Dollar at 90-95 Rubles

Alexander Shokhin, head of the Russian Union of Industrialists and Entrepreneurs (RSPP), has advocated for potentially using "non-market" tools to keep the ruble within a target exchange rate corridor. This, he argues on August 21, would help avoid excessive volatility, though he called the topic a separate discussion. Shokhin had previously raised the idea of a currency corridor in late May, noting the ruble's current exchange rate is not fully market-driven due to a limited currency segment and reduced foreign currency demand. He stated that many business community colleagues propose fixing a corridor, even through non-market methods, to ensure predictability. The business community's key targets, as outlined by Shokhin in late December 2025, are a Central Bank key rate of 12%, inflation of 4–5%, and a US dollar exchange rate of 90–95 rubles by the end of 2026. A turning point for investment, he said, would be lowering the rate to 12% with 6% inflation, though truly comfortable business conditions would require a rate below 10%. He stressed the critical importance of currency predictability for corporate investment decisions. From a data analysis perspective, the idea of a ruble corridor is not new. A similar mechanism was used in Russia from 1995 to 1998, where the central bank held the dollar within fixed boundaries through regular interventions. This regime lasted three years before ending abruptly during the 1998 default, illustrating the fragility of rigid targets under external shocks. The macro-economic link is clear: stricter corridors require more reserves to defend against currency pressure. The key unresolved technical aspect is the specific sources and volume of such interventions given the current market's limited liquidity. Whether this discussion remains theoretical or leads to concrete corridor parameters will be seen in the coming months.

cryptonews.ru3 h fa

Alexander Shokhin: Business Needs an Interest Rate Below 10% and the Dollar at 90-95 Rubles

cryptonews.ru3 h fa

Trading

Spot
活动图片