First Long-Horizon Doc2Repo Training Dataset: Code Agents Move Beyond Bug Fixing and Begin Creating Repositories

marsbitОпубликовано 2026-06-25Обновлено 2026-06-25

Введение

With the advancement of LLM Code Agents, the research focus is shifting towards long-horizon, real-world tasks, moving beyond simple bug fixes to full repository generation. To address this, researchers from Renmin University of China introduced the DeNovoSWE dataset. This dataset focuses on long-term software engineering tasks, specifically the "document-to-repository" challenge—generating an entire, executable code repository from a task description. The DeNovoSWE construction method employs a Divide & Conquer approach. It breaks down target repositories into core capabilities and uses a multi-agent Draft-Critic-Repair workflow to automatically generate high-quality, evaluation-aligned task documents. The dataset also implements difficulty-aware filtering to balance quality and diversity. The result is a high-quality, anti-leakage dataset of 4,818 instances. Experiments show that models trained on DeNovoSWE achieve significant improvements in long-horizon repository generation. For instance, Qwen3-30B-A3B-Instruct's performance on the BeyondSWE-Doc2Repo benchmark increased from 5.8% to 47.2%, and on NL2RepoBench from 4.3% to 23.0%. Similar gains were observed with stronger backbones, demonstrating that dedicated long-horizon training data is crucial for advancing Code Agents from maintainers to architects capable of planning and building complete software projects from scratch.

With the continuous improvement of LLM Code Agent capabilities, more and more researchers are realizing it's time to advance to the next stage of long-horizon tasks that are closer to real-world scenarios. Consequently, some benchmarks for evaluating long-horizon tasks have emerged, such as NL2RepoBench and BeyondSWE. The expected role of Code Agents is gradually shifting from repository maintainers to architects, capable of planning and completing long-horizon coding tasks for entire repositories.

Recently, the Gaoling School of Artificial Intelligence at Renmin University of China completed related research and officially released the DeNovoSWE dataset, focusing on long-horizon software engineering tasks, particularly repository-level code generation from scratch.

Paper link: https://arxiv.org/pdf/2606.10728

Repository link: https://github.com/AweAI-Team/DeNovoSWE

Data link: https://huggingface.co/collections/AweAI-Team/denovoswe

Through the mechanisms of Divide & Conquer and Critic & Repair, a high-quality dataset was constructed, successfully achieving scaling for long-horizon SWE tasks. This effort resulted in DeNovoSWE, an open-source, high-quality long-horizon SWE task dataset containing 4,818 real-world instances. This achievement provides large-scale data for training Code Agents' long-horizon capabilities, significantly enhancing their performance on such tasks.

The paper also proposes methods based on difficulty score filtering, effectively alleviating the trade-off between the proportion of difficult problems and trajectory quality.

Experiments show that the Qwen3-30B-A3B-Instruct model trained on DeNovoSWE improved from 5.8% to 47.2% on BeyondSWE-Doc2Repo and from 4.3% to 23.0% on NL2RepoBench, demonstrating the significant boost in repository-level code generation capabilities brought by long-horizon data.

Rebuilding an Entire Repository from a Single Document

Over the past year, with the scaling of large-scale SWE data in works like Scale-SWE, code agents have rapidly progressed on real software engineering tasks like SWE-bench. But as models become increasingly adept at "fixing an issue" or "changing a few lines of buggy code," a more critical question arises: Do agents truly possess long-horizon software engineering capabilities? Judging from the performance of frontier models on BeyondSWE-Doc2Repo and NL2RepoBench, the results are not ideal.

Real-world software development often isn't about modifying a single function or adding a conditional statement. It involves understanding requirements, planning architecture, creating files, designing APIs, handling dependencies, connecting modules, and ultimately making the entire repository run successfully in tests.

In other words, the real challenge lies in long-horizon repository-level generation: starting from a task document and generating a complete, executable, and verifiable software repository. This is precisely the problem DeNovoSWE aims to solve.

High-Quality "Generate from Scratch" Task Documents

In document-to-repository generation, the document is not just a README, nor a simple API list. It is essentially the sole task entry point for the agent to rebuild the entire repository.

A high-quality task document needs to meet at least two core standards.

First, it must be well-organized.

Repository-level tasks are inherently complex, involving multiple modules, interfaces, configurations, data structures, and interaction flows. If the document merely piles up function descriptions, the agent can easily get lost in fragmented information. Therefore, the document should first provide a clear overview of the repository, then divide into chapters based on capabilities or workflows, ensuring each part corresponds to a clear functional boundary.

Second, it must be written from the perspective of reliable evaluation.

The document cannot be too sparse, otherwise the task becomes an underdefined problem, potentially requiring the model to guess aimlessly to pass evaluation. Nor can it be too detailed, as that would directly leak implementation details, making the task unchallenging.

A truly high-quality document should describe the key behaviors on which evaluation depends: including import paths, public APIs, inputs and outputs, default parameters, exception behaviors, configuration items, pattern strings, return fields, etc., while also outlining the general functionalities to be implemented. In other words, the document should be sufficient for the agent to reproduce testable behaviors, but it should not become a copy of the implementation code.

This is also the core idea of DeNovoSWE: making documents readable, implementable, and verifiable.

The DeNovoSWE Method

DeNovoSWE frames "generating a complete repository from a document" as a large-scale, verifiable long-horizon software engineering task. It does not rely on manually written documents but automatically constructs high-quality instances through a sandboxed multi-agent workflow. The entire method can be summarized in two steps: Divide and Conquer.

In the Divide stage, the system first analyzes the target repository, decomposing it into multiple repository capabilities.

Each capability corresponds to a core function or workflow within the repository, such as authentication and connection, data reading/writing, batch processing, export flows, etc. This way, the originally massive repository generation problem is split into several structurally clear document chapters.

Simultaneously, DeNovoSWE runs the original unit tests and collects execution traces to identify which functions, classes, and interfaces actually impact evaluation. This further distinguishes between direct components, core indirect components, and non-core indirect components: interfaces directly called by tests must be documented in detail; core indirect components that affect observable behaviors also need coverage; while non-core internal implementations can be left for the agent to handle freely.

In the Conquer stage, DeNovoSWE uses a Draft-Critic-Repair mechanism to generate documents for each capability one by one. The Draft agent writes an initial draft; the Critic agent checks for omissions of key APIs, behavioral contracts, or structural information; the Repair agent then fixes the document based on the feedback. This cycle iterates until each capability chapter is clear, complete, and aligned with evaluation.

Finally, the documents for different capabilities are merged into a single, comprehensive task document, serving as the sole basis for the agent to generate the repository from scratch.

Difficulty: Why is This a Long-Horizon Task?

The difficulty of DeNovoSWE tasks stems from a fundamental change: it's no longer issue-level fixing, but whole-repository generation.

In traditional SWE tasks, agents typically face an existing repository, needing only to locate a bug, modify local code, and pass tests.

In DeNovoSWE, the agent faces a cleaned environment: the original source code and tests are removed, git history is reset, and potential leakage channels like caches, site-packages residues, pip wheels, temporary compilation artifacts, etc., are also cleared. This means the agent must truly rely on the document to complete the entire repository rebuild. It needs to plan the project structure, create module files, define public interfaces, implement cross-file interactions, handle dependencies and configurations, and continuously fix errors across multiple rounds of editing and test feedback.

Any deviation in an API signature, return field, exception type, or default behavior can cause test failures. Errors can also accumulate over the long horizon: an early poorly designed module can affect multiple subsequent files and call chains.

To further address the difficulty variance across different repositories, DeNovoSWE also proposes difficulty-aware trajectory filtering. In simple terms, easy tasks should require a higher pass rate, while difficult tasks should not be entirely discarded just for failing to achieve a perfect score. DeNovoSWE sets different filtering thresholds for different difficulty intervals based on structural complexity and LLM difficulty assessment, thereby balancing quality and diversity.

This is particularly important for long-horizon tasks: the more complex the repository, the harder it is to pass all tests in one go. Yet, the trajectories from challenging repositories, even with low scores or partial success, still contain valuable long-horizon planning and implementation capabilities.

Experimental Results

DeNovoSWE ultimately constructed 4,818 high-quality document-to-repository task instances. It is an executable, evaluable, and trainable long-horizon software engineering environment.

Experimental results show that DeNovoSWE brings significant improvements to models' long-horizon repository generation capabilities. For Qwen3-30B-A3B-Instruct, the original model scored only 5.8% on BeyondSWE-Doc2Repo and 4.3% on NL2RepoBench. Training with conventional issue-level SWE data like Scale-SWE-Agent improved these to 29.2% and 18.3%, indicating that general SWE data does have transfer effects. However, when the model was trained using DeNovoSWE, performance further increased to 47.2% and 23.0%.

This demonstrates that data oriented towards "fixing bugs" cannot fully replace data oriented towards "generating complete repositories" for long-horizon tasks. To truly teach agents repository-level engineering, specialized training environments built for long-horizon tasks are needed.

On the stronger Qwen3.5-35B-A3B backbone, DeNovoSWE similarly brought stable gains: BeyondSWE-Doc2Repo improved from 43.8% to 50.0%, and NL2RepoBench from 23.5% to 27.1%. This further indicates that the benefits of DeNovoSWE are not due to accidental adaptation to a specific model, but stem from the high-quality long-horizon data itself.

Conclusion

The next stage for code agents is not just about fixing individual issues faster, but about understanding documents, planning architecture, organizing modules, implementing interfaces, and ultimately generating a complete, runnable software repository.

DeNovoSWE systematically frames this goal into a trainable, verifiable, and scalable dataset. It answers a key question: What kind of data can truly train agents with long-horizon software engineering capabilities?

The answer is not more fragmented code, nor simpler problems, but high-quality, structured, evaluation-aligned, anti-leakage, full-repository generation tasks.

Starting from a single document, rebuild the entire repository. This is the threshold that long-horizon code agents need to cross.

Reference: https://arxiv.org/pdf/2606.10728

This article is from the WeChat public account "AI Era," edited by LRST.

Трендовые криптовалюты

Связанные с этим вопросы

QWhat is the main contribution of the research from Renmin University of China's Gaoling School of Artificial Intelligence?

AThe research introduces and releases the DeNovoSWE dataset, which is the first long-horizon Doc2Repo training set focused on repository-level code generation from scratch in software engineering tasks.

QHow does the DeNovoSWE dataset address the challenge of long-horizon software engineering tasks?

AIt uses a sandboxed multi-agent workflow based on 'Divide & Conquer' and 'Draft-Critic-Repair' mechanisms to automatically construct high-quality task documents, ensuring they are well-organized and evaluation-aligned for whole-repository generation.

QWhat were the performance improvements observed after training a model on the DeNovoSWE dataset?

AThe Qwen3-30B-A3B-Instruct model showed significant improvement, increasing its performance on BeyondSWE-Doc2Repo from 5.8% to 47.2% and on NL2RepoBench from 4.3% to 23.0%.

QWhat are the two core standards mentioned for a high-quality task document in document-to-repository generation?

AFirst, the document must be well-organized, providing a clear overview and structured chapters. Second, it must be evaluation-aligned, describing key behaviors for verification without leaking implementation details.

QWhat is the purpose of the difficulty-aware trajectory filtering mechanism in DeNovoSWE?

AIt sets different filtering thresholds for tasks of varying difficulty levels to balance quality and diversity, ensuring that valuable but partially successful trajectories from complex repositories are not discarded.

Похожее

После трёх кварталов падения: получится ли у крипторынка стабилизироваться в третьем квартале?

Криптовалютный рынок пережил худший квартал с 2022 года, общая капитализация упала на 12,6% до $2,1 трлн, а дневные объемы торгов снизились на 20,9%. Впервые за три года сократилась и капитализация стейблкоинов. Основные факторы спада — отток средств из биткоин-ETF, распродажа биткоинов корпоративным казначейством Strategy и ужесточение монетарной политики ФРС. Второй квартал завершился чистым оттоком $46,7 млрд из ETF, при этом в июне отток достиг рекордных $45 млрд. Ключевым событием третьего квартала станет заседание ФРС 28–29 июля. Мягкий сигнал может поддержать биткоин в диапазоне $68 000–84 000 и вернуть приток в ETF, тогда как жесткая риторика способна сместить торговый диапазон к $50 000–56 000. Законодательная неопределенность также давит на рынок: прогресс по закону CLARITY Act, определяющему регуляторные границы, замедлился, и вероятность его принятия в 2026 году упала до 40–45%. Несмотря на общий спад, два сегмента показали рост: объемы на рынках предсказаний выросли на 48,7%, а торговля токенизированными коллекционными предметами увеличилась на 143%. Сектор RWA продолжает стабильно развиваться, достигнув $28,1 млрд. Рынок, вероятно, миновал фазу острой распродажи, но для устойчивого восстановления необходимы ясность от ФРС и регуляторный прогресс.

marsbit15 ч. назад

После трёх кварталов падения: получится ли у крипторынка стабилизироваться в третьем квартале?

marsbit15 ч. назад

BIT Торговые часы: BTC по-прежнему под давлением 200 EMA на недельном графике, после отскока возможен перезапуск нисходящего движения; секторы хранения данных и полупроводников, выросшие ночью, начали падение в вечерней сессии

**Краткий обзор рынка: BTC под давлением, коррекция на рынке акций, внимание к данным и событиям** Рынок криптовалют демонстрирует осторожное восстановление. Bitcoin торгуется около $66 000, сталкиваясь со значительным сопротивлением в районе $68 000, где сосредоточены объемные "застрявшие" позиции. Ключевые технические уровни — 200-недельная скользящая средняя (~$63 333) и 200-недельная EMA (~$68 328). Аналитики отмечают низкую ликвидность, характерную для летнего периода. На фондовом рынке после сильного роста во вторник наблюдается коррекция. Фьючерсы на основные индексы США снижаются. Акции полупроводниковой и памяти, которые резко выросли накануне, падают в ночных торгах. Исключением стал SMCI, который вырос более чем на 15% после оптимистичного прогноза. На общий настрой негативно влияют рост цен на нефть (более $91 за баррель Brent) и доходности государственных облигаций США (10-летние — около 4,64%), что возрождает инфляционные опасения. Азиатские рынки показали нестабильную динамику. Индекс KOSPI в Корее вырос на 0,74%, а японский Nikkei 225 снизился на 0,18%. Основной риск для региона — ослабление японской иены до минимумов с 1986 года, что повышает вероятность вмешательства властей. **Ключевые события для наблюдения:** * **24 июля:** Финансовые отчеты Alphabet (Google), Tesla, IBM. Событие AMD, посвященное ИИ. * **25 июля:** Решение по процентной ставке ЕЦБ и пресс-конференция Кристин Лагард. Финансовые отчеты Intel, American Airlines, Honeywell и других. Данные по числу первичных заявок на пособие по безработице в США.

marsbit15 ч. назад

BIT Торговые часы: BTC по-прежнему под давлением 200 EMA на недельном графике, после отскока возможен перезапуск нисходящего движения; секторы хранения данных и полупроводников, выросшие ночью, начали падение в вечерней сессии

marsbit15 ч. назад

Бывший глава CFTC и президент Circle Тарберт: призывает к долгосрочной стратегии, но сам выводит $30 млн

Бывший председатель CFTC и президент Circle Хит Тарберт, публично пропагандируя долгосрочное видение компании для инвесторов на фоне падения акций на 70% с пиковых значений, сам активно распродавал свои акции CRCL. С момента IPO Circle он по плану 10b5-1 продал более 360 тысяч акций, выручив около 30 миллионов долларов, и при этом ни разу не докупал акции на открытом рынке. Эта разница между его публичными заявлениями и личными действиями вызвала критику. Карьера Тарберта демонстрирует классическое использование «вращающейся двери» между регулирующими органами и частным сектором. Уйдя с поста председателя CFTC в 2021 году, через 27 дней он занял должность главного юрисконсульта в маркет-мейкере Citadel Securities, который в тот момент находился под пристальным вниманием из-за скандала с акциями GameStop. Позже, уже работая в Citadel, Тарберт выступал за расширение полномочий CFTC на крипторынок, в то время как его работодатель планировал выход на этот рынок. В 2023 году Тарберт присоединился к Circle, где его опыт и связи сыграли ключевую роль в успешном проведении IPO компании в 2025 году. Однако его последующие массовые продажи акций, совпавшие с падением котировок и его же призывами к долгосрочным инвестициям, ставят под сомнение искренность его уверенности в будущем компании. Критики видят в его карьере образец превращения регуляторного опыта и политических связей в личную выгоду, в то время как риски несут обычные инвесторы, верящие его нарративам.

marsbit15 ч. назад

Бывший глава CFTC и президент Circle Тарберт: призывает к долгосрочной стратегии, но сам выводит $30 млн

marsbit15 ч. назад

Gate Research: Взлет «уолл-стритизации» криптофинансовых продуктов — это конкуренция или интеграция?

**Аналитический обзор Gate Research Institute: Волна "уолл-стритизации" криптофинансовых продуктов — конкуренция или интеграция?** Создание биткоина в 2009 году было ответом на финансовый кризис и стремлением построить децентрализованную систему без доверенных посредников. Однако к 2026 году значительная часть биткоинов (около 7,14%) хранится через ETF таких гигантов, как BlackRock. Это символизирует глубокую интеграцию традиционных финансов (TradFi) и крипторынка. С появлением биткоин-ETF, фьючерсов, RWA (токенизированных реальных активов) и государственных облигаций на блокчейне, традиционные институты получают всё больше влияния на выпуск, ценообразование, кастодию и дистрибуцию криптоактивов. Однако это не одностороннее поглощение, а взаимодополняющая конвергенция. Криптосфера приносит TradFi глобальную 24/7 ликвидность и программируемость, а TradFi предоставляет криптосфере регулируемые каналы, институциональное доверие и массовый доступ. Яркий пример — две противоположные, но ведущие к одной цели траектории: * **Путь А:** Криптобиржи, такие как Gate, начинают с токенизированных акций и CFD, а затем напрямую подключаются к инфраструктуре традиционных брокеров, предлагая реальную торговлю акциями США, Гонконга и Кореи за стейблкоины. * **Путь Б:** Традиционные брокеры, такие как Robinhood, интегрируют криптоактивы, а затем создают собственные Layer 2 для токенизации своих акций, стремясь к круглосуточной торговле. Обе стратегии нацелены на создание **универсального финансового аккаунта будущего** — единой точки доступа к акциям, криптовалютам, ETF, токенизированным облигациям и другим активам. Ключевой конкурентной борьбой становится не между CEX и брокерами, а между этими "супераккаунтами". Параллельно, слой RWA и токенизированных гособлигаций растёт даже на медвежьем рынке, становясь "прослойкой" для объединения капиталов. Хотя этот рынок (около $150 млрд токенизированных казначейских облигаций) ещё мал по сравнению с традиционным ($30 трлн), он демонстрирует устойчивый структурный тренд, привлекающий крупнейшие финансовые институты (JPMorgan, DTCC и др.). **Вывод:** "Уолл-стритизация" — это не поражение идеалов децентрализации, а формирование новой гибридной модели. Децентрализованные протоколы продолжают работать на базовом уровне, в то время как на уровне приложений и пользовательского опыта формируется более эффективный, глобальный и свободный объединённый рынок капитала, где активы TradFi и DeFi торгуются бок о бок в одном интерфейсе. Уолл-стрит не завоевала криптосферу, а криптосфера не обошла Уолл-стрит — они совместно строят новые финансовые рельсы.

marsbit15 ч. назад

Gate Research: Взлет «уолл-стритизации» криптофинансовых продуктов — это конкуренция или интеграция?

marsbit15 ч. назад

S&P Dow Jones и Pantera выпускают криптоиндекс, биткоин оставлен за бортом из-за «отсутствия прибыли»

Стандард энд Пурс (S&P Dow Jones Indices) совместно с Pantera Capital запускает индекс S&P Pantera Digital Asset Index. Новый индекс включает 18 токенов, но исключает биткоин, XRP и мем-токены. Критерием отбора послужила «финансовая жизнеспособность», по аналогии с S&P 500: протокол должен демонстрировать положительный доход в течение нескольких кварталов и распределять стоимость среди держателей токенов. Это первое применение фундаментального подхода к оценке криптоактивов со стороны крупнейшего в мире провайдера индексов. В первую пятерку активов вошли Ethereum (ETH), BNB, Solana (SOL), TRON (TRX) и Hyperliquid (HYPE). Pantera отмечает, что совокупный годовой доход всех протоколов индекса превышает $30 млрд. Биткоин был исключен по трем причинам: он рассматривается как монетарный актив, доступный через отдельные ETF; он не генерирует доход протокола; а смешивание его в индексе с доходными активами затрудняет фундаментальный анализ для институциональных инвесторов. Пока индекс является эталонным, но Pantera ведет переговоры о создании на его основе ETF и других инвестиционных продуктов. Компания также отмечает значительный разрыв на рынке: многие институциональные инвесторы готовы к аллокации в криптоактивы, но им не хватает структурированных продуктов, фокусирующихся на активах с реальной экономической деятельностью.

marsbit16 ч. назад

S&P Dow Jones и Pantera выпускают криптоиндекс, биткоин оставлен за бортом из-за «отсутствия прибыли»

marsbit16 ч. назад

Торговля

Спот

Популярные статьи

Как купить RE

Добро пожаловать на HTX.com! Мы сделали приобретение Re (RE) простым и удобным. Следуйте нашему пошаговому руководству и отправляйтесь в свое крипто-путешествие.Шаг 1: Создайте аккаунт на HTXИспользуйте свой адрес электронной почты или номер телефона, чтобы зарегистрироваться и бесплатно создать аккаунт на HTX. Пройдите удобную регистрацию и откройте для себя весь функционал.Создать аккаунтШаг 2: Перейдите в Купить криптовалюту и выберите свой способ оплатыКредитная/Дебетовая Карта: Используйте свою карту Visa или Mastercard для мгновенной покупки Re (RE).Баланс: Используйте средства с баланса вашего аккаунта HTX для простой торговли.Третьи Лица: Мы добавили популярные способы оплаты, такие как Google Pay и Apple Pay, для повышения удобства.P2P: Торгуйте напрямую с другими пользователями на HTX.Внебиржевая Торговля (OTC): Мы предлагаем индивидуальные услуги и конкурентоспособные обменные курсы для трейдеров.Шаг 3: Хранение Re (RE)После приобретения вами Re (RE) храните их в своем аккаунте на HTX. В качестве альтернативы вы можете отправить их куда-либо с помощью перевода в блокчейне или использовать для торговли с другими криптовалютами.Шаг 4: Торговля Re (RE)С легкостью торгуйте Re (RE) на спотовом рынке HTX. Просто зайдите в свой аккаунт, выберите торговую пару, совершайте сделки и следите за ними в режиме реального времени. Мы предлагаем удобный интерфейс как для начинающих, так и для опытных трейдеров.

270 просмотров всегоОпубликовано 2026.06.18Обновлено 2026.06.29

Как купить RE

Обсуждения

Добро пожаловать в Сообщество HTX. Здесь вы сможете быть в курсе последних новостей о развитии платформы и получить доступ к профессиональной аналитической информации о рынке. Мнения пользователей о цене на RE (RE) представлены ниже.

活动图片