Drzero
Dr. Zero Self-Evolving Search Agents without Training Data
About
This repository contains the code for Dr. Zero: Self-Evolving Search Agents without Training Data. In this work, we introduce Dr. Zero, a framework enabling search agents to effectively self-evolve without any training data. In particular, we design a self-evolution feedback loop where a proposer generates diverse questions to train a solver initialized from the same base model. As the solver evolves, it incentivizes the proposer to produce increasingly difficult yet solvable tasks, thus establishing an automated curriculum to refine both agents. To enhance training efficiency, we also introduce hop-grouped relative policy optimization (HRPO). This method clusters structurally similar questions to construct group-level baselines, effectively minimizing the sampling overhead in evaluating each query's individual difficulty and solvability. Consequently, HRPO significantly reduces the compute requirements for solver training without compromising performance or stability. Extensive experiment results demonstrate that the data-free Dr.
Open Source Health
- Stars
- 523
- Forks
- 65
- License
- Not stated
- Last commit
- 7 months ago
Resources & Links
Related Categories
Vendor
Facebookresearch
Publisher of Fairseq
Quick Links
Open Source
More by Facebookresearch
Related Products
RPG_KDD2025
This repository provides the code for implementing RPG described in our KDD'25 paper "Generating Long Semantic IDs in Parallel for Recommendation".
More from this vendor
Seamless Communication
Foundational Models for State-of-the-Art Speech and Text Translation
More from this vendor
OrienterNet
Source Code for Paper "OrienterNet Visual Localization in 2D Public Maps with Neural Matching"
More from this vendor
Matrix
Matrix (Multi-Agent daTa geneRation Infra and eXperimentation framework) is a versatile engine for multi-agent conversational data generation.
More from this vendor
Spider
A general physic-based retargeting framework.
More from this vendor
