Skip to main content

𝗩đ—Čđ—źđ—żđ—°đ—”-đ—„đŸ­ – the first 𝗿đ—Čđ—œđ—żđ—Œđ—±đ˜‚đ—°đ˜đ—¶đ—Œđ—» đ—Œđ—ł 𝗗đ—Čđ—Čđ—œđ˜€đ—Čđ—Č𝗾-đ—„đŸ­ (𝘇đ—Čđ—żđ—Œ) with reinforcement learning

What do you think of something like this?

For training reasoning and search-augmented LLM agents with reinforcement learning.

This is a step towards training an đ—Œđ—œđ—Čđ—»-đ˜€đ—Œđ˜‚đ—żđ—°đ—Č đ—ąđ—œđ—Čđ—»đ—”đ—œ “𝗗đ—Čđ—Čđ—œ 𝗿đ—Č𝘀đ—Čđ—źđ—żđ—°đ—”â€ via RL.

𝟯𝗕 𝗯𝗼𝘀đ—Č 𝗟𝗟𝗠𝘀—including not just đ—€đ˜„đ—Čđ—» 𝟼.đŸ± but also 𝗟đ—č𝗼đ—ș𝗼 𝟯.𝟼—learn to 𝗿đ—Čđ—źđ˜€đ—Œđ—» and 𝗰𝗼đ—čđ—č 𝘀đ—Čđ—źđ—żđ—°đ—” đ—Čđ—»đ—Žđ—¶đ—»đ—Č𝘀 all on their own.

We follow Deepseek R1-zero, starting with a base LLM, prompts, and ground-truth rewards. Then, we apply 𝗿đ—Čđ—¶đ—»đ—łđ—Œđ—żđ—°đ—Čđ—șđ—Čđ—»đ˜ đ—čđ—Čđ—źđ—żđ—»đ—¶đ—»đ—Ž (RL). Our experiments are conducted on 𝗡𝗼𝘁𝘂𝗿𝗼đ—č đ—€đ˜‚đ—Čđ˜€đ˜đ—¶đ—Œđ—»đ˜€ (đ—Ąđ—€), a factual QA dataset in which LLMs struggle with direct answers, making search engine calls crucial. The only supervision? A 𝗿𝘂đ—čđ—Č-𝗯𝗼𝘀đ—Čđ—± đ—Œđ˜‚đ˜đ—°đ—Œđ—șđ—Č 𝗿đ—Čđ˜„đ—źđ—żđ—± (string exact match) to determine correctness.

We first experiment with đ—„đ—Ÿ đ˜đ˜‚đ—»đ—¶đ—»đ—Ž đ‘€đ‘–đ‘Ąâ„Žđ‘œđ‘ąđ‘Ą search engine access, letting the 𝗟𝗟𝗠 (𝗟đ—č𝗼đ—ș𝗼 𝟯.𝟼-𝟯𝗕-𝗯𝗼𝘀đ—Č) answer questions on its own. Initially, the model produces đ—±đ˜‚đ—șđ—ș𝘆 đ—Œđ˜‚đ˜đ—œđ˜‚đ˜đ˜€, but through RL, it đ—Žđ—żđ—źđ—±đ˜‚đ—źđ—čđ—č𝘆 đ—čđ—Čđ—źđ—żđ—»đ˜€ to generate meaningful answers.

Image

Next, we đ—¶đ—»đ˜€đ˜đ—żđ˜‚đ—°đ˜ the 𝗟𝗟𝗠 (𝗟đ—č𝗼đ—ș𝗼 𝟯.𝟼-𝟯𝗕-𝗯𝗼𝘀đ—Č) that it can 𝗰𝗼đ—čđ—č 𝗼 𝘀đ—Čđ—źđ—żđ—°đ—” đ—Čđ—»đ—Žđ—¶đ—»đ—Č to retrieve relevant information. đ—Šđ˜‚đ—żđ—œđ—żđ—¶đ˜€đ—¶đ—»đ—Žđ—č𝘆, even đ˜„đ—¶đ˜đ—”đ—Œđ˜‚đ˜ any supervised fine-tuning (SFT), the 𝗯𝗼𝘀đ—Č 𝗟𝗟𝗠 đ—čđ—Čđ—źđ—żđ—»đ˜€ đ˜đ—Œ 𝗰𝗼đ—čđ—č đ˜đ—”đ—Č 𝘀đ—Čđ—źđ—żđ—°đ—” đ—Čđ—»đ—Žđ—¶đ—»đ—Č, đ—¶đ—»đ˜đ—Čđ—żđ—œđ—żđ—Č𝘁 𝘀đ—Čđ—źđ—żđ—°đ—” 𝗿đ—Č𝘀𝘂đ—č𝘁𝘀, đ—źđ—»đ—± đ—źđ—»đ˜€đ˜„đ—Č𝗿 đ—Ÿđ˜‚đ—Čđ˜€đ˜đ—¶đ—Œđ—»đ˜€â€”đ—źđ—čđ—č đ˜đ—”đ—żđ—Œđ˜‚đ—Žđ—” đ—„đ—Ÿ!

Image

We compare the performance of the 𝗟𝗟𝗠 đ˜„đ—¶đ˜đ—”đ—Œđ˜‚đ˜ 𝘀đ—Čđ—źđ—żđ—°đ—” đ—Čđ—»đ—Žđ—¶đ—»đ—Č 𝗼𝗰𝗰đ—Č𝘀𝘀 vs. 𝗟𝗟𝗠 đ˜„đ—¶đ˜đ—” 𝘀đ—Čđ—źđ—żđ—°đ—”-𝗼𝘂𝗮đ—șđ—Čđ—»đ˜đ—Čđ—± đ—čđ—Čđ—źđ—żđ—»đ—¶đ—»đ—Ž via RL. The 𝘀đ—Čđ—źđ—żđ—°đ—”-đ—Čđ—»đ—źđ—Żđ—čđ—Čđ—± đ—șđ—Œđ—±đ—Čđ—č đ˜„đ—¶đ—»đ˜€!

Image

When training 𝗟đ—č𝗼đ—ș𝗼 𝟯.𝟼-𝟯𝗕-𝗯𝗼𝘀đ—Č with 𝘀đ—Čđ—źđ—żđ—°đ—” đ—Čđ—»đ—Žđ—¶đ—»đ—Č 𝗰𝗼đ—čđ—čđ—¶đ—»đ—Ž, the response length follows an interesting trend:

đ—™đ—¶đ—żđ˜€đ˜, đ—¶đ˜ đ—±đ—Č𝗰𝗿đ—Č𝗼𝘀đ—Č𝘀—the model learns to đ—źđ˜ƒđ—Œđ—¶đ—± đ—Č𝘅𝗰đ—Čđ˜€đ˜€đ—¶đ˜ƒđ—Č đ—±đ˜‚đ—șđ—ș𝘆 đ˜„đ—Œđ—żđ—±đ˜€. đ—§đ—”đ—Čđ—», đ—¶đ˜ đ—¶đ—»đ—°đ—żđ—Č𝗼𝘀đ—Č𝘀—as it learns to 𝗰𝗼đ—čđ—č đ˜đ—”đ—Č 𝘀đ—Čđ—źđ—żđ—°đ—” đ—Čđ—»đ—Žđ—¶đ—»đ—Č đ—źđ—»đ—± 𝗿đ—Čđ—źđ˜€đ—Œđ—» effectively. Since đ—Ąđ—€ đ—¶đ˜€ 𝗼 𝗿đ—Čđ—čđ—źđ˜đ—¶đ˜ƒđ—Čđ—č𝘆 đ˜€đ—¶đ—șđ—œđ—čđ—Č 𝘁𝗼𝘀𝗾, the response length đ˜€đ˜đ—źđ—Żđ—¶đ—čđ—¶đ˜‡đ—Č𝘀 𝗼𝘁 ~đŸ±đŸŹđŸŹ đ˜đ—Œđ—žđ—Čđ—»đ˜€.

Image

We experiment with đ—€đ˜„đ—Čđ—»đŸź.đŸ±-𝟯𝗕-𝗯𝗼𝘀đ—Č and đ—€đ˜„đ—Čđ—»đŸź.đŸ±-𝟳𝗕-𝗯𝗼𝘀đ—Č under both with/without search engine RL settings. 𝗜𝘁 đ˜„đ—Œđ—żđ—žđ˜€ đ—łđ—Œđ—ż đ—Żđ—Œđ˜đ—”! Interestingly, in the 𝘀đ—Čđ—źđ—żđ—°đ—”-𝗼𝘂𝗮đ—șđ—Čđ—»đ˜đ—Čđ—± 𝘀đ—Čđ˜đ˜đ—¶đ—»đ—Ž, the 𝟯𝗕 đ—șđ—Œđ—±đ—Čđ—č đ—źđ—°đ—”đ—¶đ—Č𝘃đ—Č𝘀 đ—œđ—Čđ—żđ—łđ—Œđ—żđ—șđ—źđ—»đ—°đ—Č đ—°đ—Œđ—șđ—œđ—źđ—żđ—źđ—Żđ—čđ—Č đ˜đ—Œ đ˜đ—”đ—Č 𝟳𝗕 đ—șđ—Œđ—±đ—Čđ—č. đ—›đ˜†đ—œđ—Œđ˜đ—”đ—Čđ˜€đ—¶đ˜€: When an 𝗟𝗟𝗠 đ—¶đ˜€ đ—°đ—Œđ—»đ—»đ—Č𝗰𝘁đ—Čđ—± đ˜đ—Œ đ—Č𝘅𝘁đ—Čđ—żđ—»đ—źđ—č đ—¶đ—»đ—łđ—Œđ—żđ—șđ—źđ˜đ—¶đ—Œđ—», its 𝗿đ—Čđ—źđ˜€đ—Œđ—»đ—¶đ—»đ—Ž đ—źđ—Żđ—¶đ—čđ—¶đ˜đ˜† đ—ș𝗼𝘆 đ—»đ—Œđ˜ đ—»đ—Č𝗰đ—Čđ˜€đ˜€đ—źđ—żđ—¶đ—č𝘆 𝗿đ—Čđ—Ÿđ˜‚đ—¶đ—żđ—Č 𝗼 đ—č𝗼𝗿𝗮đ—Č đ—șđ—Œđ—±đ—Čđ—č đ˜€đ—¶đ˜‡đ—Č.

Image

đ—•đ—Œđ˜đ—” 𝗯𝗼𝘀đ—Č đ—źđ—»đ—± đ—¶đ—»đ˜€đ˜đ—żđ˜‚đ—°đ˜đ—¶đ—Œđ—» đ—șđ—Œđ—±đ—Čđ—č𝘀 đ˜„đ—Œđ—żđ—ž! The đ—¶đ—»đ˜€đ˜đ—żđ˜‚đ—°đ˜đ—¶đ—Œđ—» đ—șđ—Œđ—±đ—Čđ—č converges 𝗳𝗼𝘀𝘁đ—Č𝗿 and starts from 𝗼 𝗯đ—Č𝘁𝘁đ—Č𝗿 đ—¶đ—»đ—¶đ˜đ—¶đ—źđ—č đ—œđ—Čđ—żđ—łđ—Œđ—żđ—șđ—źđ—»đ—°đ—Č. However, the đ—łđ—¶đ—»đ—źđ—č đ—œđ—Čđ—żđ—łđ—Œđ—żđ—șđ—źđ—»đ—°đ—Č of both models is 𝘃đ—Č𝗿𝘆 đ˜€đ—¶đ—șđ—¶đ—č𝗼𝗿. This suggests that while đ—¶đ—»đ˜€đ˜đ—żđ˜‚đ—°đ˜đ—¶đ—Œđ—» đ˜đ˜‚đ—»đ—¶đ—»đ—Ž 𝗼𝗰𝗰đ—Čđ—čđ—Č𝗿𝗼𝘁đ—Č𝘀 đ—čđ—Čđ—źđ—żđ—»đ—¶đ—»đ—Ž, 𝗿đ—Čđ—¶đ—»đ—łđ—Œđ—żđ—°đ—Čđ—șđ—Čđ—»đ˜ đ—čđ—Čđ—źđ—żđ—»đ—¶đ—»đ—Ž đ—°đ—źđ—» đ—Żđ—żđ—¶đ—±đ—Žđ—Č đ˜đ—”đ—Č đ—Žđ—źđ—œ đ—Œđ˜ƒđ—Č𝗿 đ˜đ—¶đ—șđ—Č.

Image

We experiment with đ—€đ˜„đ—Čđ—»đŸź.đŸ±-𝟯𝗕-𝗯𝗼𝘀đ—Č, 𝗟đ—č𝗼đ—ș𝗼𝟯.𝟼-𝟯𝗕-𝗯𝗼𝘀đ—Č, đ—źđ—»đ—± đ—€đ˜„đ—Čđ—»đŸź.đŸ±-𝟳𝗕-𝗯𝗼𝘀đ—Č—and đ˜đ—”đ—Č𝘆 𝗼đ—čđ—č đ˜„đ—Œđ—żđ—ž! This is đ—»đ—Œđ˜đ—źđ—Żđ—č𝘆 đ—±đ—¶đ—łđ—łđ—Č𝗿đ—Čđ—»đ˜ đ—łđ—żđ—Œđ—ș đ—șđ—źđ˜đ—” 𝗿đ—Čđ—źđ˜€đ—Œđ—»đ—¶đ—»đ—Ž, where only the đ—€đ˜„đ—Čđ—»đŸź.đŸ± 𝘀đ—Čđ—żđ—¶đ—Č𝘀 models succeed.

Image

The 𝗟𝗟𝗠 đ—čđ—Čđ—źđ—żđ—»đ˜€ đ˜đ—Œ đ—œđ—Čđ—żđ—łđ—Œđ—żđ—ș đ—ș𝘂đ—čđ˜đ—¶-đ˜đ˜‚đ—żđ—» 𝘀đ—Čđ—źđ—żđ—°đ—” đ—Čđ—»đ—Žđ—¶đ—»đ—Č 𝗰𝗼đ—čđ—č𝘀, refining its queries step by step to gather more relevant information. This showcases its ability to đ—¶đ˜đ—Čđ—żđ—źđ˜đ—¶đ˜ƒđ—Čđ—č𝘆 đ—¶đ—șđ—œđ—żđ—Œđ˜ƒđ—Č 𝗿đ—Čđ˜đ—żđ—¶đ—Č𝘃𝗼đ—č đ—źđ—»đ—± 𝗿đ—Čđ—źđ˜€đ—Œđ—»đ—¶đ—»đ—Žâ€”a key capability for real-world research agents!

Image

Our framework supports 𝗳đ—čđ—Čđ˜…đ—¶đ—Żđ—čđ—Č 𝘀đ—Čđ—źđ—żđ—°đ—” đ—Čđ—»đ—Žđ—¶đ—»đ—Č đ—°đ—”đ—Œđ—¶đ—°đ—Č𝘀, including: đ—Ÿđ—Œđ—°đ—źđ—č 𝗿đ—Čđ˜đ—żđ—¶đ—Č𝘃đ—Č𝗿𝘀 (sparse/dense) đ—ąđ—»đ—čđ—¶đ—»đ—Č 𝘀đ—Čđ—źđ—żđ—°đ—” đ—Čđ—»đ—Žđ—¶đ—»đ—Č𝘀 (Google, Bing, etc.) đ—–đ˜‚đ˜€đ˜đ—Œđ—ș 𝘀đ—Čđ—źđ—żđ—°đ—” đ—Čđ—»đ—Žđ—¶đ—»đ—Č𝘀—Launch your own on any corpus and integrate it with RL effortlessly!


The pipeline is based on verl (https://github.com/volcengine/verl), a highly efficient RL framework.

Fully open source

Experimental logs
Github

Status: Rejected

Log in to comment and vote

No comments yet

Be the first to share your thoughts.