THUDM / THUDM/slime

Questions about the Integration of Off-policy Improvements

Open
#1,019 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
8.5k
Forks
1.3k
Avg merge
5h 36m
Merged PRs (30d)
22

Description

Dear Slime Team,

First and foremost, I would like to express my sincere appreciation for your outstanding work on the Slime framework. It has been an invaluable resource for the community.

I am writing to inquire about the potential roadmap for integrating off-policy improvements or experience replay mechanisms into the framework. Specifically, I am interested in whether there are plans to adopt techniques such as the Decoupled PPO proposed in AReal, or the Experience Replay strategy demonstrated in RLEP.

Could you kindly provide some insight into the following:

  • Feasibility & Complexity: Based on the current architecture of Slime, how would you assess the technical feasibility and complexity of implementing such experience replay or off-policy algorithms?
  • Strategic Perspective: I would greatly value your team's professional perspective on the role of off-policy algorithms within the current landscape of LLM post-training. Do you view these as a priority direction for the framework?

Thank you very much for your time and your continued contributions to the field.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

No files, tests, or entry points are identified in the issue. Clarify whether the request is for a roadmap discussion or an implementation, then define the targeted off-policy or experience-replay approach and what completion would require.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
ai, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.