microsoft / microsoft/qlib

Question about abnormal validation reward of PPO baseline in order execution code

Open
#2,211 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

question
Dominant language
Python
Stars
48.7k
Forks
7.7k
PR merge metrics
No merged PRs in 30d

Description

Specifically, when running the baseline PPO strategy, the reward on the validation set remains constant every time. After checking the training log, I found that the model takes a "sell all" action every time. May I ask what might be causing this issue? How should I do?

The code is almost unchanged, except for modifications to workflow.py and order_gen.py. Because running the original code kept throwing this error: TypeError: Cannot compare Timestamp with datetime.date. Use ts == pd.Timestamp(date) or ts.date() == date instead.

Image

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reviewing the reported changes in workflow.py and order_gen.py, then reproduce the original TypeError: Cannot compare Timestamp with datetime.date. Compare the training log and validation behavior to determine why the PPO baseline repeatedly chooses "sell all." Done means the timestamp error is accounted for and the validation reward and action behavior are no longer constant.

Written by the indexing model from the issue text.

Assessment

Tech stack
pandas, python
Domain
fintech-quant, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
42/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.