AI4Finance-Foundation / AI4Finance-Foundation/FinRL

ElegantRL training on paper trading notebook doesn't show the model learning

Ouverte
#1,112 1 commentaire 0 réactions 1 personne assignée Réclamée par @Yonv1943 Voir sur GitHub
bug
Langage dominant
Jupyter Notebook
Étoiles
16.3k
Forks
3.5k
Métriques de merge des PR
Aucune PR mergée en 30 j

Description

Using the following ERL parameters:

ERL_PARAMS = {"learning_rate": 3e-6,"batch_size": 2048,"gamma": 0.985,
"seed":312,"net_dimension":[128,64], "target_step":50000, "eval_gap":30,
"eval_times":5}

When I run training on a larger dataset as seen below

train(start_date = '2005-01-01',
end_date = '2022-12-31',
ticker_list = ticker_list,
data_source = 'alpaca',
time_interval= '1Min',
technical_indicator_list= INDICATORS,
drl_lib='elegantrl',
env=env,
model_name='ppo',
if_vix=True,
API_KEY = API_KEY,
API_SECRET = API_SECRET,
API_BASE_URL = API_BASE_URL,
erl_params=ERL_PARAMS,
cwd='./papertrading_erl_orig', #current_working_dir
break_step=1e7)

My output for the training is:

| `step`: Number of samples, or total training steps, or running times of `env.step()`.
| `time`: Time spent from the start of training to this moment.
| `avgR`: Average value of cumulative rewards, which is the sum of rewards in an episode.
| `stdR`: Standard dev of cumulative rewards, which is the sum of rewards in an episode.
| `avgS`: Average of steps in an episode.
| `objC`: Objective of Critic network. Or call it loss function of critic network.
| `objA`: Objective of Actor network. It is the average Q value of the critic network.
| step time | avgR stdR avgS | objC objA
| 2.00e+04 11 | -0.49 0.02 12345 | 0.05 0.19
| 4.00e+04 22 | -0.49 0.02 12345 | 0.00 0.18
| 6.00e+04 33 | -0.49 0.02 12345 | 0.00 0.19
| 8.00e+04 44 | -0.49 0.01 12345 | 0.00 0.19
| 1.00e+05 55 | -0.48 0.03 12345 | 0.00 0.18
| 1.20e+05 66 | -0.49 0.03 12345 | 0.00 0.17
| 1.40e+05 77 | -0.48 0.02 12345 | 0.00 0.18
| 1.60e+05 88 | -0.50 0.02 12345 | 0.00 0.19
| 1.80e+05 99 | -0.48 0.02 12345 | 0.00 0.18
| 2.00e+05 111 | -0.48 0.03 12345 | 0.00 0.18
| 2.20e+05 122 | -0.48 0.01 12345 | 0.00 0.19
| 2.40e+05 133 | -0.49 0.02 12345 | 0.00 0.18
| 2.60e+05 144 | -0.48 0.03 12345 | 0.00 0.19
| 2.80e+05 155 | -0.49 0.01 12345 | 0.00 0.19
| 3.00e+05 166 | -0.48 0.02 12345 | 0.00 0.19

this output continues even after the training has ran for hours. Shouldn't the avgR and objA values increase slowly over time?

Is this output normal? I have tweaked the ERL params and changed the batch size, learning rate and other setting but I always get the same results. If I change the dataset to a smaller interval, for example 2021-01-01 to 2022-12-31 the avgR increases to over 50 points but again it stays pretty constant.

When I run the test on unseen data I get mixed results. When using SB3 I can see through the explained_variance if the model is learning, in this case I have no clue, is this a bug or normal behavior?

Guide de contribution

Aucun guide de contribution indexé pour ce dépôt

Évaluation

Cette issue n'a pas encore été évaluée.

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.