QuantConnect / QuantConnect/Lean

Improves PandasConverter

Open
#6,286 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

feature performance
Dominant language
C#
Stars
21.7k
Forks
5.3k
Avg merge
2d 22h
Merged PRs (30d)
34

Description

Expected Behavior

Faster and less computationally expensive pandas converter.

Actual Behavior

The converter will hold all the information/dataframe in the memory by Symbol and then concatenate the dataframe. Concatenating hundreds of Symbols with thousands of rows is a cheap operation.

Potential Solution

N/A. Research alternatives.

Reproducing the Problem

History request of minute resolution data of hundred of Symbols for 1000 bars.

Checklist
  • I have completely filled out this template
  • I have confirmed that this issue exists on the current master branch
  • I have confirmed that this is not a duplicate issue by searching issues

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Locate the PandasConverter entry point and reproduce the stated history request using minute-resolution data for hundreds of Symbols and 1,000 bars. Research alternatives to retaining each dataframe before concatenation, then compare memory use and runtime against the current behavior. Done means the converter is measurably faster and less computationally expensive without changing its results.

Written by the indexing model from the issue text.

Assessment

Tech stack
csharp, pandas
Domain
data, performance
Issue type
Refactor
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.