CSV.Chunks splits file into uneven chunks
Open
Nobody has claimed this yet.
- Dominant language
- Julia
- Stars
- 506
- Forks
- 150
- Avg merge
- 4d 16h
- Merged PRs (30d)
- 7
Description
In the following example the "a" variable does not have a consistent size.
using CSV, DataFrames
number_of_lines = 10^6
CSV.write("data.csv", DataFrame(rand(number_of_lines, 10), :auto))
steps = 10
@time for chunk in CSV.Chunks("data.csv"; ntasks=steps)
a = chunk |> DataFrame
display(size((a)))
end
Please also have a look at this discourse post, for context.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the CSV.Chunks("data.csv"; ntasks=steps) reproduction in the issue and compare the displayed DataFrame sizes across iterations. Read the linked Discourse post for context on skipped lines and allocation, then consider the issue complete when the chunk sizes are consistent for the example without regressing CSV parsing behavior.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- julia
- Domain
- data-engineering
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 38/100