JuliaData / JuliaData/CSV.jl

CSV.Chunks splits file into uneven chunks

Open
#1,122 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Julia
Stars
506
Forks
150
Avg merge
4d 16h
Merged PRs (30d)
7

Description

In the following example the "a" variable does not have a consistent size.

using CSV, DataFrames

number_of_lines = 10^6
CSV.write("data.csv", DataFrame(rand(number_of_lines, 10), :auto))

steps = 10

@time for chunk in CSV.Chunks("data.csv"; ntasks=steps)
    a = chunk |> DataFrame
    display(size((a)))
end

Please also have a look at this discourse post, for context.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the CSV.Chunks("data.csv"; ntasks=steps) reproduction in the issue and compare the displayed DataFrame sizes across iterations. Read the linked Discourse post for context on skipped lines and allocation, then consider the issue complete when the chunk sizes are consistent for the example without regressing CSV parsing behavior.

Written by the indexing model from the issue text.

Assessment

Tech stack
julia
Domain
data-engineering
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.