fread for complex input

Open
#4,466 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
35/100
Issue type
Feature
Clarity
Mostly clear
Activity status
Stale
Tech stack
c, r
Domain
data

Research direction

Review the related discussions in #3690 and #4464, then inspect parse_double_regular_core in fread.c and the Rcomplex handling in freadR.c. Determine how complex values should be stored and how type auto-detection affects the sample scan; done means fread can recognize and return complex columns without an unacceptable efficiency cost.

Written by the indexing model from the issue text.

Description

fread

Follow-up to #3690 and #4464.

Originally I said it would take significant effort but #4464 would make adding a parser for complex straightforward -- parse_double_regular_core followed by + or -, then parse_double_regular_core followed by i. The only complication is about storage -- fread.c returns a struct and freadR.c applies Rcomplex?

It also raises a question of efficiency -- how much of an impact does it have to add auto-detect types? Since plain numeric and complex columns look identical for 10-15 characters (maybe), it could be inefficient to offer this "lookalike" parser. OTOH it's only applied to a small sample of rows.

Dominant language
R
Stars
3.9k
Forks
1.1k
Avg merge
14h 4m
Merged PRs (30d)
4

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Rdatatable/data.table

All issues in Rdatatable/data.table

Similar issues

More R issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.