Slow .SD[.N] compare to last(.SD) with groupby

Open
#4,809 6 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
35/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Stale
Tech stack
r
Domain
data, performance

Research direction

Start by reproducing the grouped expressions DT[,.SD[.N], by=col] and DT[,.last(.SD), by=col] on the large data described, then compare them with the corresponding .SD[1] and first(.SD) cases. Use the reported performance difference as the investigation target; done means its cause is explained and any confirmed regression is addressed without worsening the matching first-row case.

Written by the indexing model from the issue text.

Description

GForce

I noticed there is significant speed difference between
DT[,.SD[.N], by=col] and
DT[,.last(.SD), by=col]

for relatively large data (10 * 1.8M rows).

image

DT[,.SD[1],by=col] and DT[,first(.SD), by=col] did not show difference in performance, however.

image
Dominant language
R
Stars
3.9k
Forks
1.1k
Avg merge
14h 4m
Merged PRs (30d)
4

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Rdatatable/data.table

All issues in Rdatatable/data.table

Similar issues

More R issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.