lance-format / lance-format/lance

Support byte-sized batch limits in file reader

Open
#6,387 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

enhancement
Dominant language
Rust
Stars
7.1k
Forks
852
Avg merge
3d 18h
Merged PRs (30d)
272

Description

Summary

Allow the batch size to be specified in bytes instead of rows when reading lance files. This gives callers control over memory usage per batch, which is especially important for variable-width data where row count is a poor proxy for memory consumption.

Approach

Add an estimate_decoded_bytes method to the page and field decoder traits. Before each drain, the batch stream queries this to compute a row count that fits the byte budget. A post-decode feedback loop measures actual batch sizes and corrects the estimate for subsequent batches. No file format changes required — all estimates use data already available at decode time.

See investigation notes and worktree at feat-byte-sized-batches-file-reader for full design details.

Tasks

  • Add batch_size_bytes config option — wire batch_size_bytes: Option<u64> through SchedulerDecoderConfig into BatchDecodeStream. When set, replaces fixed rows_per_batch. When unset, existing row-based behavior unchanged.
  • Modify BatchDecodeStream::next_batch_task() for byte-based row selection — query estimate_decoded_bytes across children, compute row count that fits the budget, use cross-column coordination loop (estimate → check → adjust → converge).
  • Add post-decode feedback loop — after into_batch() measures actual batch size (decoder.rs:2551), feed measured bytes-per-row back to refine row count for subsequent batches.
  • Add estimate_decoded_bytes to StructuralPageDecoder trait — default implementation returns conservative fallback so the system works before all encodings are covered.
  • Add estimate_decoded_bytes to StructuralFieldDecoder trait — field-level decoder spans pages, knows data type, delegates to page decoders. Struct decoder aggregates across children.
  • Implement estimates for exact fixed-width encodings — Flat, Constant, InlineBitpacking, OutOfLineBitpacking, RLE, ByteStreamSplit, PackedStruct. All num_rows * known_width.
  • Implement estimate for General compression (LZ4/Zstd) — read length prefix from compressed frame header (8-byte u64 for Zstd, 4-byte u32 for LZ4) to get exact decompressed size without decompressing.
  • Implement estimate for Dictionary encoding — inspect loaded dictionary DataBlock. Fixed-width: use value width. Variable-width: scan offsets for max value size, bound = num_rows * max_value_size.
  • Implement estimate for Variable encoding (no wrapper) — read offsets from loaded chunk data for exact size of N rows.
  • Implement estimate for FSST — apply 8x algorithmic bound on compressed data size. If wrapped in General, read General length prefix first then apply 8x.
  • Implement estimation for list types with miniblock rep index — binary search rep index to map rows to chunks, estimate at chunk granularity, overestimate boundary chunks.
  • Testing — estimation accuracy per encoding, end-to-end byte-sized batch tests, edge cases (empty/single-row pages, extreme variance), backward compatibility.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Read rust/lance-encoding/AGENTS.md and inspect BatchDecodeStream::next_batch_task(), the decoder.rs:2551 feedback point, and the StructuralPageDecoder and StructuralFieldDecoder traits. Implement the remaining encoding estimates and tests listed in the issue, then verify byte-sized batches, edge cases, and unchanged row-based behavior.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust
Domain
backend, data
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.