TimelyDataflow / TimelyDataflow/differential-dataflow

Better Scaling for Parallel Workers

Open
#273 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Rust
Stars
3k
Forks
211
Avg merge
10h 42m
Merged PRs (30d)
34

Description

Hello,
my name is Omar and this summer I am doing a internship at VMWare with @ryzhyk and @mbudiu-vmw.

My project is pretty general right now: get better parallel worker scaling for differential datalog programs (I don't have concrete numbers yet on how ddlog programs scale).

I have spend the last couple weeks poking around the differential-dataflow and timely-dataflow repositories seeing how things fit together and reading relevant blogs: https://github.com/frankmcsherry/blog. There is a lot of information so I can't say I have absorbed it all.

I was hoping to get some advice or thoughts on this project. It seems that batch size and timestamp granularity have subtle interplay with latency and throughput. As I see it, there is two general ways to approach the project:

  1. Top-down starting at ddlog: Understand what kind of workflows we're interested in achieving better scaling out of and profile them. Then tune the number of workers, timestamp granularity, operators, etc, to best exploit differentail-dataflow's parallelism. Perhaps there will be common cases among the ddlog programs that I can optimize the differential or timely implementation for.

  2. Working at lower layers: Work at the timely or differential level and try to achieve better parallelism by profiling timely/differential programs respectively exhibiting poor scaling. I don't understand a lot of the timely internals so it is currently unclear to me how feasible this approach is. And even if it is, it may be the case that the workflows we're interested in may not see much speed up from these changes. Any thoughts on possible bottlenecks or places where we could hope optimize the execution for better parallel scaling?

Thank you for any feedback or thoughts.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by profiling representative differential Datalog programs and reading the timely-dataflow and differential-dataflow internals, along with the linked Frank McSherry blog posts. Compare the top-down DDlog approach with lower-layer profiling, focusing on worker count, batch size, timestamp granularity, latency, and throughput. The issue does not define concrete workloads, bottlenecks, or completion criteria.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust
Domain
distributed-systems, performance
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
15/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.