google / google/dpsynth

Pipeline SWIFT query selection appears to use exact marginal errors

Open
#35 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
29
Forks
10
Avg merge
2d 2h
Merged PRs (30d)
36

Description

## Problem

The scalable pipeline SWIFT path appears to use exact-error-driven query selection.

In `dpsynth/pipeline_transformations/swift.py`, the pipeline computes exact candidate marginals, converts them into errors with `marginals_computations.compute_errors(...)`, requests a budget named `Swift Select Queries`, and then passes the errors into `swift.select_queries(...)`.

However, the selection scores do not appear to be noised before `swift.select_queries(...)`; noise is added only later to the selected marginal measurements.

## Why this matters

The selected clique tree / selected workload is itself data-dependent output. Concretely, the junction-tree topology, the selected clique set, and (when diagnostics are enabled) the exact error scores are all released and all depend on exact high-order marginals. If selection is driven by exact marginal errors, the later noisy measurement step does not protect the information leaked by which queries were selected.

This is separate from the local `discrete_mechanisms.swift` path, which has its own score-noising logic (`_compute_initial_errors` adds noise funded by a dedicated selection budget). The issue here is the scalable pipeline transformation path, which has no equivalent noising step.

## Local evidence

Reviewed at commit `18c2c951bd2923f889f6e3b2b757e01aaae398ee`; re-verified still present at current `main` (`91e9181`) — the pipeline path still feeds unnoised errors from `compute_errors` into `swift.select_queries`.

Relevant lines in the current tree:

- `dpsynth/pipeline_transformations/swift.py`: `exact_marginals = marginals_computations.compute_exact_marginals(...)`
- `dpsynth/pipeline_transformations/swift.py`: `errors = marginals_computations.compute_errors(...)`
- `dpsynth/pipeline_transformations/swift.py`: budget request named `Swift Select Queries`
- `dpsynth/pipeline_transformations/swift.py`: `return swift.select_queries(errors_dict, ...)`
- `dpsynth/pipeline_transformations/swift.py`: noise is added at the later `Add noise to selected marginals` stage
- `dpsynth/pipeline_transformations/marginals_computations.py`: `compute_errors(...)` uses `exact_vals` from exact marginals

## Possible fix

Account separately for selection and measurement. Add DP noise to the vector of SWIFT candidate error scores before clique-tree/query selection, and use the remaining measurement budget only for selected marginal measurement. Diagnostic output should avoid publishing exact errors.

## Draft PR

I opened a draft fix here: https://github.com/google/dpsynth/pull/31

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.