microsoft / microsoft/dp-transformers

Gradients haven't been cleared since the last optimizer step. In order to obtain privacy guarantees you must call optimizer.zero_grad()on each step

Open
#47 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
149
Forks
33
Avg merge
22h 6m
Merged PRs (30d)
1

Description

System Info

  • transformers version == 4.42.4
  • dp-transformers version == 1.0.1
  • accelerate version == 0.29.3
  • llm == Distil-gpt2
  • torch version == 2.0.0
  • The model is runned under ddp with 2 GPUs

Error
in line 2330 in the function self.optimizer.step()
Gradients haven't been cleared since the last optimizer step. In order to obtain privacy guarantees you must call optimizer.zero_grad()on each step

Explanation
Hello, I am trying to train a Tabula (https://github.com/zhao-zilong/Tabula) model using differential privacy. I rewrote my Tabula trainer to use the OpacusDPTrainer, but I encountered an error stating that optimizer.step() requires a call to optimizer.zero_grad() at each step to obtain privacy guarantees. I tried to resolve the problem by reimplementing the _inner_training_loop function from the transformers trainer and forcing self.optimizer.zero_grad() at each step, but I still get the same error.

Code
Function start at line 1939, the modification is added to line 2032
`
def _inner_training_loop(

    self, batch_size=None, args=None, resume_from_checkpoint=None, trial=None, ignore_keys_for_eval=None

):

    self.accelerator.free_memory()

    self._train_batch_size = batch_size

    if self.args.auto_find_batch_size:

    ....

    step = -1

        for step, inputs in enumerate(epoch_iterator):

            total_batched_samples += 1

            # My function 

            self.optimizer.zero_grad()

            if self.args.include_num_input_tokens_seen:

                main_input_name = getattr(self.model, "main_input_name", "input_ids")

                if main_input_name not in inputs:

                    logger.warning(

                        "Tried to track the number of tokens seen, however the current model is "

                        "not configured properly to know what item is the input. To fix this, add "

                        "a `main_input_name` attribute to the model class you are using."

                    )

                else:`

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the reported self.optimizer.step() failure around line 2330 and the modified _inner_training_loop beginning around line 1939, including the zero_grad() call added near line 2032. Reproduce with the listed Transformers, dp-transformers, Accelerate, PyTorch, DistilGPT-2, and two-GPU DDP setup; done means the training step no longer raises the gradient-clearing error while retaining the required privacy behavior.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning, security
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.