NVIDIA / NVIDIA/apex

DDP Failed when using the parameters directly to calculate the loss.

Open
#436 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
9k
Forks
1.5k
Avg merge
2d 4h
Merged PRs (30d)
3

Description

# the 'WORLD_SIZE' environment variable will also be set automatically.
 args.distributed = False
 if 'WORLD_SIZE' in os.environ:
     args.distributed = int(os.environ['WORLD_SIZE']) > ### ### **1**
 
 if args.distributed:
     # FOR DISTRIBUTED:  Set the device according to local_rank.
     torch.cuda.set_device(args.local_rank)
  
     # FOR DISTRIBUTED:  Initialize the backend.  torch.distributed.launch will provide
     # environment variables, and requires that you use init_method=`env://`.
     torch.distributed.init_process_group(backend='nccl',
                                          init_method='env://')
 
 torch.backends.cudnn.benchmark = True
 
 N, D_in, D_out = 64, 1024, 1
 
 # Each process receives its own batch of "fake input data" and "fake target data."
 # The "training loop" in each process just uses this fake batch over and over.
 # https://github.com/NVIDIA/apex/tree/master/examples/imagenet provides a more realistic
 # example of distributed data sampling for both training and validation.
 x = torch.randn(D_in, device='cuda')
 y = torch.randn(D_out, device='cuda')
 
 model = torch.nn.Linear(D_in, D_out).cuda()
 optimizer = torch.optim.SGD(model.parameters(), lr=1e-3)
 
 model, optimizer = amp.initialize(model, optimizer, opt_level="O1")
 
 if args.distributed:
     # FOR DISTRIBUTED:  After amp.initialize, wrap the model with
     # apex.parallel.DistributedDataParallel.
     model = DistributedDataParallel(model,delay_allreduce =False)
     # torch.nn.parallel.DistributedDataParallel is also fine, with some added args:
     # model = torch.nn.parallel.DistributedDataParallel(model,
     #                                                   device_ids=[args.local_rank],
     #                                                   output_device=args.local_rank)
 
     # print(model.callback_queued)
 loss_fn = torch.nn.MSELoss()
 print(y)
 
 for t in range(500):
     optimizer.zero_grad()
     #y_pred = model(x)
     loss = loss_fn(model.module.weight, x.view(1,-1))
     print(loss)
     with amp.scale_loss(loss, optimizer) as scaled_loss:
         scaled_loss.backward(retain_graph=False)
     optimizer.step()
 
 if args.local_rank == 0:
     print("final loss = ", loss)

This code leads to an attributeError.

AttributeError: 'DistributedDataParallel' object has no attribute 'needs_refresh'

Main reason is that no explicit feedforward in this code, so there is no chance to collect the parameters.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by running the provided distributed example and trace the interaction between amp.initialize, DistributedDataParallel, and amp.scale_loss when the loss uses model.module.weight without a forward pass. Done means the reproducer no longer raises the needs_refresh AttributeError while direct parameter-based loss calculation remains supported.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
distributed-systems, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.