{
  "id": 155422,
  "title": "Has anyone tried mixed-precision training?",
  "url": "/competitions/alaska2-image-steganalysis/discussion/155422",
  "author_name": "",
  "post_date": "2020-06-01T16:09:49.730377800Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>When training eb7 with mixed precision i have encountered the problem, when the dynamic scaler goes to zero scale, which leads to division by zero. Any suggestions what could go wrong? The optimizer i am using is torch.AdamW</p>\n\n<pre><code>net, optimizer = amp.initialize(net, optimizer, opt_level=\"O1\", verbosity=True)\n\ntraining_start_time = time.time()\n\nfor epoch in range(n_epochs): \n    net.train()\n    for inputs, labels in tqdm(train_loader):\n        inputs = inputs[\"image\"].to(device, dtype=torch.float)\n        labels = Variable(labels.view(-1)).to(device)\n\n        optimizer.zero_grad()\n\n        outputs = net(inputs).cuda()\n        loss_size = loss(outputs, labels)\n\n        with amp.scale_loss(loss_size, optimizer) as scaled_loss:\n            scaled_loss.backward()\n        optimizer.step()\n        del loss_size, inputs, labels, outputs\n</code></pre>",
  "messages": [
    {
      "id": "870309",
      "postDate": "06/01/2020 16:09:49",
      "content": "<p>When training eb7 with mixed precision i have encountered the problem, when the dynamic scaler goes to zero scale, which leads to division by zero. Any suggestions what could go wrong? The optimizer i am using is torch.AdamW</p>\n\n<pre><code>net, optimizer = amp.initialize(net, optimizer, opt_level=\"O1\", verbosity=True)\n\ntraining_start_time = time.time()\n\nfor epoch in range(n_epochs): \n    net.train()\n    for inputs, labels in tqdm(train_loader):\n        inputs = inputs[\"image\"].to(device, dtype=torch.float)\n        labels = Variable(labels.view(-1)).to(device)\n\n        optimizer.zero_grad()\n\n        outputs = net(inputs).cuda()\n        loss_size = loss(outputs, labels)\n\n        with amp.scale_loss(loss_size, optimizer) as scaled_loss:\n            scaled_loss.backward()\n        optimizer.step()\n        del loss_size, inputs, labels, outputs\n</code></pre>",
      "rawMarkdown": "When training eb7 with mixed precision i have encountered the problem, when the dynamic scaler goes to zero scale, which leads to division by zero. Any suggestions what could go wrong? The optimizer i am using is torch.AdamW\n\n    net, optimizer = amp.initialize(net, optimizer, opt_level=\"O1\", verbosity=True)\n    \n    training_start_time = time.time()\n    \n    for epoch in range(n_epochs): \n        net.train()\n        for inputs, labels in tqdm(train_loader):\n            inputs = inputs[\"image\"].to(device, dtype=torch.float)\n            labels = Variable(labels.view(-1)).to(device)\n            \n            optimizer.zero_grad()\n            \n            outputs = net(inputs).cuda()\n            loss_size = loss(outputs, labels)\n            \n            with amp.scale_loss(loss_size, optimizer) as scaled_loss:\n                scaled_loss.backward()\n            optimizer.step()\n            del loss_size, inputs, labels, outputs",
      "votes": null
    },
    {
      "id": "870852",
      "postDate": "06/02/2020 01:45:05",
      "content": "<p>Does the model train without fp16? The first thing to check would be this. Possible that you're getting nans in your loss/gradient. If your network weights are nan after the error of course that is what happened.</p>",
      "rawMarkdown": "Does the model train without fp16? The first thing to check would be this. Possible that you're getting nans in your loss/gradient. If your network weights are nan after the error of course that is what happened.",
      "votes": null
    },
    {
      "id": "871282",
      "postDate": "06/02/2020 08:42:40",
      "content": "<p>Yes, it works perfectly well if i turn off mixed-precision training. With it turned on i have on of this options:\n-Model does not converge.\n-Zero division error.\nI found no answer searching the internet:(</p>",
      "rawMarkdown": "Yes, it works perfectly well if i turn off mixed-precision training. With it turned on i have on of this options:\n-Model does not converge.\n-Zero division error.\nI found no answer searching the internet:(",
      "votes": null
    },
    {
      "id": "901561",
      "postDate": "06/25/2020 15:09:25",
      "content": "<p>Hey! I was facing similar issues and using torch.cuda.amp worked well for me.\nIn the latest build of PyTorch( dev version i.e 1.6+) , there's an option called amp is built into torch core.\nscaler = torch.cuda.amp.GradScaler()\nwith torch.cuda.amp.autocast():\n        outputs = net(inputs).cuda()\n        loss_size = loss(outputs, labels)</p>\n\n<p>scaler.scale(loss_size).backward()\nscaler.step(optimizer)\nscaler.update()</p>",
      "rawMarkdown": "Hey! I was facing similar issues and using torch.cuda.amp worked well for me.\nIn the latest build of PyTorch( dev version i.e 1.6+) , there's an option called amp is built into torch core.\nscaler = torch.cuda.amp.GradScaler()\nwith torch.cuda.amp.autocast():\n        outputs = net(inputs).cuda()\n        loss_size = loss(outputs, labels)\n\nscaler.scale(loss_size).backward()\nscaler.step(optimizer)\nscaler.update()",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 870852,
      "author_name": "nickouellet",
      "author_url": "",
      "post_date": "06/02/2020 01:45:05",
      "content": "<p>Does the model train without fp16? The first thing to check would be this. Possible that you're getting nans in your loss/gradient. If your network weights are nan after the error of course that is what happened.</p>",
      "votes": null,
      "replies": [
        {
          "id": 871282,
          "author_name": "vovanf98",
          "author_url": "",
          "post_date": "06/02/2020 08:42:40",
          "content": "<p>Yes, it works perfectly well if i turn off mixed-precision training. With it turned on i have on of this options:\n-Model does not converge.\n-Zero division error.\nI found no answer searching the internet:(</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 901561,
      "author_name": "sumukhaithal",
      "author_url": "",
      "post_date": "06/25/2020 15:09:25",
      "content": "<p>Hey! I was facing similar issues and using torch.cuda.amp worked well for me.\nIn the latest build of PyTorch( dev version i.e 1.6+) , there's an option called amp is built into torch core.\nscaler = torch.cuda.amp.GradScaler()\nwith torch.cuda.amp.autocast():\n        outputs = net(inputs).cuda()\n        loss_size = loss(outputs, labels)</p>\n\n<p>scaler.scale(loss_size).backward()\nscaler.step(optimizer)\nscaler.update()</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "870309": "When training eb7 with mixed precision i have encountered the problem, when the dynamic scaler goes to zero scale, which leads to division by zero. Any suggestions what could go wrong? The optimizer i am using is torch.AdamW\n\n    net, optimizer = amp.initialize(net, optimizer, opt_level=\"O1\", verbosity=True)\n    \n    training_start_time = time.time()\n    \n    for epoch in range(n_epochs): \n        net.train()\n        for inputs, labels in tqdm(train_loader):\n            inputs = inputs[\"image\"].to(device, dtype=torch.float)\n            labels = Variable(labels.view(-1)).to(device)\n            \n            optimizer.zero_grad()\n            \n            outputs = net(inputs).cuda()\n            loss_size = loss(outputs, labels)\n            \n            with amp.scale_loss(loss_size, optimizer) as scaled_loss:\n                scaled_loss.backward()\n            optimizer.step()\n            del loss_size, inputs, labels, outputs",
    "870852": "Does the model train without fp16? The first thing to check would be this. Possible that you're getting nans in your loss/gradient. If your network weights are nan after the error of course that is what happened.",
    "871282": "Yes, it works perfectly well if i turn off mixed-precision training. With it turned on i have on of this options:\n-Model does not converge.\n-Zero division error.\nI found no answer searching the internet:(",
    "901561": "Hey! I was facing similar issues and using torch.cuda.amp worked well for me.\nIn the latest build of PyTorch( dev version i.e 1.6+) , there's an option called amp is built into torch core.\nscaler = torch.cuda.amp.GradScaler()\nwith torch.cuda.amp.autocast():\n        outputs = net(inputs).cuda()\n        loss_size = loss(outputs, labels)\n\nscaler.scale(loss_size).backward()\nscaler.step(optimizer)\nscaler.update()"
  },
  "source": "meta"
}