{
  "id": 44789,
  "title": "Training with gradient accumulation doesn't work the same ",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/44789",
  "author_name": "",
  "post_date": "2017-12-02T13:48:12.913046Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Usually I do my trainings on a machine with two GTX 1060 6GB RAM, however from time to time I would spin up a beefier machine with 3 1080ti, to speed up the training and see results.\nHowever I noticed that when training with 3 1080ti (e.g: batch_size~430) I was able to get to better results (accuracy) than when training on 2x1060 (batch_size~110).\nI did use gradients accumulation or alternatively lower LR, but still the training on 3x1080ti resulted in higher accuracy on the same script rather than the 2x1060.</p>\n\n<p>And so my question - does my training loop for gradient accumulation has a bug?\nAnybody else generally experiencing these types of issues?</p>\n\n<p><a href=\"https://gist.github.com/burgalon/31360161193cd8065c51db5cc1e9b5f6\">https://gist.github.com/burgalon/31360161193cd8065c51db5cc1e9b5f6</a></p>",
  "messages": [
    {
      "id": "252211",
      "postDate": "12/02/2017 13:48:12",
      "content": "<p>Usually I do my trainings on a machine with two GTX 1060 6GB RAM, however from time to time I would spin up a beefier machine with 3 1080ti, to speed up the training and see results.\nHowever I noticed that when training with 3 1080ti (e.g: batch_size~430) I was able to get to better results (accuracy) than when training on 2x1060 (batch_size~110).\nI did use gradients accumulation or alternatively lower LR, but still the training on 3x1080ti resulted in higher accuracy on the same script rather than the 2x1060.</p>\n\n<p>And so my question - does my training loop for gradient accumulation has a bug?\nAnybody else generally experiencing these types of issues?</p>\n\n<p><a href=\"https://gist.github.com/burgalon/31360161193cd8065c51db5cc1e9b5f6\">https://gist.github.com/burgalon/31360161193cd8065c51db5cc1e9b5f6</a></p>",
      "rawMarkdown": "Usually I do my trainings on a machine with two GTX 1060 6GB RAM, however from time to time I would spin up a beefier machine with 3 1080ti, to speed up the training and see results.\nHowever I noticed that when training with 3 1080ti (e.g: batch_size~430) I was able to get to better results (accuracy) than when training on 2x1060 (batch_size~110).\nI did use gradients accumulation or alternatively lower LR, but still the training on 3x1080ti resulted in higher accuracy on the same script rather than the 2x1060.\n\nAnd so my question - does my training loop for gradient accumulation has a bug?\nAnybody else generally experiencing these types of issues?\n\nhttps://gist.github.com/burgalon/31360161193cd8065c51db5cc1e9b5f6",
      "votes": null
    },
    {
      "id": "252212",
      "postDate": "12/02/2017 13:50:18",
      "content": "<p>Separately I'm exieriencing this weird problem <a href=\"https://datascience.stackexchange.com/questions/25207/resuming-from-checkpoint-accuracy-drops-for-one-cycle\">https://datascience.stackexchange.com/questions/25207/resuming-from-checkpoint-accuracy-drops-for-one-cycle</a></p>",
      "rawMarkdown": "Separately I'm exieriencing this weird problem https://datascience.stackexchange.com/questions/25207/resuming-from-checkpoint-accuracy-drops-for-one-cycle",
      "votes": null
    },
    {
      "id": "252534",
      "postDate": "12/03/2017 07:03:27",
      "content": "<p>Gradient accumulation can't be a complete substitute for increasing batchsize.\nBecause Batch Normalization layer in you network feed the original minibatch every iteration, so BN behaves differently in your two cases. </p>\n\n<p>By the way, would you tell me how much different between your two results ?</p>",
      "rawMarkdown": "Gradient accumulation can't be a complete substitute for increasing batchsize.\nBecause Batch Normalization layer in you network feed the original minibatch every iteration, so BN behaves differently in your two cases. \n\nBy the way, would you tell me how much different between your two results ?",
      "votes": null
    },
    {
      "id": "252536",
      "postDate": "12/03/2017 07:11:28",
      "content": "<p>0.69 vs 0.727</p>",
      "rawMarkdown": "0.69 vs 0.727",
      "votes": null
    },
    {
      "id": "253532",
      "postDate": "12/05/2017 05:55:24",
      "content": "<p>Same problem for me using tensorflow, have you figure out what reason might be? By the way,I use gradient accumulation too.</p>",
      "rawMarkdown": "Same problem for me using tensorflow, have you figure out what reason might be? By the way,I use gradient accumulation too.",
      "votes": null
    },
    {
      "id": "253547",
      "postDate": "12/05/2017 07:09:05",
      "content": "<p>You'r right, gradient of BN is dependent on each mini-batch's mean and variance, so the sum of the two gradient not equals the gradient of two sum! </p>\n\n<p>But is this related with the problem that accuracy droped everytime restored from a checkpoint?</p>",
      "rawMarkdown": "You'r right, gradient of BN is dependent on each mini-batch's mean and variance, so the sum of the two gradient not equals the gradient of two sum! \n\nBut is this related with the problem that accuracy droped everytime restored from a checkpoint?",
      "votes": null
    },
    {
      "id": "254055",
      "postDate": "12/06/2017 05:32:59",
      "content": "<p>Sorry, I did not look the problem carefully. Probably this is not related about that.</p>",
      "rawMarkdown": "Sorry, I did not look the problem carefully. Probably this is not related about that.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 252212,
      "author_name": "burgalon",
      "author_url": "",
      "post_date": "12/02/2017 13:50:18",
      "content": "<p>Separately I'm exieriencing this weird problem <a href=\"https://datascience.stackexchange.com/questions/25207/resuming-from-checkpoint-accuracy-drops-for-one-cycle\">https://datascience.stackexchange.com/questions/25207/resuming-from-checkpoint-accuracy-drops-for-one-cycle</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 253532,
          "author_name": "plainsailing",
          "author_url": "",
          "post_date": "12/05/2017 05:55:24",
          "content": "<p>Same problem for me using tensorflow, have you figure out what reason might be? By the way,I use gradient accumulation too.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 252534,
      "author_name": "lyakaap",
      "author_url": "",
      "post_date": "12/03/2017 07:03:27",
      "content": "<p>Gradient accumulation can't be a complete substitute for increasing batchsize.\nBecause Batch Normalization layer in you network feed the original minibatch every iteration, so BN behaves differently in your two cases. </p>\n\n<p>By the way, would you tell me how much different between your two results ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 252536,
          "author_name": "burgalon",
          "author_url": "",
          "post_date": "12/03/2017 07:11:28",
          "content": "<p>0.69 vs 0.727</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 253547,
          "author_name": "plainsailing",
          "author_url": "",
          "post_date": "12/05/2017 07:09:05",
          "content": "<p>You'r right, gradient of BN is dependent on each mini-batch's mean and variance, so the sum of the two gradient not equals the gradient of two sum! </p>\n\n<p>But is this related with the problem that accuracy droped everytime restored from a checkpoint?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 254055,
          "author_name": "lyakaap",
          "author_url": "",
          "post_date": "12/06/2017 05:32:59",
          "content": "<p>Sorry, I did not look the problem carefully. Probably this is not related about that.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "252211": "Usually I do my trainings on a machine with two GTX 1060 6GB RAM, however from time to time I would spin up a beefier machine with 3 1080ti, to speed up the training and see results.\nHowever I noticed that when training with 3 1080ti (e.g: batch_size~430) I was able to get to better results (accuracy) than when training on 2x1060 (batch_size~110).\nI did use gradients accumulation or alternatively lower LR, but still the training on 3x1080ti resulted in higher accuracy on the same script rather than the 2x1060.\n\nAnd so my question - does my training loop for gradient accumulation has a bug?\nAnybody else generally experiencing these types of issues?\n\nhttps://gist.github.com/burgalon/31360161193cd8065c51db5cc1e9b5f6",
    "252212": "Separately I'm exieriencing this weird problem https://datascience.stackexchange.com/questions/25207/resuming-from-checkpoint-accuracy-drops-for-one-cycle",
    "252534": "Gradient accumulation can't be a complete substitute for increasing batchsize.\nBecause Batch Normalization layer in you network feed the original minibatch every iteration, so BN behaves differently in your two cases. \n\nBy the way, would you tell me how much different between your two results ?",
    "252536": "0.69 vs 0.727",
    "253532": "Same problem for me using tensorflow, have you figure out what reason might be? By the way,I use gradient accumulation too.",
    "253547": "You'r right, gradient of BN is dependent on each mini-batch's mean and variance, so the sum of the two gradient not equals the gradient of two sum! \n\nBut is this related with the problem that accuracy droped everytime restored from a checkpoint?",
    "254055": "Sorry, I did not look the problem carefully. Probably this is not related about that."
  },
  "source": "meta"
}