{
  "id": 133217,
  "title": "CUTMIX: Train Loss and Validation Loss Not Improving.  ",
  "url": "/competitions/bengaliai-cv19/discussion/133217",
  "author_name": "",
  "post_date": "2020-03-01T12:08:57.177246100Z",
  "votes": 3,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I am using SeResNext50 model, with a batch size of 1024, 64x64 train imgs with Cutmix (on only 40% of train img) alone as the augmentation method. My train acc got to around 60% within 4-5 epochs but after that my training accuracy is stuck at 60% and validation accuracy at 70%, even after 35 epochs!</p>\n\n<p>Can anyone suggest what would be the issue or is it normal and the accuracy would improve after more number of epochs?</p>",
  "messages": [
    {
      "id": "760520",
      "postDate": "03/01/2020 12:08:57",
      "content": "<p>I am using SeResNext50 model, with a batch size of 1024, 64x64 train imgs with Cutmix (on only 40% of train img) alone as the augmentation method. My train acc got to around 60% within 4-5 epochs but after that my training accuracy is stuck at 60% and validation accuracy at 70%, even after 35 epochs!</p>\n\n<p>Can anyone suggest what would be the issue or is it normal and the accuracy would improve after more number of epochs?</p>",
      "rawMarkdown": "I am using SeResNext50 model, with a batch size of 1024, 64x64 train imgs with Cutmix (on only 40% of train img) alone as the augmentation method. My train acc got to around 60% within 4-5 epochs but after that my training accuracy is stuck at 60% and validation accuracy at 70%, even after 35 epochs!\n\nCan anyone suggest what would be the issue or is it normal and the accuracy would improve after more number of epochs?",
      "votes": null
    },
    {
      "id": "760525",
      "postDate": "03/01/2020 12:16:50",
      "content": "<p>Does it help if you reduce your batch_size to say 128 (or even lower let's say 64)?\nIs the 40% of train image random? does it change at every epoch (it should)?\nWhen doing cutmix, it's hard to compute a clear comparable metric on the train set since you are generating some really weird examples (and the real label is unclear), so I would not worry too much about your training scores, but your validation score should not be stuck to such a low score, maybe you should also try to reduce your learning rate.\nHope it helps!</p>",
      "rawMarkdown": "Does it help if you reduce your batch_size to say 128 (or even lower let's say 64)?\nIs the 40% of train image random? does it change at every epoch (it should)?\nWhen doing cutmix, it's hard to compute a clear comparable metric on the train set since you are generating some really weird examples (and the real label is unclear), so I would not worry too much about your training scores, but your validation score should not be stuck to such a low score, maybe you should also try to reduce your learning rate.\nHope it helps!",
      "votes": null
    },
    {
      "id": "760568",
      "postDate": "03/01/2020 13:29:51",
      "content": "<p>Thanks for your advise <a href=\"/optimo\">@optimo</a> .\nYes, the train images are shuffled at every epoch.\nNow, I have changed the batch size from 1024 to 128, hoping for improvement!</p>\n\n<p>Can you please explain why bigger batch size of 1024 is hampering the accuracy? I do understand the fact that with that large batch size, every train image in a particular batch would have the same cutmix ratio, making difficult to train, but we still have 60% of the remaining train data to train, which are not cutmixed.</p>\n\n<p>Also, i am using 0.0001 as lr, isn't is small already?</p>\n\n<p>Thanks again :)</p>",
      "rawMarkdown": "Thanks for your advise @optimo .\nYes, the train images are shuffled at every epoch.\nNow, I have changed the batch size from 1024 to 128, hoping for improvement!\n\nCan you please explain why bigger batch size of 1024 is hampering the accuracy? I do understand the fact that with that large batch size, every train image in a particular batch would have the same cutmix ratio, making difficult to train, but we still have 60% of the remaining train data to train, which are not cutmixed.\n\nAlso, i am using 0.0001 as lr, isn't is small already?\n\nThanks again :)",
      "votes": null
    },
    {
      "id": "760590",
      "postDate": "03/01/2020 13:49:53",
      "content": "<p>your LR should be alright.</p>\n\n<p>About the batch size, I guess it's theoretically a bit unclear why a very small BS should work better than a reasonably large one, but since here you have 3 different tasks, with more than a hundred labels maybe it's hard to converge with averaging 1024 gradients (I might be wrong and it might work here, I personally did not try) but some people including Yann Lecun have very strong feelings about this (see here : <a href=\"https://twitter.com/ylecun/status/989610208497360896?lang=en\">https://twitter.com/ylecun/status/989610208497360896?lang=en</a>)</p>\n\n<p>I think large mini batches is alright for tabular data (from my experience) but might be dangerous with images. (I think this mostly depends whether you happen to have BatchNorm as one of your layers as well)</p>\n\n<p>Anyway let us know about the results to see if I was right or not! </p>",
      "rawMarkdown": "your LR should be alright.\n\nAbout the batch size, I guess it's theoretically a bit unclear why a very small BS should work better than a reasonably large one, but since here you have 3 different tasks, with more than a hundred labels maybe it's hard to converge with averaging 1024 gradients (I might be wrong and it might work here, I personally did not try) but some people including Yann Lecun have very strong feelings about this (see here : https://twitter.com/ylecun/status/989610208497360896?lang=en)\n\nI think large mini batches is alright for tabular data (from my experience) but might be dangerous with images. (I think this mostly depends whether you happen to have BatchNorm as one of your layers as well)\n\nAnyway let us know about the results to see if I was right or not!",
      "votes": null
    },
    {
      "id": "760673",
      "postDate": "03/01/2020 15:51:51",
      "content": "<p><a href=\"/optimo\">@optimo</a> My val accuracy did improve to 88% after 11 epochs of training :)</p>",
      "rawMarkdown": "optimo My val accuracy did improve to 88% after 11 epochs of training :)",
      "votes": null
    },
    {
      "id": "760683",
      "postDate": "03/01/2020 16:03:06",
      "content": "<p>hehe good to know! glad it helped!</p>",
      "rawMarkdown": "hehe good to know! glad it helped!",
      "votes": null
    },
    {
      "id": "760705",
      "postDate": "03/01/2020 16:23:40",
      "content": "<p>refer to facebook paper:</p>\n\n<p>\"Linear Scaling Rule: When the minibatch size is\nmultiplied by k, multiply the learning rate by k.\"</p>\n\n<p><a href=\"https://research.fb.com/wp-content/uploads/2017/06/imagenet1kin1h5.pdf\">https://research.fb.com/wp-content/uploads/2017/06/imagenet1kin1h5.pdf</a></p>\n\n<p>see also this : <a href=\"https://medium.com/mini-distill/effect-of-batch-size-on-training-dynamics-21c14f7a716e\">https://medium.com/mini-distill/effect-of-batch-size-on-training-dynamics-21c14f7a716e</a></p>\n\n<p>\"Finding: large batch size means the model makes very large gradient updates and very small gradient updates. The size of the update depends heavily on which particular samples are drawn from the dataset. On the other hand using small batch size means the model makes updates that are all about the same size. The size of the update only weakly depends on which particular samples are drawn from the dataset.\"</p>\n\n<p></p>",
      "rawMarkdown": "refer to facebook paper:\n\n\"Linear Scaling Rule: When the minibatch size is\nmultiplied by k, multiply the learning rate by k.\"\n\nhttps://research.fb.com/wp-content/uploads/2017/06/imagenet1kin1h5.pdf\n\n\nsee also this : https://medium.com/mini-distill/effect-of-batch-size-on-training-dynamics-21c14f7a716e\n\n\"Finding: large batch size means the model makes very large gradient updates and very small gradient updates. The size of the update depends heavily on which particular samples are drawn from the dataset. On the other hand using small batch size means the model makes updates that are all about the same size. The size of the update only weakly depends on which particular samples are drawn from the dataset.\"\n\n\n![](https://miro.medium.com/max/707/1*kJKRottm7aJsAKcNUjwHkA.png)",
      "votes": null
    },
    {
      "id": "760722",
      "postDate": "03/01/2020 16:49:30",
      "content": "<p>Thanks <a href=\"/hengck23\">@hengck23</a> did not know about that rule.\n<a href=\"/ankitsajwan\">@ankitsajwan</a> it would be interesting to see what results you get if you go back to BS=1024 and set your LR=1e-4*8, see if you get similar results.</p>\n\n<p><a href=\"/hengck23\">@hengck23</a> I've got a kind of related question, with a multihead network would you say increasing the loss like this <code>loss = N*loss1+loss2+loss3</code> is somehow similar to increasing the learning rate of head1 by N ? Do you have any reference I could read? Thanks!</p>",
      "rawMarkdown": "Thanks @hengck23 did not know about that rule.\n@ankitsajwan it would be interesting to see what results you get if you go back to BS=1024 and set your LR=1e-4*8, see if you get similar results.\n\n@hengck23 I've got a kind of related question, with a multihead network would you say increasing the loss like this `loss = N*loss1+loss2+loss3` is somehow similar to increasing the learning rate of head1 by N ? Do you have any reference I could read? Thanks!",
      "votes": null
    },
    {
      "id": "761029",
      "postDate": "03/02/2020 03:50:12",
      "content": "<p><a href=\"/optimo\">@optimo</a> Yes, using <code>N*loss</code> with <code>LR = 1.0 * a</code> is the same as using <code>1.0 * loss</code> with <code>LR = N * a</code>. This is because the update to the weights is equal to the learning rate multiplied by the gradient of weights with respect to loss. Therefore update increases by <code>N</code> when you change <code>loss</code> to <code>N*loss</code> and update increases by <code>N</code> when you change LR from <code>a</code> to <code>N*a</code>.</p>",
      "rawMarkdown": "optimo Yes, using `N*loss` with `LR = 1.0 * a` is the same as using `1.0 * loss` with `LR = N * a`. This is because the update to the weights is equal to the learning rate multiplied by the gradient of weights with respect to loss. Therefore update increases by `N` when you change `loss` to `N*loss` and update increases by `N` when you change LR from `a` to `N*a`.",
      "votes": null
    },
    {
      "id": "761173",
      "postDate": "03/02/2020 08:00:14",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> yep that's exactly true for a single headed model, but when you've got a common basis with shared layers I guess this does not quite hold in the end (the math is more complicated and the multiplication does not back propagate directly through the shared layers no?).\nI mean if you have different ways of representing images in first layers, some will help loss1 but not loss2 then changing the loss directly will make the network to improve only (or mostly) for loss1, while with a basic loss but different learning rates you're doing something that looks really different to me - pursuing a global optimization but more fiercely on the first head. \nAm I clear why I still doubt that it would actually be the same? </p>",
      "rawMarkdown": "cdeotte yep that's exactly true for a single headed model, but when you've got a common basis with shared layers I guess this does not quite hold in the end (the math is more complicated and the multiplication does not back propagate directly through the shared layers no?).\nI mean if you have different ways of representing images in first layers, some will help loss1 but not loss2 then changing the loss directly will make the network to improve only (or mostly) for loss1, while with a basic loss but different learning rates you're doing something that looks really different to me - pursuing a global optimization but more fiercely on the first head. \nAm I clear why I still doubt that it would actually be the same?",
      "votes": null
    },
    {
      "id": "761183",
      "postDate": "03/02/2020 08:22:38",
      "content": "<p>I think it works for multiple losses too. Because if <code>loss = loss1 + loss2 + loss3</code> then the weight update is <code>weights -= LR * grad(loss1) + LR * grad(loss2) + LR * grad(loss3)</code>. </p>\n\n<p>And then <code>LR * grad(N * loss1) + LR * grad(loss2) + LR * grad(loss3)</code> equals <code>N * LR * grad(loss1) + LR * grad(loss2) + LR * grad(loss3)</code></p>",
      "rawMarkdown": "I think it works for multiple losses too. Because if `loss = loss1 + loss2 + loss3` then the weight update is `weights -= LR * grad(loss1) + LR * grad(loss2) + LR * grad(loss3)`. \n  \nAnd then `LR * grad(N * loss1) + LR * grad(loss2) + LR * grad(loss3)` equals `N * LR * grad(loss1) + LR * grad(loss2) + LR * grad(loss3)`",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 760525,
      "author_name": "optimo",
      "author_url": "",
      "post_date": "03/01/2020 12:16:50",
      "content": "<p>Does it help if you reduce your batch_size to say 128 (or even lower let's say 64)?\nIs the 40% of train image random? does it change at every epoch (it should)?\nWhen doing cutmix, it's hard to compute a clear comparable metric on the train set since you are generating some really weird examples (and the real label is unclear), so I would not worry too much about your training scores, but your validation score should not be stuck to such a low score, maybe you should also try to reduce your learning rate.\nHope it helps!</p>",
      "votes": null,
      "replies": [
        {
          "id": 760568,
          "author_name": "ankitsajwan",
          "author_url": "",
          "post_date": "03/01/2020 13:29:51",
          "content": "<p>Thanks for your advise <a href=\"/optimo\">@optimo</a> .\nYes, the train images are shuffled at every epoch.\nNow, I have changed the batch size from 1024 to 128, hoping for improvement!</p>\n\n<p>Can you please explain why bigger batch size of 1024 is hampering the accuracy? I do understand the fact that with that large batch size, every train image in a particular batch would have the same cutmix ratio, making difficult to train, but we still have 60% of the remaining train data to train, which are not cutmixed.</p>\n\n<p>Also, i am using 0.0001 as lr, isn't is small already?</p>\n\n<p>Thanks again :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 760590,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "03/01/2020 13:49:53",
          "content": "<p>your LR should be alright.</p>\n\n<p>About the batch size, I guess it's theoretically a bit unclear why a very small BS should work better than a reasonably large one, but since here you have 3 different tasks, with more than a hundred labels maybe it's hard to converge with averaging 1024 gradients (I might be wrong and it might work here, I personally did not try) but some people including Yann Lecun have very strong feelings about this (see here : <a href=\"https://twitter.com/ylecun/status/989610208497360896?lang=en\">https://twitter.com/ylecun/status/989610208497360896?lang=en</a>)</p>\n\n<p>I think large mini batches is alright for tabular data (from my experience) but might be dangerous with images. (I think this mostly depends whether you happen to have BatchNorm as one of your layers as well)</p>\n\n<p>Anyway let us know about the results to see if I was right or not! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 760673,
          "author_name": "ankitsajwan",
          "author_url": "",
          "post_date": "03/01/2020 15:51:51",
          "content": "<p><a href=\"/optimo\">@optimo</a> My val accuracy did improve to 88% after 11 epochs of training :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 760683,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "03/01/2020 16:03:06",
          "content": "<p>hehe good to know! glad it helped!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 760705,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/01/2020 16:23:40",
          "content": "<p>refer to facebook paper:</p>\n\n<p>\"Linear Scaling Rule: When the minibatch size is\nmultiplied by k, multiply the learning rate by k.\"</p>\n\n<p><a href=\"https://research.fb.com/wp-content/uploads/2017/06/imagenet1kin1h5.pdf\">https://research.fb.com/wp-content/uploads/2017/06/imagenet1kin1h5.pdf</a></p>\n\n<p>see also this : <a href=\"https://medium.com/mini-distill/effect-of-batch-size-on-training-dynamics-21c14f7a716e\">https://medium.com/mini-distill/effect-of-batch-size-on-training-dynamics-21c14f7a716e</a></p>\n\n<p>\"Finding: large batch size means the model makes very large gradient updates and very small gradient updates. The size of the update depends heavily on which particular samples are drawn from the dataset. On the other hand using small batch size means the model makes updates that are all about the same size. The size of the update only weakly depends on which particular samples are drawn from the dataset.\"</p>\n\n<p></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 760722,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "03/01/2020 16:49:30",
          "content": "<p>Thanks <a href=\"/hengck23\">@hengck23</a> did not know about that rule.\n<a href=\"/ankitsajwan\">@ankitsajwan</a> it would be interesting to see what results you get if you go back to BS=1024 and set your LR=1e-4*8, see if you get similar results.</p>\n\n<p><a href=\"/hengck23\">@hengck23</a> I've got a kind of related question, with a multihead network would you say increasing the loss like this <code>loss = N*loss1+loss2+loss3</code> is somehow similar to increasing the learning rate of head1 by N ? Do you have any reference I could read? Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 761029,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "03/02/2020 03:50:12",
          "content": "<p><a href=\"/optimo\">@optimo</a> Yes, using <code>N*loss</code> with <code>LR = 1.0 * a</code> is the same as using <code>1.0 * loss</code> with <code>LR = N * a</code>. This is because the update to the weights is equal to the learning rate multiplied by the gradient of weights with respect to loss. Therefore update increases by <code>N</code> when you change <code>loss</code> to <code>N*loss</code> and update increases by <code>N</code> when you change LR from <code>a</code> to <code>N*a</code>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 761173,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "03/02/2020 08:00:14",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> yep that's exactly true for a single headed model, but when you've got a common basis with shared layers I guess this does not quite hold in the end (the math is more complicated and the multiplication does not back propagate directly through the shared layers no?).\nI mean if you have different ways of representing images in first layers, some will help loss1 but not loss2 then changing the loss directly will make the network to improve only (or mostly) for loss1, while with a basic loss but different learning rates you're doing something that looks really different to me - pursuing a global optimization but more fiercely on the first head. \nAm I clear why I still doubt that it would actually be the same? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 761183,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "03/02/2020 08:22:38",
          "content": "<p>I think it works for multiple losses too. Because if <code>loss = loss1 + loss2 + loss3</code> then the weight update is <code>weights -= LR * grad(loss1) + LR * grad(loss2) + LR * grad(loss3)</code>. </p>\n\n<p>And then <code>LR * grad(N * loss1) + LR * grad(loss2) + LR * grad(loss3)</code> equals <code>N * LR * grad(loss1) + LR * grad(loss2) + LR * grad(loss3)</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "760520": "I am using SeResNext50 model, with a batch size of 1024, 64x64 train imgs with Cutmix (on only 40% of train img) alone as the augmentation method. My train acc got to around 60% within 4-5 epochs but after that my training accuracy is stuck at 60% and validation accuracy at 70%, even after 35 epochs!\n\nCan anyone suggest what would be the issue or is it normal and the accuracy would improve after more number of epochs?",
    "760525": "Does it help if you reduce your batch_size to say 128 (or even lower let's say 64)?\nIs the 40% of train image random? does it change at every epoch (it should)?\nWhen doing cutmix, it's hard to compute a clear comparable metric on the train set since you are generating some really weird examples (and the real label is unclear), so I would not worry too much about your training scores, but your validation score should not be stuck to such a low score, maybe you should also try to reduce your learning rate.\nHope it helps!",
    "760568": "Thanks for your advise @optimo .\nYes, the train images are shuffled at every epoch.\nNow, I have changed the batch size from 1024 to 128, hoping for improvement!\n\nCan you please explain why bigger batch size of 1024 is hampering the accuracy? I do understand the fact that with that large batch size, every train image in a particular batch would have the same cutmix ratio, making difficult to train, but we still have 60% of the remaining train data to train, which are not cutmixed.\n\nAlso, i am using 0.0001 as lr, isn't is small already?\n\nThanks again :)",
    "760590": "your LR should be alright.\n\nAbout the batch size, I guess it's theoretically a bit unclear why a very small BS should work better than a reasonably large one, but since here you have 3 different tasks, with more than a hundred labels maybe it's hard to converge with averaging 1024 gradients (I might be wrong and it might work here, I personally did not try) but some people including Yann Lecun have very strong feelings about this (see here : https://twitter.com/ylecun/status/989610208497360896?lang=en)\n\nI think large mini batches is alright for tabular data (from my experience) but might be dangerous with images. (I think this mostly depends whether you happen to have BatchNorm as one of your layers as well)\n\nAnyway let us know about the results to see if I was right or not!",
    "760673": "optimo My val accuracy did improve to 88% after 11 epochs of training :)",
    "760683": "hehe good to know! glad it helped!",
    "760705": "refer to facebook paper:\n\n\"Linear Scaling Rule: When the minibatch size is\nmultiplied by k, multiply the learning rate by k.\"\n\nhttps://research.fb.com/wp-content/uploads/2017/06/imagenet1kin1h5.pdf\n\n\nsee also this : https://medium.com/mini-distill/effect-of-batch-size-on-training-dynamics-21c14f7a716e\n\n\"Finding: large batch size means the model makes very large gradient updates and very small gradient updates. The size of the update depends heavily on which particular samples are drawn from the dataset. On the other hand using small batch size means the model makes updates that are all about the same size. The size of the update only weakly depends on which particular samples are drawn from the dataset.\"\n\n\n![](https://miro.medium.com/max/707/1*kJKRottm7aJsAKcNUjwHkA.png)",
    "760722": "Thanks @hengck23 did not know about that rule.\n@ankitsajwan it would be interesting to see what results you get if you go back to BS=1024 and set your LR=1e-4*8, see if you get similar results.\n\n@hengck23 I've got a kind of related question, with a multihead network would you say increasing the loss like this `loss = N*loss1+loss2+loss3` is somehow similar to increasing the learning rate of head1 by N ? Do you have any reference I could read? Thanks!",
    "761029": "optimo Yes, using `N*loss` with `LR = 1.0 * a` is the same as using `1.0 * loss` with `LR = N * a`. This is because the update to the weights is equal to the learning rate multiplied by the gradient of weights with respect to loss. Therefore update increases by `N` when you change `loss` to `N*loss` and update increases by `N` when you change LR from `a` to `N*a`.",
    "761173": "cdeotte yep that's exactly true for a single headed model, but when you've got a common basis with shared layers I guess this does not quite hold in the end (the math is more complicated and the multiplication does not back propagate directly through the shared layers no?).\nI mean if you have different ways of representing images in first layers, some will help loss1 but not loss2 then changing the loss directly will make the network to improve only (or mostly) for loss1, while with a basic loss but different learning rates you're doing something that looks really different to me - pursuing a global optimization but more fiercely on the first head. \nAm I clear why I still doubt that it would actually be the same?",
    "761183": "I think it works for multiple losses too. Because if `loss = loss1 + loss2 + loss3` then the weight update is `weights -= LR * grad(loss1) + LR * grad(loss2) + LR * grad(loss3)`. \n  \nAnd then `LR * grad(N * loss1) + LR * grad(loss2) + LR * grad(loss3)` equals `N * LR * grad(loss1) + LR * grad(loss2) + LR * grad(loss3)`"
  },
  "source": "meta"
}