{
  "id": 337278,
  "title": "Optimal LR depends on dataset size / augmentations",
  "url": "/competitions/hubmap-organ-segmentation/discussion/337278",
  "author_name": "",
  "post_date": "2022-07-15T09:27:27.232296200Z",
  "votes": 6,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Made an interesting observation. When adding a large amount of new data to the training set, or adding strong augmentations like RandomCrop, it looks like an increase in Learning Rate on the training curves, that is, the error starts to jump strongly from epoch to epoch. And I have seen this several times.</p>\n<p>Apparently, with an increase in the number of data, the network has more diverse gradients that knock the network out of a stable state (otherwise, the network quickly converges to some kind of stable state in which there are few gradients. And since there are few gradients in it, then the network in this and remains in this state). In particular, a gradient appears on those neurons that previously did not have it at all.</p>\n<p>Accordingly, when adding data / augmentations, Learning Rate must be decreased</p>",
  "messages": [
    {
      "id": "1856339",
      "postDate": "07/15/2022 09:27:27",
      "content": "<p>Made an interesting observation. When adding a large amount of new data to the training set, or adding strong augmentations like RandomCrop, it looks like an increase in Learning Rate on the training curves, that is, the error starts to jump strongly from epoch to epoch. And I have seen this several times.</p>\n<p>Apparently, with an increase in the number of data, the network has more diverse gradients that knock the network out of a stable state (otherwise, the network quickly converges to some kind of stable state in which there are few gradients. And since there are few gradients in it, then the network in this and remains in this state). In particular, a gradient appears on those neurons that previously did not have it at all.</p>\n<p>Accordingly, when adding data / augmentations, Learning Rate must be decreased</p>",
      "rawMarkdown": "Made an interesting observation. When adding a large amount of new data to the training set, or adding strong augmentations like RandomCrop, it looks like an increase in Learning Rate on the training curves, that is, the error starts to jump strongly from epoch to epoch. And I have seen this several times.\n\nApparently, with an increase in the number of data, the network has more diverse gradients that knock the network out of a stable state (otherwise, the network quickly converges to some kind of stable state in which there are few gradients. And since there are few gradients in it, then the network in this and remains in this state). In particular, a gradient appears on those neurons that previously did not have it at all.\n\nAccordingly, when adding data / augmentations, Learning Rate must be decreased",
      "votes": null
    },
    {
      "id": "1856426",
      "postDate": "07/15/2022 10:37:11",
      "content": "<p><a href=\"https://www.kaggle.com/zavodrobotov\" target=\"_blank\">@zavodrobotov</a>  interesting observation … I am curious to know a bit further on this .. is it the same way on  Cutout, mixup, CutMix, and AugMix augmentations as well ? Have you got a chance to experiment on these as well ?</p>",
      "rawMarkdown": "zavodrobotov  interesting observation ... I am curious to know a bit further on this .. is it the same way on  Cutout, mixup, CutMix, and AugMix augmentations as well ? Have you got a chance to experiment on these as well ?",
      "votes": null
    },
    {
      "id": "1856502",
      "postDate": "07/15/2022 11:21:25",
      "content": "<p>I try CutOut as one of these augmentations</p>",
      "rawMarkdown": "I try CutOut as one of these augmentations",
      "votes": null
    },
    {
      "id": "1863547",
      "postDate": "07/20/2022 11:31:46",
      "content": "<p>This is a great observation. I have seen this happen as well. With more data, the model seems to need a slower learning rate in order to converge properly. I think this is because the model is seeing more varied data and needs to be trained more slowly in order to avoid overfitting. Thanks for sharing!</p>",
      "rawMarkdown": "This is a great observation. I have seen this happen as well. With more data, the model seems to need a slower learning rate in order to converge properly. I think this is because the model is seeing more varied data and needs to be trained more slowly in order to avoid overfitting. Thanks for sharing!",
      "votes": null
    },
    {
      "id": "2599941",
      "postDate": "01/13/2024 08:31:30",
      "content": "<p><a href=\"https://www.kaggle.com/arunpurakkatt\" target=\"_blank\">@arunpurakkatt</a> Yes, I did experiment with MixUp and other semi-correct augmentations, and it worked. This is because image is at least 100*100, and CNN have millions or billion parameters. So if we taking into account only corners of feature space, we have space of size 2^10000 or 2^1000000. And most of this space will be empty (with no data). And model predict in these regions will be random. But if we fill this empty space with semi-correct data, we fill some portions of this empty space, and model predict will be less random (and it is better that make predict fully random at these regions).</p>",
      "rawMarkdown": "arunpurakkatt Yes, I did experiment with MixUp and other semi-correct augmentations, and it worked. This is because image is at least 100*100, and CNN have millions or billion parameters. So if we taking into account only corners of feature space, we have space of size 2^10000 or 2^1000000. And most of this space will be empty (with no data). And model predict in these regions will be random. But if we fill this empty space with semi-correct data, we fill some portions of this empty space, and model predict will be less random (and it is better that make predict fully random at these regions).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1856426,
      "author_name": "arunpurakkatt",
      "author_url": "",
      "post_date": "07/15/2022 10:37:11",
      "content": "<p><a href=\"https://www.kaggle.com/zavodrobotov\" target=\"_blank\">@zavodrobotov</a>  interesting observation … I am curious to know a bit further on this .. is it the same way on  Cutout, mixup, CutMix, and AugMix augmentations as well ? Have you got a chance to experiment on these as well ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1856502,
          "author_name": "zavodrobotov",
          "author_url": "",
          "post_date": "07/15/2022 11:21:25",
          "content": "<p>I try CutOut as one of these augmentations</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2599941,
          "author_name": "zavodrobotov",
          "author_url": "",
          "post_date": "01/13/2024 08:31:30",
          "content": "<p><a href=\"https://www.kaggle.com/arunpurakkatt\" target=\"_blank\">@arunpurakkatt</a> Yes, I did experiment with MixUp and other semi-correct augmentations, and it worked. This is because image is at least 100*100, and CNN have millions or billion parameters. So if we taking into account only corners of feature space, we have space of size 2^10000 or 2^1000000. And most of this space will be empty (with no data). And model predict in these regions will be random. But if we fill this empty space with semi-correct data, we fill some portions of this empty space, and model predict will be less random (and it is better that make predict fully random at these regions).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1863547,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "07/20/2022 11:31:46",
      "content": "<p>This is a great observation. I have seen this happen as well. With more data, the model seems to need a slower learning rate in order to converge properly. I think this is because the model is seeing more varied data and needs to be trained more slowly in order to avoid overfitting. Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1856339": "Made an interesting observation. When adding a large amount of new data to the training set, or adding strong augmentations like RandomCrop, it looks like an increase in Learning Rate on the training curves, that is, the error starts to jump strongly from epoch to epoch. And I have seen this several times.\n\nApparently, with an increase in the number of data, the network has more diverse gradients that knock the network out of a stable state (otherwise, the network quickly converges to some kind of stable state in which there are few gradients. And since there are few gradients in it, then the network in this and remains in this state). In particular, a gradient appears on those neurons that previously did not have it at all.\n\nAccordingly, when adding data / augmentations, Learning Rate must be decreased",
    "1856426": "zavodrobotov  interesting observation ... I am curious to know a bit further on this .. is it the same way on  Cutout, mixup, CutMix, and AugMix augmentations as well ? Have you got a chance to experiment on these as well ?",
    "1856502": "I try CutOut as one of these augmentations",
    "1863547": "This is a great observation. I have seen this happen as well. With more data, the model seems to need a slower learning rate in order to converge properly. I think this is because the model is seeing more varied data and needs to be trained more slowly in order to avoid overfitting. Thanks for sharing!",
    "2599941": "arunpurakkatt Yes, I did experiment with MixUp and other semi-correct augmentations, and it worked. This is because image is at least 100*100, and CNN have millions or billion parameters. So if we taking into account only corners of feature space, we have space of size 2^10000 or 2^1000000. And most of this space will be empty (with no data). And model predict in these regions will be random. But if we fill this empty space with semi-correct data, we fill some portions of this empty space, and model predict will be less random (and it is better that make predict fully random at these regions)."
  },
  "source": "meta"
}