{
  "id": 167217,
  "title": "I have a question about LRscheduler and Image augmentation",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/167217",
  "author_name": "",
  "post_date": "2020-07-15T16:44:13.675830400Z",
  "votes": 3,
  "comment_count": 12,
  "views": 0,
  "content": "<ol>\n<li>When I found a code for Image classification.\nI'm almost watch this lr code.</li>\n</ol>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2701710%2F3d97f637694af75183694605a893bdf7%2Fchrome_8saxN1jIJW.png?generation=1594831058076840&amp;alt=media\" alt=\"\"></p>\n\n<p>But as far as I know, keras supports the following code, why use the first code and what are the advantages?\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2701710%2Fa5ab1f7c18743046ab37f6708eda7e2a%2Fpycharm64_PTH9thaOCX.png?generation=1594831191354691&amp;alt=media\" alt=\"\"></p>\n\n<ol>\n<li>How to define Image augmentation code?</li>\n</ol>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2701710%2F3e82a8e5d082774f544a706ea59eb060%2Fchrome_40aPjP0xnC.png?generation=1594831383395100&amp;alt=media\" alt=\"\"></p>\n\n<p>I'm also watch upper Image augmentation code.\nSo How to define that code? \nYou've experimented with each function multiple times?</p>",
  "messages": [
    {
      "id": "930694",
      "postDate": "07/15/2020 16:44:13",
      "content": "<ol>\n<li>When I found a code for Image classification.\nI'm almost watch this lr code.</li>\n</ol>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2701710%2F3d97f637694af75183694605a893bdf7%2Fchrome_8saxN1jIJW.png?generation=1594831058076840&amp;alt=media\" alt=\"\"></p>\n\n<p>But as far as I know, keras supports the following code, why use the first code and what are the advantages?\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2701710%2Fa5ab1f7c18743046ab37f6708eda7e2a%2Fpycharm64_PTH9thaOCX.png?generation=1594831191354691&amp;alt=media\" alt=\"\"></p>\n\n<ol>\n<li>How to define Image augmentation code?</li>\n</ol>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2701710%2F3e82a8e5d082774f544a706ea59eb060%2Fchrome_40aPjP0xnC.png?generation=1594831383395100&amp;alt=media\" alt=\"\"></p>\n\n<p>I'm also watch upper Image augmentation code.\nSo How to define that code? \nYou've experimented with each function multiple times?</p>",
      "rawMarkdown": "1. When I found a code for Image classification.\nI'm almost watch this lr code.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2701710%2F3d97f637694af75183694605a893bdf7%2Fchrome_8saxN1jIJW.png?generation=1594831058076840&amp;alt=media)\n\nBut as far as I know, keras supports the following code, why use the first code and what are the advantages?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2701710%2Fa5ab1f7c18743046ab37f6708eda7e2a%2Fpycharm64_PTH9thaOCX.png?generation=1594831191354691&amp;alt=media)\n\n\n2. How to define Image augmentation code?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2701710%2F3e82a8e5d082774f544a706ea59eb060%2Fchrome_40aPjP0xnC.png?generation=1594831383395100&amp;alt=media)\n\nI'm also watch upper Image augmentation code.\nSo How to define that code? \nYou've experimented with each function multiple times?",
      "votes": null
    },
    {
      "id": "932853",
      "postDate": "07/17/2020 10:52:35",
      "content": "<p>keras fit method also supports callback</p>",
      "rawMarkdown": "keras fit method also supports callback",
      "votes": null
    },
    {
      "id": "932980",
      "postDate": "07/17/2020 12:05:15",
      "content": "<p>Oh. I mean that Keras is already made second picture code , But Many people used first picture code when they want to scheduling learining rate.<br>\nSo I wondering why people used second picture code.</p>",
      "rawMarkdown": "Oh. I mean that Keras is already made second picture code , But Many people used first picture code when they want to scheduling learining rate.\nSo I wondering why people used second picture code.",
      "votes": null
    },
    {
      "id": "933003",
      "postDate": "07/17/2020 12:17:42",
      "content": "<ul>\n<li><strong>LearningRateScheduler</strong>— The learning rate will be modified whenever a new epoch starts (based on a function).</li>\n<li><strong>ReduceLROnPlateau</strong>— When a specific metric stop improving, decrease the learning rate.\n<a href=\"https://towardsdatascience.com/tensorflow-learn-how-to-use-callbacks-efficiently-b13e0df89de3\"><strong>check here</strong></a></li>\n</ul>",
      "rawMarkdown": "* **LearningRateScheduler**— The learning rate will be modified whenever a new epoch starts (based on a function).\n* **ReduceLROnPlateau**— When a specific metric stop improving, decrease the learning rate.\n[**check here**](https://towardsdatascience.com/tensorflow-learn-how-to-use-callbacks-efficiently-b13e0df89de3)",
      "votes": null
    },
    {
      "id": "933051",
      "postDate": "07/17/2020 13:00:35",
      "content": "<p>You can use multiple callbacks and <a href=\"/vatsalparsaniya\">@vatsalparsaniya</a> already mentioned use of each thing</p>",
      "rawMarkdown": "You can use multiple callbacks and @vatsalparsaniya already mentioned use of each thing",
      "votes": null
    },
    {
      "id": "933168",
      "postDate": "07/17/2020 14:35:53",
      "content": "<p>ThankYou!!!!!</p>",
      "rawMarkdown": "ThankYou!!!!!",
      "votes": null
    },
    {
      "id": "933247",
      "postDate": "07/17/2020 15:38:27",
      "content": "<p>In you pictures above, the biggest diffence between the custom <code>LearningRateScheduler</code> and <code>ReduceOnPlateau</code> is that ROP trains the first epoch with a large learning rate and then decreases it as time goes on. With LRS, the first epoch has a small learning rate then it slowly increases until it reaches a maximum in epoch 5, then it decreases as time goes on.</p>\n\n<p>So you see the LRS has a \"ramp-up\" phase before \"decay\" phase. And LOR only has a \"decay\" phase. In many transfer learning tasks, \"ramp-up\" helps.</p>",
      "rawMarkdown": "In you pictures above, the biggest diffence between the custom `LearningRateScheduler` and `ReduceOnPlateau` is that ROP trains the first epoch with a large learning rate and then decreases it as time goes on. With LRS, the first epoch has a small learning rate then it slowly increases until it reaches a maximum in epoch 5, then it decreases as time goes on.\n\nSo you see the LRS has a \"ramp-up\" phase before \"decay\" phase. And LOR only has a \"decay\" phase. In many transfer learning tasks, \"ramp-up\" helps.",
      "votes": null
    },
    {
      "id": "933252",
      "postDate": "07/17/2020 15:41:51",
      "content": "<p>If you search public notebooks, i think there is an example of ImageDataGenerator. It is my understanding that you can use that with TensorFlow GPU but you cannot use that with TensorFlow TPU</p>",
      "rawMarkdown": "If you search public notebooks, i think there is an example of ImageDataGenerator. It is my understanding that you can use that with TensorFlow GPU but you cannot use that with TensorFlow TPU",
      "votes": null
    },
    {
      "id": "936327",
      "postDate": "07/20/2020 06:31:52",
      "content": "<p>Thanks your comment.<br>\nCan you give me  a example of ImageGenerator link?<br>\nI'm searching now, but I can't find proper public notebook.</p>",
      "rawMarkdown": "Thanks your comment.\nCan you give me  a example of ImageGenerator link?\nI'm searching now, but I can't find proper public notebook.",
      "votes": null
    },
    {
      "id": "936350",
      "postDate": "07/20/2020 06:49:10",
      "content": "<p>LRS prevents early training from blowing up and destroying weights when you start with a pre trained model.  Fastai showed the benefits of this type of schedule in version 1 a couple of years ago for pytorch - the code your seeing is a nice implementation of that for tensorflow models.</p>\n\n<p>If your not using pre trained weights or freezing the full pre trained model than not sure it has any benefits and reduceonplateau might get you done in fewer epochs.</p>",
      "rawMarkdown": "LRS prevents early training from blowing up and destroying weights when you start with a pre trained model.  Fastai showed the benefits of this type of schedule in version 1 a couple of years ago for pytorch - the code your seeing is a nice implementation of that for tensorflow models.\n\nIf your not using pre trained weights or freezing the full pre trained model than not sure it has any benefits and reduceonplateau might get you done in fewer epochs.",
      "votes": null
    },
    {
      "id": "940067",
      "postDate": "07/22/2020 16:53:49",
      "content": "<p>Is there any paper that describes the same ?</p>\n\n<p><a href=\"/pcjimmmy\">@pcjimmmy</a> </p>",
      "rawMarkdown": "Is there any paper that describes the same ?\n\n@pcjimmmy",
      "votes": null
    },
    {
      "id": "940094",
      "postDate": "07/22/2020 17:11:10",
      "content": "<p><a href=\"https://www.kaggle.com/nandanam\">Nandanam</a></p>\n\n<p>I learned about the method doing <a href=\"https://course.fast.ai/\">Jermery Howards fastai</a> v1 training.  My recollection was that he had developed it - and that he's not a paper writing kind of guy - but at my age it takes a lot of epochs for data to be firmly and correctly added to the recollection site.  The scheduler was a one liner in fastai and I never looked under the hood at the code.  But the LR curve looks similar so my assumption is that LRS is an adoption.  </p>\n\n<p>In the recollection site the one liner to use the schedule was <a href=\"https://docs.fast.ai/basic_train.html#fit_one_cycle\">fit_one_cycle.</a>.</p>\n\n<p>That link takes you to a <a href=\"https://docs.fast.ai/callbacks.one_cycle.html#What-is-1cycle?\">description page</a> if you follow the crumbs where a <a href=\"https://arxiv.org/pdf/1803.09820.pdf\">paper</a> is referenced.</p>",
      "rawMarkdown": "[Nandanam](https://www.kaggle.com/nandanam)\n\nI learned about the method doing [Jermery Howards fastai](https://course.fast.ai/) v1 training.  My recollection was that he had developed it - and that he's not a paper writing kind of guy - but at my age it takes a lot of epochs for data to be firmly and correctly added to the recollection site.  The scheduler was a one liner in fastai and I never looked under the hood at the code.  But the LR curve looks similar so my assumption is that LRS is an adoption.  \n\nIn the recollection site the one liner to use the schedule was [fit_one_cycle.](https://docs.fast.ai/basic_train.html#fit_one_cycle).\n\nThat link takes you to a [description page](https://docs.fast.ai/callbacks.one_cycle.html#What-is-1cycle?) if you follow the crumbs where a [paper](https://arxiv.org/pdf/1803.09820.pdf) is referenced.",
      "votes": null
    },
    {
      "id": "940662",
      "postDate": "07/23/2020 04:32:53",
      "content": "<p>I thought that approach was just used to find the optimal learning rate. Did not notice the part about the increase and decrease part.</p>\n\n<p>Thanks for the response.</p>",
      "rawMarkdown": "I thought that approach was just used to find the optimal learning rate. Did not notice the part about the increase and decrease part.\n\nThanks for the response.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 932853,
      "author_name": "ajax0564",
      "author_url": "",
      "post_date": "07/17/2020 10:52:35",
      "content": "<p>keras fit method also supports callback</p>",
      "votes": null,
      "replies": [
        {
          "id": 932980,
          "author_name": "zxzxs9182",
          "author_url": "",
          "post_date": "07/17/2020 12:05:15",
          "content": "<p>Oh. I mean that Keras is already made second picture code , But Many people used first picture code when they want to scheduling learining rate.<br>\nSo I wondering why people used second picture code.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 933051,
          "author_name": "kurianbenoy",
          "author_url": "",
          "post_date": "07/17/2020 13:00:35",
          "content": "<p>You can use multiple callbacks and <a href=\"/vatsalparsaniya\">@vatsalparsaniya</a> already mentioned use of each thing</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 933003,
      "author_name": "vatsalparsaniya",
      "author_url": "",
      "post_date": "07/17/2020 12:17:42",
      "content": "<ul>\n<li><strong>LearningRateScheduler</strong>— The learning rate will be modified whenever a new epoch starts (based on a function).</li>\n<li><strong>ReduceLROnPlateau</strong>— When a specific metric stop improving, decrease the learning rate.\n<a href=\"https://towardsdatascience.com/tensorflow-learn-how-to-use-callbacks-efficiently-b13e0df89de3\"><strong>check here</strong></a></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 933168,
          "author_name": "zxzxs9182",
          "author_url": "",
          "post_date": "07/17/2020 14:35:53",
          "content": "<p>ThankYou!!!!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 933247,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "07/17/2020 15:38:27",
      "content": "<p>In you pictures above, the biggest diffence between the custom <code>LearningRateScheduler</code> and <code>ReduceOnPlateau</code> is that ROP trains the first epoch with a large learning rate and then decreases it as time goes on. With LRS, the first epoch has a small learning rate then it slowly increases until it reaches a maximum in epoch 5, then it decreases as time goes on.</p>\n\n<p>So you see the LRS has a \"ramp-up\" phase before \"decay\" phase. And LOR only has a \"decay\" phase. In many transfer learning tasks, \"ramp-up\" helps.</p>",
      "votes": null,
      "replies": [
        {
          "id": 936327,
          "author_name": "zxzxs9182",
          "author_url": "",
          "post_date": "07/20/2020 06:31:52",
          "content": "<p>Thanks your comment.<br>\nCan you give me  a example of ImageGenerator link?<br>\nI'm searching now, but I can't find proper public notebook.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 936350,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "07/20/2020 06:49:10",
          "content": "<p>LRS prevents early training from blowing up and destroying weights when you start with a pre trained model.  Fastai showed the benefits of this type of schedule in version 1 a couple of years ago for pytorch - the code your seeing is a nice implementation of that for tensorflow models.</p>\n\n<p>If your not using pre trained weights or freezing the full pre trained model than not sure it has any benefits and reduceonplateau might get you done in fewer epochs.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 940067,
          "author_name": "nandanam",
          "author_url": "",
          "post_date": "07/22/2020 16:53:49",
          "content": "<p>Is there any paper that describes the same ?</p>\n\n<p><a href=\"/pcjimmmy\">@pcjimmmy</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 940094,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "07/22/2020 17:11:10",
          "content": "<p><a href=\"https://www.kaggle.com/nandanam\">Nandanam</a></p>\n\n<p>I learned about the method doing <a href=\"https://course.fast.ai/\">Jermery Howards fastai</a> v1 training.  My recollection was that he had developed it - and that he's not a paper writing kind of guy - but at my age it takes a lot of epochs for data to be firmly and correctly added to the recollection site.  The scheduler was a one liner in fastai and I never looked under the hood at the code.  But the LR curve looks similar so my assumption is that LRS is an adoption.  </p>\n\n<p>In the recollection site the one liner to use the schedule was <a href=\"https://docs.fast.ai/basic_train.html#fit_one_cycle\">fit_one_cycle.</a>.</p>\n\n<p>That link takes you to a <a href=\"https://docs.fast.ai/callbacks.one_cycle.html#What-is-1cycle?\">description page</a> if you follow the crumbs where a <a href=\"https://arxiv.org/pdf/1803.09820.pdf\">paper</a> is referenced.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 940662,
          "author_name": "nandanam",
          "author_url": "",
          "post_date": "07/23/2020 04:32:53",
          "content": "<p>I thought that approach was just used to find the optimal learning rate. Did not notice the part about the increase and decrease part.</p>\n\n<p>Thanks for the response.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 933252,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "07/17/2020 15:41:51",
      "content": "<p>If you search public notebooks, i think there is an example of ImageDataGenerator. It is my understanding that you can use that with TensorFlow GPU but you cannot use that with TensorFlow TPU</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "930694": "1. When I found a code for Image classification.\nI'm almost watch this lr code.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2701710%2F3d97f637694af75183694605a893bdf7%2Fchrome_8saxN1jIJW.png?generation=1594831058076840&amp;alt=media)\n\nBut as far as I know, keras supports the following code, why use the first code and what are the advantages?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2701710%2Fa5ab1f7c18743046ab37f6708eda7e2a%2Fpycharm64_PTH9thaOCX.png?generation=1594831191354691&amp;alt=media)\n\n\n2. How to define Image augmentation code?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2701710%2F3e82a8e5d082774f544a706ea59eb060%2Fchrome_40aPjP0xnC.png?generation=1594831383395100&amp;alt=media)\n\nI'm also watch upper Image augmentation code.\nSo How to define that code? \nYou've experimented with each function multiple times?",
    "932853": "keras fit method also supports callback",
    "932980": "Oh. I mean that Keras is already made second picture code , But Many people used first picture code when they want to scheduling learining rate.\nSo I wondering why people used second picture code.",
    "933003": "* **LearningRateScheduler**— The learning rate will be modified whenever a new epoch starts (based on a function).\n* **ReduceLROnPlateau**— When a specific metric stop improving, decrease the learning rate.\n[**check here**](https://towardsdatascience.com/tensorflow-learn-how-to-use-callbacks-efficiently-b13e0df89de3)",
    "933051": "You can use multiple callbacks and @vatsalparsaniya already mentioned use of each thing",
    "933168": "ThankYou!!!!!",
    "933247": "In you pictures above, the biggest diffence between the custom `LearningRateScheduler` and `ReduceOnPlateau` is that ROP trains the first epoch with a large learning rate and then decreases it as time goes on. With LRS, the first epoch has a small learning rate then it slowly increases until it reaches a maximum in epoch 5, then it decreases as time goes on.\n\nSo you see the LRS has a \"ramp-up\" phase before \"decay\" phase. And LOR only has a \"decay\" phase. In many transfer learning tasks, \"ramp-up\" helps.",
    "933252": "If you search public notebooks, i think there is an example of ImageDataGenerator. It is my understanding that you can use that with TensorFlow GPU but you cannot use that with TensorFlow TPU",
    "936327": "Thanks your comment.\nCan you give me  a example of ImageGenerator link?\nI'm searching now, but I can't find proper public notebook.",
    "936350": "LRS prevents early training from blowing up and destroying weights when you start with a pre trained model.  Fastai showed the benefits of this type of schedule in version 1 a couple of years ago for pytorch - the code your seeing is a nice implementation of that for tensorflow models.\n\nIf your not using pre trained weights or freezing the full pre trained model than not sure it has any benefits and reduceonplateau might get you done in fewer epochs.",
    "940067": "Is there any paper that describes the same ?\n\n@pcjimmmy",
    "940094": "[Nandanam](https://www.kaggle.com/nandanam)\n\nI learned about the method doing [Jermery Howards fastai](https://course.fast.ai/) v1 training.  My recollection was that he had developed it - and that he's not a paper writing kind of guy - but at my age it takes a lot of epochs for data to be firmly and correctly added to the recollection site.  The scheduler was a one liner in fastai and I never looked under the hood at the code.  But the LR curve looks similar so my assumption is that LRS is an adoption.  \n\nIn the recollection site the one liner to use the schedule was [fit_one_cycle.](https://docs.fast.ai/basic_train.html#fit_one_cycle).\n\nThat link takes you to a [description page](https://docs.fast.ai/callbacks.one_cycle.html#What-is-1cycle?) if you follow the crumbs where a [paper](https://arxiv.org/pdf/1803.09820.pdf) is referenced.",
    "940662": "I thought that approach was just used to find the optimal learning rate. Did not notice the part about the increase and decrease part.\n\nThanks for the response."
  },
  "source": "meta"
}